Ingeniería Industrial · 2025
Caracterización de información con el apoyo de herramientas Web Scraping y PLN sobre educación e industria publicados en medios digitales para las ciudades de Medellín y Neiva
Solo ficha. El repositorio marca este trabajo como «Acceso abierto»: el PDF no es público, así que aquí no hay texto para leer. Ver la ficha en el repositorio.
Resumen
This project addresses the lack of efficient systems to analyze higher education access needs and industrial development trends in Medellín and Neiva. The objective is to characterize data from both sectors to develop a sentiment analysis for future strategic business studies for the Universidad Antonio Nariño (UAN). The methodology includes identifying key digital sources, real-time data extraction using Web Scraping (using Scrapy and BeautifulSoup), and applying the Sentiment VADER model of Natural Language Processing (NLP tool). Expected results include generating quantifiable insights into the emotional polarity of the topics and mapping contextualized strategic opportunities. Preliminary findings identify barriers such as dropout, lack of inclusion (Medellín), and economic and industrial standardization limitations (Neiva).