Sporala red del conocimiento
Universidad Antonio Nariño

Ingeniería Industrial · 2025

Caracterización de información con el apoyo de herramientas Web Scraping y PLN sobre educación e industria publicados en medios digitales para las ciudades de Medellín y Neiva

Realpe Rios, Nicol Valentina · Vargas Rojas, Carlos AndresAsesor: Martínez, Yenny Alexandra · Delgadillo Leguizamón, David Santiago

Solo ficha. El repositorio marca este trabajo como «Acceso abierto»: el PDF no es público, así que aquí no hay texto para leer. Ver la ficha en el repositorio.

Resumen

This project addresses the lack of efficient systems to analyze higher education access needs and industrial development trends in Medellín and Neiva. The objective is to characterize data from both sectors to develop a sentiment analysis for future strategic business studies for the Universidad Antonio Nariño (UAN). The methodology includes identifying key digital sources, real-time data extraction using Web Scraping (using Scrapy and BeautifulSoup), and applying the Sentiment VADER model of Natural Language Processing (NLP tool). Expected results include generating quantifiable insights into the emotional polarity of the topics and mapping contextualized strategic opportunities. Preliminary findings identify barriers such as dropout, lack of inclusion (Medellín), and economic and industrial standardization limitations (Neiva).

Palabras clave