Sporala red del conocimiento
Universidad de Antioquia

Maestrías de la Facultad de Ingeniería · 2022

Demographic information retrieval from text for subject characterization and market segmentation

Escobar Grisales, DanielAsesor: Orozco Arroyave, Juan Rafael · Vásquez Correa, Juan Camilo

ABSTRACT : In recent years, the most important trends to improve customer services in the e-commerce industry are focused on customer customization and the use of automated dialogue systems to enhance the support experience. On one hand, demographic traits from a subject/customer such as gender, nationality, and age can help to strengthen marketing strategies or even improve customer empathy with the product or the advisor. On the other hand, the service of automated dialogue can help to improve the ability to serve multiple users. However, in order to improve the customer service support, the dialogue system should correctly recognize the customer requirements. This research work aims to improve customer services based on text data from the subject/customer, considering both scenarios, demographic trait recognition, and evaluation of effectiveness in conversations between humans and chatbots. For demographic trait recognition, this work proposes the use of recurrent and convolutional neural networks and a transfer learning strategy to recognize three demographic traits: gender, variety of language according to nationality, and age. Models are tested in two different document types, Tweets (documents written in informal language) and call-center conversations (documents written in formal language). In documents in informal language, accuracies of up to 75% and 92% are achieved for the recognition of gender and language variety, respectively, and an unweighted average recall of up to 50% is achieved for age recognition. In documents in formal language, accuracies of up to 70%, 72%, and 68% are achieved for the recognition of gender, variety language, and age respectively. Results indicate that for the traits of gender and variety of language it is possible to transfer the knowledge from a system trained on a specific type of expression to another, where the structure is completely different, and its amount of data is scarcer. In addition, the learning acquired by the models to recognize language varieties in Spanish-speaking countries can be successfully used to fine-tune models to recognize more subtle language varieties, such as the ones within the same country. For evaluation of effectiveness in conversations with chatbots, we pro- pose a new methodology for automatic evaluation of chatbot effectiveness in real production environments. The analysis considers convolutional neural networks, using two parallel convolutional layers to evaluate questions and answers independently. This methodology is tested upon real conversations of chatbots that provide service to two different companies. The results are compared to baseline models based on classical techniques with different pre-trained word embedding models. According to our results, the proposed approach provides accuracies between 78% and 80%, which outperforms the best result of the baseline models by 2.9%.

Texto completo 89 páginas con texto

Leer la tesis completa Ficha en el repositorio

Tesauro de su biblioteca