Sporala red del conocimiento
Universidad de Antioquia

Maestrías de la Facultad de Ingeniería · 2022

Automatic personality estimation from text in two languages using natural language processing techniques

López Pabón, Felipe OrlandoAsesor: Orozco Arroyave, Juan Rafael

ABSTRACT : According to the literature review, there is a lack of works in the field of automatic personality recognition based on machine learning or deep learning models using word embeddings that do not rely on dictionaries or lexicons that are highly dependent on human intervention. Similarly, there is a lack of works using texts in a language other than English. This study is focused on the use of natural language processing techniques that allow the extraction of word embeddings that are useful for the estimation of the 5 personality traits defined in the OCEAN model of psychology (the Emotional Stability trait is considered instead of the Neuroticism trait) in two languages: English and Spanish. The main contribution of this work includes: 1) different experiments are explored: i) regression methods, to predict the personality scores on the traits, classification methods, such as ii) two-class classification: weak vs. strong presence of each trait, and iii) three-class classification: low vs. medium vs. high presence of each trait; 2) implementation of word embeddings based on classical methods: Word2Vec and GloVe as well as word embeddings based on state-of-the-art methods such as BERT and BETO to train machine learning methods; 3) use of deep learning methods that allow to extract word embeddings from texts and train the embeddings layer from scratch or use pre-trained embeddings to improve the performance of the architectures; and 4) evaluation of the different methods in Spanish language taking into account text signals coming from YouTube and Twitter.

Texto completo 161 páginas con texto

Leer la tesis completa Ficha en el repositorio

Contenido

Tesauro de su biblioteca