Maestrías de la Facultad de Ingeniería · 2022
Automatic personality estimation from text in two languages using natural language processing techniques
ABSTRACT : According to the literature review, there is a lack of works in the field of automatic personality recognition based on machine learning or deep learning models using word embeddings that do not rely on dictionaries or lexicons that are highly dependent on human intervention. Similarly, there is a lack of works using texts in a language other than English. This study is focused on the use of natural language processing techniques that allow the extraction of word embeddings that are useful for the estimation of the 5 personality traits defined in the OCEAN model of psychology (the Emotional Stability trait is considered instead of the Neuroticism trait) in two languages: English and Spanish. The main contribution of this work includes: 1) different experiments are explored: i) regression methods, to predict the personality scores on the traits, classification methods, such as ii) two-class classification: weak vs. strong presence of each trait, and iii) three-class classification: low vs. medium vs. high presence of each trait; 2) implementation of word embeddings based on classical methods: Word2Vec and GloVe as well as word embeddings based on state-of-the-art methods such as BERT and BETO to train machine learning methods; 3) use of deep learning methods that allow to extract word embeddings from texts and train the embeddings layer from scratch or use pre-trained embeddings to improve the performance of the architectures; and 4) evaluation of the different methods in Spanish language taking into account text signals coming from YouTube and Twitter.
Texto completo 161 páginas con texto
Leer la tesis completa Ficha en el repositorio