Maestrías de la Facultad de Ingeniería · 2022
Robust Automatic Speech Recognition
ABSTRACT : In contact center organizations, customer satisfaction (CS) analysis is an important issue since the organization's reputation is strongly impacted by the customer's perception of the quality of service (QoS) provided. Service agents must provide exceptional service for the continual growth of the organization in today's dynamic market. In order to improve the service, these companies have human experts to evaluate the QoS. This practice is commonly based on the customer's opinion of the service after the conversation with the agent. However, such practice has two main disadvantages: 1) double cost and effort, i.e., human experts are needed to answer the calls as well as to evaluate them; and 2) only a small sample of the total number of calls is rated due to human limitations. Given these difficulties, these organizations have promoted research into the development of different systems based on acoustic and linguistic analyses that help to automatically evaluate CS. The acoustic-based system detects abnormal changes on the speech signal such as: poorly-articulated speech, increase in speech rate, increase in voice volume, and others. The linguistic-based system searches for keywords that reflect satisfaction/dissatisfaction. This approach requires an Automatic Speech Recognition (ASR) system to convert the speech signal into a text transcriptions. The ASR system must be designed in such a way that its performance is minimally dependent of the acoustic conditions. This thesis proposes a methodology to robustly recognize speech in non-controlled acoustic conditions using recordings collected by a call center. It also proposes a methodology to recognize emotion from speech and to evaluate CS based on acoustic and linguistic analysis. The acoustic features include articulation, prosody and phonation features. The linguistic features consist of word embeddings extracted from the transcriptions generated by the proposed ASR system. Deep learning approaches are considered for both speech recognition and CS evaluation and they are compared with traditional techniques.