Speech emotion recognition using transfer learning methods

Loading...
Thumbnail Image

Supplementary material

Other Title

Author ORCID Profiles (clickable)

Degree

Master of Applied Technologies (Computing)

Grantor

Unitec, Te Pūkenga – New Zealand Institute of Skills and Technology

Date

Ngā Upoko Tukutuku (Māori subject headings)

Keyword

emotions
speech analysis
speech processing systems
transfer learning (machine learning)
data augmentation (machine learning)
machine learning

Citation

Kumarage, L.V.A. (2025). Speech emotion recognition using transfer learning methods (Unpublished document submitted in partial fulfilment of the requirements for the degree of Master of Applied Technologies (Computing)). Unitec, Te Pūkenga - New Zealand Institute of Skills and Technology https://hdl.handle.net/10652/6818

Abstract

RESEARCH QUESTIONS •What are the current limitations of the existing speech-emotion recognition models? • How data augmentation methods can be used to generate synthetic data that could improve the performance of SER systems? • How to use the Transfer learning methods with a fine-tuning process to utilize its benefits? ABSTRACT Identifying human emotions using machine learning techniques in this technological advancement era is more popular, and it offers practical contributions to different fields such as health and medicine. Identifying human emotions can be made possible by analyzing their facial expressions, body gestures, and speech. With the advancement of automatic speech recognition, emotion recognition using speech is more effective and accurate than other methods. Most existing systems trained with insufficient amounts of data face the issue of performing effectively when confronted with unseen data. These systems mostly used traditional machine learning techniques. This thesis focuses on identifying speech emotions using transfer learning, which uses pre-trained models with large-scale data. Therefore, due to the lessening of training duration and the computational cost, the efficiency of the process of speech emotion recognition will increase. Combining several datasets and using augmented data can help the existing limitation of lack of data for training.

Publisher

Link to ePress publication

DOI

Copyright holder

Author

Copyright notice

All rights reserved

Copyright license

Available online at