Speech emotion recognition using transfer learning methods
Loading...
Supplementary material
Other Title
Author ORCID Profiles (clickable)
Degree
Master of Applied Technologies (Computing)
Grantor
Unitec, Te Pūkenga – New Zealand Institute of Skills and Technology
Date
Supervisors
Type
Ngā Upoko Tukutuku (Māori subject headings)
Keyword
emotions
speech analysis
speech processing systems
transfer learning (machine learning)
data augmentation (machine learning)
machine learning
speech analysis
speech processing systems
transfer learning (machine learning)
data augmentation (machine learning)
machine learning
ANZSRC Field of Research Code (2020)
Citation
Kumarage, L.V.A. (2025). Speech emotion recognition using transfer learning methods (Unpublished document submitted in partial fulfilment of the requirements for the degree of Master of Applied Technologies (Computing)). Unitec, Te Pūkenga - New Zealand Institute of Skills and Technology
https://hdl.handle.net/10652/6818
Abstract
RESEARCH QUESTIONS
•What are the current limitations of the existing speech-emotion recognition models?
• How data augmentation methods can be used to generate synthetic data that could improve the performance of SER systems?
• How to use the Transfer learning methods with a fine-tuning process to utilize its benefits?
ABSTRACT
Identifying human emotions using machine learning techniques in this technological advancement era is more popular, and it offers practical contributions to different fields such as health and medicine. Identifying human emotions can be made possible by analyzing their facial expressions, body gestures, and speech. With the advancement of automatic speech recognition, emotion recognition using speech is more effective and accurate than other methods. Most existing systems trained with insufficient amounts of data face the issue of performing effectively when confronted with unseen data. These systems mostly used traditional machine learning techniques.
This thesis focuses on identifying speech emotions using transfer learning, which uses pre-trained models with large-scale data. Therefore, due to the lessening of training duration and the computational cost, the efficiency of the process of speech emotion recognition will increase. Combining several datasets and using augmented data can help the existing limitation of lack of data for training.
Publisher
Permanent link
Link to ePress publication
DOI
Copyright holder
Author
Copyright notice
All rights reserved
