Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/95326
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Electronic and Information Engineering | en_US |
| dc.creator | Yi, L | en_US |
| dc.creator | Mak, MW | en_US |
| dc.date.accessioned | 2022-09-19T01:59:41Z | - |
| dc.date.available | 2022-09-19T01:59:41Z | - |
| dc.identifier.issn | 2162-237X | en_US |
| dc.identifier.uri | http://hdl.handle.net/10397/95326 | - |
| dc.language.iso | en | en_US |
| dc.publisher | Institute of Electrical and Electronics Engineers | en_US |
| dc.rights | © 2020 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. | en_US |
| dc.rights | The following publication L. Yi and M. -W. Mak, "Improving Speech Emotion Recognition With Adversarial Data Augmentation Network," in IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 1, pp. 172-184, Jan. 2022 is available at https://doi.org/10.1109/TNNLS.2020.3027600. | en_US |
| dc.subject | Data augmentation | en_US |
| dc.subject | Generative adversarial networks (GANs) | en_US |
| dc.subject | Speech emotion recognition | en_US |
| dc.subject | Wasserstein divergence | en_US |
| dc.title | Improving speech emotion recognition with adversarial data augmentation network | en_US |
| dc.type | Journal/Magazine Article | en_US |
| dc.identifier.spage | 172 | en_US |
| dc.identifier.epage | 184 | en_US |
| dc.identifier.volume | 33 | en_US |
| dc.identifier.issue | 1 | en_US |
| dc.identifier.doi | 10.1109/TNNLS.2020.3027600 | en_US |
| dcterms.abstract | When training data are scarce, it is challenging to train a deep neural network without causing the overfitting problem. For overcoming this challenge, this article proposes a new data augmentation network-namely adversarial data augmentation network (ADAN)-based on generative adversarial networks (GANs). The ADAN consists of a GAN, an autoencoder, and an auxiliary classifier. These networks are trained adversarially to synthesize class-dependent feature vectors in both the latent space and the original feature space, which can be augmented to the real training data for training classifiers. Instead of using the conventional cross-entropy loss for adversarial training, the Wasserstein divergence is used in an attempt to produce high-quality synthetic samples. The proposed networks were applied to speech emotion recognition using EmoDB and IEMOCAP as the evaluation data sets. It was found that by forcing the synthetic latent vectors and the real latent vectors to share a common representation, the gradient vanishing problem can be largely alleviated. Also, results show that the augmented data generated by the proposed networks are rich in emotion information. Thus, the resulting emotion classifiers are competitive with state-of-The-Art speech emotion recognition systems. | en_US |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | IEEE transactions on neural networks and learning systems, Jan. 2022, v. 33, no. 1, p. 172-184 | en_US |
| dcterms.isPartOf | IEEE transactions on neural networks and learning systems | en_US |
| dcterms.issued | 2022-01 | - |
| dc.identifier.scopus | 2-s2.0-85116598107 | - |
| dc.identifier.pmid | 33035171 | - |
| dc.identifier.eissn | 2162-2388 | en_US |
| dc.description.validate | 202209 bcvc | en_US |
| dc.description.oa | Accepted Manuscript | en_US |
| dc.identifier.FolderNumber | RGC-B2-0262, a1720 | - |
| dc.identifier.SubFormID | 45835 | - |
| dc.description.fundingSource | RGC | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.description.oaCategory | Green (AAM) | en_US |
| Appears in Collections: | Journal/Magazine Article | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| T-NNLS-emotion.pdf | Pre-Published version | 10.19 MB | Adobe PDF | View/Open |
Page views
96
Last Week
0
0
Last month
Citations as of Apr 14, 2025
Downloads
464
Citations as of Apr 14, 2025
SCOPUSTM
Citations
93
Citations as of Dec 19, 2025
WEB OF SCIENCETM
Citations
82
Citations as of Dec 18, 2025
Google ScholarTM
Check
Altmetric
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



