Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Rethinking Auditory Affective Descriptors Through Zero-Shot Emotion Recognition in Speech

Bibliographic Data

ID22107293
AuthorsXinzhou Xu (0000-0002-4017-5919, Nanjing University of Posts and Telecommunications), Jun Deng (0000-0002-8916-6503, Agile Robots AG, Munich, Germany), Zixing Zhang (0000-0001-8487-0561, Hunan University), Xijian Fan (0000-0002-7017-7667, Nanjing Forestry University), Li Zhao (0009-0004-7836-6878, Southeast University), Laurence Devillers (0000-0001-9894-172X, Centre National de la Recherche Scientifique), Björn W Schuller (0000-0002-6478-8699, University of Augsburg)
Year2022
Volume9
Issue5
Pages1530-1541
Publication date2022-10-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueIEEE Transactions on Computational Social Systems (JOURNAL)
Journal identifiersISSN: 2329-924X • E-ISSN: 2373-7476
PublisherInstitute of Electrical and Electronics Engineers (IEEE) (PUBLISHER)
DOI10.1109/tcss.2021.3130401
OpenAlexW4205734456
LanguageEN
Citations received3
References cited65

Zero-shot speech emotion recognition (SER) endows machines with the ability of sensing unseen-emotional states in speech, compared with conventional SER endeavors on supervised cases. On addressing the zero-shot SER task, auditory affective descriptors (AADs) are typically employed to transfer affective knowledge from seen- to unseen-emotional states. However, it remains unknown which types of AADs can well describe emotional states in speech during the transfer. In this regard, we define and research on three types of AADs, namely, per-emotion semantic-embedding, per-emotion manually annotated, and per-sample manually annotated AADs, through zero-shot emotion recognition in speech. This leads to a systematic design including prototype- and annotation-based zero-shot SER modules, relying on the input from per-emotion and per-sample AADs, respectively. We then perform extensive experimental comparisons between human and machines’ AADs on the French emotional speech corpus CINEMO for positive-negative (PN) and within-negative (WN) tasks. The experimental results indicate that semantic-embedding prototypes from pretrained models can outperform manually annotated emotional dimensions in zero-shot SER. The results further demonstrate that it is possible for machines to understand and describe affective information in speech better than human beings, with the help of sufficient pretrained models

Cognitive psychology · Emotion recognition · Linguistics · Speech recognition · Computer Science · Emotion and Mood Recognition · Psychology · Speech and Audio Processing · Speech Recognition and Synthesis

  • RobinNet

    Open Access•Yash Khurana, Swamita Gupta et al.•IEEE Transactions on Computational…•2024

  • Piezoelectric Touch Sensing and Random-Forest-Based Technique for Emotion Recognition

    Open Access•Yuqing Qi, Weichen Jia et al.•IEEE Transactions on Computational…•2024

  • SIA-Net

    Open Access•Shuzhen Li, Tong Zhang et al.•IEEE Transactions on Computational…•2024

  • Enriching Word Vectors with Subword Information

    Open Access•Piotr Bojanowski, Edouard Grave et al.•Transactions of the Association…•2017

  • Toward Artificial Emotional Intelligence for Cooperative Social Human–Machine Interaction

    Open Access•Berat A Erol, Abhijit Majumdar et al.•IEEE Transactions on Computational…•2020

  • A Review on Five Recent and Near-Future Developments in Computational Processing of Emotion in the Human Voice

    Open Access•Dagmar Schuller, Dagmar M Schuller et al.•Emotion Review•2021

  • Negative emotions in informal feedback

    Open Access•Genevieve Fuji Johnson, Genevieve Johnson et al.•Human Relations•2014

  • Mapping 24 emotions conveyed by brief human vocalization

    Alan S Cowen, Hillary Anger Elfenbein et al.•American Psychologist•2018

  • Emotion knowledge

    P R Shaver, Phillip Shaver et al.•Journal of Personality and Social…•1987

Unique citing works3
Citations per year1,5
Citation span2024 - 2024 (1)
Citation velocityrecent
Highly citedNo
Citation typesNeutral: 3

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae