Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

A digital “flat affect”? Popular speech compression codecs and their effects on emotional prosody

Bibliographic Data

ID22092467
AuthorsOliver Niebuhr (0000-0002-8623-1680, University of Southern Denmark, corresponding author), Ingo Siegert (0000-0001-7447-7141, Otto-von-Guericke-Universität Magdeburg, corresponding author)
Year2023
Volume8
Publication date2023-03-23
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueFrontiers in Communication (JOURNAL)
Journal identifiersISSN: 2297-900X • E-ISSN: 2297-900X
PublisherFrontiers Media SA (PUBLISHER • CH)
DOI10.3389/fcomm.2023.972182
OpenAlexW4360616662
LanguageEN
References cited73

Introduction Calls via video apps, mobile phones and similar digital channels are a rapidly growing form of speech communication. Such calls are not only— and perhaps less and less— about exchanging content, but about creating, maintaining, and expanding social and business networks. In the phonetic code of speech, these social and emotional signals are considerably shaped by (or encoded in) prosody. However, according to previous studies, it is precisely this prosody that is significantly distorted by modern compression codecs. As a result, the identification of emotions becomes blurred and can even be lost to the extent that opposing emotions like joy and anger or disgust and sadness are no longer differentiated on the recipients' side. The present study searches for the acoustic origins of these perceptual findings. Method A set of 108 sentences from the Berlin Database of Emotional Speech served as speech material in our study. The sentences were realized by professional actors (2m, 2f) with seven different emotions (neutral, fear, disgust, joy, boredom, anger, sadness) and acoustically analyzed in the original uncompressed (WAV) version and as well as in strongly compressed versions based on the four popular codecs AMR-WB, MP3, OPUS, and SPEEX. The analysis included 6 tonal (i.e. f0-related) and 7 non-tonal prosodic parameters (e.g., formants as well as acoustic-energy and spectral-slope estimates). Results Results show significant, codec-specific distortion effects on all 13 prosodic parameter measurements compared to the WAV reference condition. Means values of automatic measurement can, across sentences, deviate by up to 20% from the values of the WAV reference condition. Moreover, the effects go in opposite directions for tonal and non-tonal parameters. While tonal parameters are distorted by speech compression such that the acoustic differences between emotions are increased, compressing non-tonal parameters make the acoustic-prosodic profiles of emotions more similar to each other, particularly under MP3 and SPEEX compression. Discussion The term “flat affect” comes from the medical field and describes a person's inability to express or display emotions. So, does strong compression of emotional speech create a “digital flat affect”? The answer to this question is a conditional “yes”. We provided clear evidence for a “digital flat affect”. However, it seems less strongly pronounced in the present acoustic measurements than in previous perception data, and it manifests itself more strongly in non-tonal than in tonal parameters. We discuss the practical implications of our findings for the everyday use of digital communication devices and critically reflect on the generalizability of our findings, also with respect to their origins in the codecs' inner mechanics

Codec · Physics · Prosody · Speech recognition · Telecommunications · Communication · Computer Science · Digital Communication and Language · Emotion and Mood Recognition · Phonetics and Phonology Research · Psychology

  • Fundamental Aspects in the Perception of f0

    Oliver Niebuhr, Henning Reetz et al.•Oxford Handbook of Language Prosody•2020

  • Funders' positive affective reactions to entrepreneurs' crowdfunding pitches

    Open Access•Blakley C Davis, Keith M Hmieleski et al.•Journal of Business Venturing•2017

  • The role of voice quality in communicating emotion, mood and attitude

    Open Access•Christer Gobl•Speech Communication•2003

  • Age, sex, and vowel dependencies of acoustic measures related to the voice source

    Markus Iseli, Yen-Liang Shue et al.•The Journal of the Acoustical…•2007

  • Can Charisma Be Taught? Tests of Two Interventions

    John Antonakis, Marika Fenley et al.•Academy of Management Learning &…•2011

  • Emotion recognition and confidence ratings predicted by vocal stimulus type and prosodic parameters

    Open Access•Adi Lausen, Kurt Hammerschmidt•Humanities and Social Sciences…•2020

  • Case Report

    Open Access•Ingo Siegert, Oliver Niebuhr•Frontiers in Communication•2021

  • Effect of charismatic signaling in social media settings

    Open Access•Benjamin Tur, Jennifer Harstad et al.•The Leadership Quarterly•2022

  • Multi-modal emotion expression and online charity crowdfunding success

    Open Access•Kexin Zhao, Lina Zhou et al.•Decision Support Systems•2022

  • Honest Signals

    Alex Pentland, Tracy Heibeck•Honest Signals•2008

  • Charisma

    John Antonakis, Nicolas Bastardoz et al.•Annual Review of Organizational…•2016

  • Empathy or perceived credibility? An empirical study on individual donation behavior in charitable crowdfunding

    Open Access•Lili Liu, Ayoung Suh et al.•Internet Research•2018

  • Speech Melody Matters—How Robots Profit from Using Charismatic Speech

    Open Access•Kerstin Fischer, Oliver Niebuhr et al.•ACM Transactions on Human-Robot…•2019

  • Embodying “tech

    Open Access•Teresa Pratt•Journal of Sociolinguistics•2020

  • Spectral Analysis of Candidates' Nonverbal Vocal Communication

    Stanford W Gregory, Stanford W Gregory Jr et al.•Social Psychology Quarterly•2002

  • Processes of Opinion Change

    Herbert C Kelman•Public Opinion Quarterly•1961

  • Acoustic profiles in vocal emotion expression

    R Banse, Klaus R Scherer•Journal of Personality and Social…•1996

  • Formant frequencies and body size of speaker

    Open Access•Julio Gonzalez•Journal of Phonetics•2004

  • Phonetic differences between male and female speech

    Open Access•Adrian P Simpson, Adrian Simpson•Language and Linguistics Compass•2009

  • Articulatory-acoustic relationships during vocal tract growth for French vowels

    Open Access•Lucie Ménard, Jana L Schwartz et al.•Journal of Phonetics•2007

Citation velocityhistorical
Highly citedNo

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae