The power of multimodality
Improving speech act classification through visual cues, audio intonation, and gaze tracking
Bibliographic Data
| ID | 21500664 |
|---|---|
| Authors | Ying Li (0000-0003-0678-9535, Guizhou University, corresponding author), Wari Wongwaropakorn (0000-0003-1263-3182, Walailak University) |
| Year | 2025 |
| Publication date | 2025-08-06 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Information Development (JOURNAL) |
| Journal identifiers | ISSN: 0266-6669 • E-ISSN: 1741-6469 |
| Publisher | SAGE Publications (PUBLISHER • US) |
| DOI | 10.1177/02666669251355940 |
| OpenAlex | W4413004302 |
| Language | EN |
| References cited | 33 |
Multimodal discourse analysis enhances precision and contextual understanding in speech acts by integrating modalities such as visual cues, text, non-verbal signals, and gaze tracking. This study explores the effectiveness of multimodal discourse in improving speech act classification through combined visual, auditory, and non-verbal data. A mixed-method approach was employed, involving quantitative data from 370 communication professionals analyzed using Statistical Package For Social Sciences (SPSS), alongside qualitative insights from interviews and focus groups. Findings indicate that visual cues significantly enhance speech act classification performance, while audio intonation improves accuracy under noisy conditions. The integration of text and non-verbal data further supports deeper contextual understanding, particularly benefiting indirect speech act recognition and overall multimodal fusion effectiveness. This study's holistic approach uniquely combines multiple modalities, visual, audio, text, and gaze tracking, surpassing previous research focused on isolated speech interpretation factors. Multimodality significantly improves accuracy and contextual comprehension in speech act classification, demonstrating that communication analysis should extend beyond textual content to include audio and non-verbal traits for a fuller understanding
Audio visual · Eye tracking · Gaze · Linguistics · Multimedia · Multimodality · Speech recognition · World Wide Web · Computer Science · Hearing Impairment and Communication · Language, Metaphor, and Cognition · Subtitles and Audiovisual Media · Artificial Intelligence
Semiotic modes accentuating learners’ metafunctions
Discursive de/humanizing
A multimodal discourse analysis of English dentistry texts written by Saudi undergraduate students
Developing multimodal communicative competence in emerging academic and professional genres
Signs of understanding and turns-as-actions
Do political cartoons and illustrations have their own specialized forms for warnings, threats, and the like? Speech acts in the nonverbal mode
Multimodal constructions revisited. Testing the strength of association between spoken and non-spoken features of Tell me about it
Eye gaze and viewpoint in multimodal interaction management
Sámi tourism in marketing material
Speech Act Theory
Multimodal Coordination of Sound and Movement in Music and Speech
A Corpus Study on the Difference of Turn-Taking in Online Audio, Online Video, and Face-to-Face Conversation
| Citation velocity | historical |
|---|---|
| Highly cited | No |