Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Multimodal Spatiotemporal Semisupervised Transformer Network for Video-Based Group-Level Emotion Recognition

Bibliographic Data

ID22106965
AuthorsXiaohua Huang (0000-0001-8897-3517, Nanjing Institute of Technology), Jinke Xu (0009-0001-6032-2795, Nanjing Institute of Technology)
Year2026
Volume13
Issue2
Pages1719-1733
Publication date2026-04-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueIEEE Transactions on Computational Social Systems (JOURNAL)
Journal identifiersISSN: 2329-924X • E-ISSN: 2373-7476
PublisherInstitute of Electrical and Electronics Engineers (IEEE) (PUBLISHER)
DOI10.1109/tcss.2025.3622991
OpenAlexW4416513078
LanguageEN
Citations received1
References cited58

Group-level emotion recognition (GER) has emerged as a critical research topic for identifying collective emotions in multiperson scenarios. Despite recent advancements, such as dual branch cross-attention (CA) mechanism, existing methods struggle to differentiate ambiguous emotion categories effectively. In addition, the limited size and diversity of GER datasets hinder further performance improvements. To address these challenges, this article introduces a novel approach, the multimodal spatiotemporal semisupervised transformer (MSST). First, we propose a multimodal spatiotemporal transformer to encode spatial features, capture temporal dynamics, and fuse information from three modalities effectively. Second, a semisupervised learning (SSL) strategy leverages unlabeled data, enhancing robustness against noise and outliers. Last, a two-stage classification strategy and consistency loss are introduced to improve the model’s ability to handle category ambiguity and ensure robust predictions for similar samples. Comprehensive experiments conducted on benchmark GER datasets demonstrate that MSST either considerably outperforms or achieves competitive performance compared to state-of-the-art methods, underscoring its effectiveness in advancing GER research and overcoming the limitations of existing approaches

Ambiguity · ENCODE · Modalities · Noise immunity · Transformer · Emotion and Mood Recognition · Face and Expression Recognition · Sentiment Analysis and Opinion Mining

  • A Survey on Deep Learning for Group-Level Emotion Recognition

    Open Access•Xiaohua Huang, Xiaopeng Hong et al.•IEEE Transactions on Computational…•2026

  • Long Short-Term Memory

    Sepp Hochreiter, Jurgen Schmidhuber•Neural Computation•1997

  • A Self-Fusion Network Based on Contrastive Learning for Group Emotion Recognition

    Open Access•Xingzhi Wang, Dong Zhang et al.•IEEE Transactions on Computational…•2023

  • Perceived collective continuity

    Open Access•Fabio Sani, Mhairi Bowe et al.•European Journal of Social…•2007

  • Collective Emotions

    Open Access•Amit Goldenberg, David García et al.•Current Directions in Psychological…•2020

Unique citing works1
Citations per year1
Citation span2026 - 2026 (1)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 1

Tools

Open DOI
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae