Multimodal Spatiotemporal Semisupervised Transformer Network for Video-Based Group-Level Emotion Recognition
Bibliographic Data
| ID | 22106965 |
|---|---|
| Authors | Xiaohua Huang (0000-0001-8897-3517, Nanjing Institute of Technology), Jinke Xu (0009-0001-6032-2795, Nanjing Institute of Technology) |
| Year | 2026 |
| Volume | 13 |
| Issue | 2 |
| Pages | 1719-1733 |
| Publication date | 2026-04-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Journal identifiers | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Publisher | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2025.3622991 |
| OpenAlex | W4416513078 |
| Language | EN |
| Citations received | 1 |
| References cited | 58 |
Group-level emotion recognition (GER) has emerged as a critical research topic for identifying collective emotions in multiperson scenarios. Despite recent advancements, such as dual branch cross-attention (CA) mechanism, existing methods struggle to differentiate ambiguous emotion categories effectively. In addition, the limited size and diversity of GER datasets hinder further performance improvements. To address these challenges, this article introduces a novel approach, the multimodal spatiotemporal semisupervised transformer (MSST). First, we propose a multimodal spatiotemporal transformer to encode spatial features, capture temporal dynamics, and fuse information from three modalities effectively. Second, a semisupervised learning (SSL) strategy leverages unlabeled data, enhancing robustness against noise and outliers. Last, a two-stage classification strategy and consistency loss are introduced to improve the model’s ability to handle category ambiguity and ensure robust predictions for similar samples. Comprehensive experiments conducted on benchmark GER datasets demonstrate that MSST either considerably outperforms or achieves competitive performance compared to state-of-the-art methods, underscoring its effectiveness in advancing GER research and overcoming the limitations of existing approaches
Ambiguity · ENCODE · Modalities · Noise immunity · Transformer · Emotion and Mood Recognition · Face and Expression Recognition · Sentiment Analysis and Opinion Mining
| Unique citing works | 1 |
|---|---|
| Citations per year | 1 |
| Citation span | 2026 - 2026 (1) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 1 |