Multimodal Depression Detection Based on Self-Attention Network With Facial Expression and Pupil
Datos Bibliográficos
| ID | 22107352 |
|---|---|
| Autores | Xiang Liu (0000-0002-9541-0541, Dongguan University of Technology), Hao Shen (0000-0003-3361-6058, Lanzhou University), Huiru Li (0000-0001-7334-569X, Lanzhou University), Yongfeng Tao (0009-0009-9712-9380, Lanzhou University), Minqiang Yang (0000-0002-7571-6439, Lanzhou University) |
| Año | 2025 |
| Volumen | 12 |
| Número | 1 |
| Páginas | 64-76 |
| Fecha de publicación | 2025-02-01 |
| Peer Reviewed | Sí |
| Open Access | Sí |
| Tipo | ARTICLE |
| Revista | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Identificadores de la revista | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Editorial | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2024.3405949 |
| OpenAlex | W4400033126 |
| Idioma | EN |
| Citas recibidas | 2 |
| Referencias citadas | 61 |
Depression is a major mental health issue in contemporary society, with an estimated 350 million people affected globally. The number of individuals diagnosed with depression continues to rise each year. Currently, clinical practice relies entirely on self-reporting and clinical assessment, which carries the risk of subjective biases. In this article, we propose a multimodal method based on facial expression and pupil to detect depression more objectively and precisely. Our method first extracts the features of facial expressions and pupil diameter using residual networks and 1-D convolutional neural networks. Second, a cross-modal fusion model based on self-attention networks (CMF-SNs) is proposed, which utilizes cross-modal attention networks within modalities and parallel self-attention networks between different modalities to extract CMF features of facial expressions and pupil diameter, effectively complementing information between different modalities. Finally, the obtained features are fully connected to identify depression. Multiple controlled experiments show that compared to single modality, the multimodal fusion method based on self-attention networks has a higher ability to recognize depression, with the highest accuracy of 75.0%. In addition, we conducted comparative experiments under three different stimulation paradigms, and the results showed that the classification accuracy under negative and neutral stimuli was higher than that under positive stimuli, indicating a bias of depressed patients toward negative images. The experimental results demonstrate the superiority of our multimodal fusion method
Cognitive psychology · Facial expression · Pupil · Computer Science · Emotion and Mood Recognition · Neuroscience · Psychology · Artificial Intelligence
The pupil as a measure of emotional arousal and autonomic activation
Attentional Biases for Negative Interpersonal Stimuli in Clinical Depression.
Global prevalence and burden of depressive and anxiety disorders in 204 countries and territories in 2020 due to the Covid-19 pandemic
Orthogonal-Moment-Based Attraction Measurement With Ocular Hints in Video-Watching Task
What Does Your Bio Say? Inferring Twitter Users’ Depression Status From Multimodal Profile Information Using Deep Learning
Factors influencing suicidal tendencies during Covid-19 pandemic in Korean multicultural adolescents
| Obras citantes distintas | 2 |
|---|---|
| Citas por año | 2 |
| Intervalo de citas | 2025 - 2026 (2) |
| Velocidad de citación | current |
| Altamente citado | No |
| Tipos de cita | Neutras: 2 |