Deductively coding psychosocial autopsy interview data using a few-shot learning large language model
Bibliographic Data
| ID | 22090577 |
|---|---|
| Authors | Elias Balt (0000-0002-7022-2106, GGD Amsterdam, corresponding author), Salim Salmi (0000-0002-8342-4815, GGD Amsterdam), Sandjai Bhulai (0000-0003-1124-8821), Stefan Vrinzen (0009-0007-3551-0325, GGD Amsterdam), Merijn Eikelenboom (0000-0003-4934-6427, GGD Amsterdam), Renske Gilissen (0000-0002-8009-6326, Leiden University), Daan Creemers, Daan H M Creemers (0000-0001-8638-8971, Radboud University Nijmegen), Arne Popma (0000-0003-2170-3023, Amsterdam University Medical Centers), Saskia Mérelle (0000-0003-1748-7700, GGD Amsterdam) |
| Year | 2025 |
| Volume | 13 |
| Pages | 1512537-1512537 |
| Publication date | 2025-02-19 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Frontiers in Public Health (JOURNAL) |
| Journal identifiers | ISSN: 2296-2565 • E-ISSN: 2296-2565 |
| Publisher | Frontiers Media SA (PUBLISHER • CH) |
| DOI | 10.3389/fpubh.2025.1512537 |
| PMID | 40046117 |
| OpenAlex | W4407757890 |
| Language | EN |
| Citations received | 1 |
| References cited | 24 |
Background: Psychosocial autopsy is a retrospective study of suicide, aimed to identify emerging themes and psychosocial risk factors. It typically relies heavily on qualitative data from interviews or medical documentation. However, qualitative research has often been scrutinized for being prone to bias and is notoriously time- and cost-intensive. Therefore, the current study aimed to investigate if a Large Language Model (LLM) can be feasibly integrated with qualitative research procedures, by evaluating the performance of the model in deductively coding and coherently summarizing interview data obtained in a psychosocial autopsy. Methods: Data from 38 semi-structured interviews conducted with individuals bereaved by the suicide of a loved one was deductively coded by qualitative researchers and a server-installed LLAMA3 large language model. The model performance was evaluated in three tasks: (1) binary classification of coded segments, (2) independent classification using a sliding window approach, and (3) summarization of coded data. Intercoder agreement scores were calculated using Cohen's Kappa, and the LLM's summaries were qualitatively assessed using the Constant Comparative Method. Results: The results showed that the LLM achieved substantial agreement with the researchers for the binary classification (accuracy: 0.84) and the sliding window task (accuracy: 0.67). The performance had large variability across codes. LLM summaries were typically rich enough for subsequent analysis by the researcher, with around 80% of the summaries being rated independently by two researchers as 'adequate' or 'good.' Emerging themes in the qualitative assessment of the summaries included unsolicited elaboration and hallucination. Conclusion: State-of-the-art LLMs show great potential to support researchers in deductively coding complex interview data, which would alleviate the investment of time and resources. Integrating models with qualitative research procedures can facilitate near real-time monitoring. Based on the findings, we recommend a collaborative model, whereby the LLM's deductive coding is complemented by review, inductive coding and further interpretation by a researcher. Future research may aim to replicate the findings in different contexts and evaluate models with a larger context size
Cause of death · Natural language processing · Pathology · Psychiatry · Psychosocial · Sociology · Verbal autopsy · Computer Science · Grief, Bereavement, and Mental Health · Medicine · Mental Health via Writing · Psychology · Suicide and Self-Harm Studies · Artificial Intelligence
A Purposeful Approach to the Constant Comparative Method in the Analysis of Qualitative Interviews
The Measurement of Observer Agreement for Categorical Data
Using artificial intelligence to improve public health
Sociodemographic and psychosocial risk factors of railway suicide
Compatibility between Text Mining and Qualitative Research in the Perspectives of Grounded Theory, Content Analysis, and Reliability
Artificial Intelligence Augmented Qualitative Analysis
What is Qualitative in Qualitative Research
An Examination of the Use of Large Language Models to Aid Analysis of Textual Data
| Unique citing works | 1 |
|---|---|
| Citations per year | 1 |
| Citation span | 2026 - 2026 (1) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 1 |