Pular para o conteúdo principal

ETHNOS_APP

Início • Busca • Periódicos • Lista 0

Relative Value Encoding in Large Language Models

A Multi-Task, Multi-Model Investigation

Dados Bibliográficos

ID22157751
AutoresWilliam M Hayes (0000-0001-5378-656X, Binghamton University), Nicolas Yax (0009-0008-1176-5806, Institut national de recherche en sciences et technologies du numérique), Stefano Palminteri (0000-0001-5768-6646, Université Paris Sciences et Lettres)
Ano2025
Volume9
Páginas709-725
Data de publicação2025-05-09
Peer ReviewedSim
Open AccessSim
TipoARTICLE
PeriódicoOpen MIND (JOURNAL)
Identificadores do periódicoISSN: 2470-2986 • E-ISSN: 2470-2986
EditoraMIT Press (PUBLISHER • US)
DOI10.1162/opmi_a_00209
PMID40474931
OpenAlexW4410612370
IdiomaEN
Referências citadas46

Abtract In-context learning enables large language models (LLMs) to perform a variety of tasks, including solving reinforcement learning (RL) problems. Given their potential use as (autonomous) decision-making agents, it is important to understand how these models behave in RL tasks and the extent to which they are susceptible to biases. Motivated by the fact that, in humans, it has been widely documented that the value of a choice outcome depends on how it compares to other local outcomes, the present study focuses on whether similar value encoding biases apply to LLMs. Results from experiments with multiple bandit tasks and models show that LLMs exhibit behavioral signatures of relative value encoding. Adding explicit outcome comparisons to the prompt magnifies the bias, impairing the ability of LLMs to generalize from the outcomes presented in-context to new choice problems, similar to effects observed in humans. Computational cognitive modeling reveals that LLM behavior is well-described by a simple RL algorithm that incorporates relative values at the outcome encoding stage. Lastly, we present preliminary evidence that the observed biases are not limited to fine-tuned LLMs, and that relative value processing is detectable in the final hidden layer activations of a raw, pretrained model. These findings have important implications for the use of LLMs in decision-making applications

Machine learning · Natural language processing · Computer Science · Engineering · Explainable Artificial Intelligence (XAI · Natural Language Processing Techniques · Topic Modeling · Artificial Intelligence

  • Reinforcement Learning

    Open Access•Richard S Sutton, A G Barto et al.•IEEE Transactions on Neural…•1998

  • Generative Agents

    Open Access•Joon Sung Park, Joseph O’Brien et al.•Proceedings of the 36th Annual…•2023

  • Using cognitive psychology to understand GPT-3

    Open Access•Marcel Binz, Emily Schulz et al.•Proceedings of the National…•2023

  • Large language models in medicine

    Open Access•Arun James Thirunavukarasu, Darren Shu Jeng Ting et al.•Nature Medicine•2023

  • Testing models of context-dependent outcome encoding in reinforcement learning

    Open Access•William M Hayes, Douglas H Wedell•Cognition•2023

  • Behavioural and neural characterization of optimistic reinforcement learning

    Open Access•Germain Lefebvre, Maël Lebreton et al.•Nature Human Behaviour•2017

Velocidade de citaçãohistorical
Altamente citadoNão
Ethnos_APP • Projeto Open Source • Licença MIT • Frontend v2.0.0 • Privacidade e Cookies • Documentação da API: api.ethnos.app/docs • Código da API: GitHub • DOI: 10.5281/zenodo.17049435 • Código do Frontend: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae