Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Mechanistic indicators of understanding in large language models

Bibliographic Data

ID21370950
AuthorsPierre Beckmann (0000-0001-9247-4841, Ecole polytechnique fédérale de Lausanne, corresponding author), Matthieu Queloz (0000-0001-6644-9992, University of Bern)
Year2026
Volume183
Issue6
Pages1747-1792
Publication date2026-06-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenuePhilosophical Studies (JOURNAL)
Journal identifiersISSN: 0031-8116 • E-ISSN: 1573-0883
PublisherSpringer Science and Business Media LLC (PUBLISHER)
DOI10.1007/s11098-026-02513-1
OpenAlexW4417275656
LanguageEN
Citations received1
References cited57

Large language models are often portrayed as merely imitating linguistic patterns without genuine understanding. We argue that recent findings in mechanistic interpretability, the emerging field probing the inner workings of LLMs, render this picture increasingly untenable—but only once those findings are integrated within a theoretical account of understanding. We propose a tiered framework for thinking about understanding in LLMs and use it to synthesize the most relevant findings to date. The framework distinguishes three hierarchical varieties of understanding, each tied to a corresponding level of computational organization: conceptual understanding emerges when a model forms “features” as directions in latent space, learning connections between diverse manifestations of a single entity or property; state-of-the-world understanding emerges when a model learns contingent factual connections between features and dynamically tracks changes in the world; principled understanding emerges when a model ceases to rely on memorized facts and discovers a compact “circuit” connecting these facts. Across these tiers, MI uncovers internal organizations that can underwrite understanding-like unification. However, these also diverge from human cognition in their parallel exploitation of heterogeneous mechanisms. Fusing philosophical theory with mechanistic evidence thus allows us to transcend binary debates over whether AI understands, paving the way for a comparative, mechanistically grounded epistemology that explores how AI understanding aligns with—and diverges from—our own

Cognition · Comprehension · Computational model · Field (mathematics) · Interpretability · Language Understanding · Computational and Text Analysis Methods · Explainable Artificial Intelligence (XAI · Language and cultural evolution

  • On Large Language Models, De-anthropomorphized Narrative Production, and Deflationary Understanding

    Open Access•Warmhold Jan Thomas Mollema•Philosophy & Technology•2026

  • The Value of Knowledge and the Pursuit of Understanding

    Open Access•Jonathan L Kvanvig•Value of Knowledge and the…•2003

  • Understanding, Explanation, and Scientific Knowledge

    Kareem Khalifa•Understanding, Explanation, and…•2017

  • Understanding Scientific Understanding

    Henk W De Regt•Understanding Scientific…•2017

  • Representation Learning

    Open Access•Yoshua Bengio, Aaron Courville et al.•IEEE Transactions on Pattern…•2013

  • Explanation and Scientific Understanding

    Michael Friedman•The Journal of Philosophy•1974

  • Climbing towards NLU

    Open Access•Emily M Bender, Alexander Koller•Proceedings of the 58th Annual…•2020

  • The debate over understanding in AI’s large language models

    Open Access•Melanie Mitchell, David C Krakauer•Proceedings of the National…•2023

  • Logic, Meaning, and Conceptual Role

    HHF HHF•The Journal of Philosophy•1977

  • Beyond Concepts

    Ruth Garrett Millikan•Beyond Concepts•2017

  • A Mark of the Mental

    Karen Neander•A Mark of the Mental•2017

  • A Study of Concepts

    Christopher Peacocke•A Study of Concepts•1992

  • Can structural correspondences ground real‐world representational content in large language models

    Open Access•Iwan Williams•Mind & Language•2026

  • Moving beyond content‐specific computation in artificial neural networks

    Open Access•Nicholas Shea•Mind & Language•2023

  • Simulacra as conscious exotica

    Open Access•Murray Shanahan•Inquiry•2026

  • Understanding phenomena

    Open Access•Christoph Kelp•Synthese•2015

  • No understanding without explanation

    Open Access•Michael Strevens•Studies in History and Philosophy…•2013

  • The Epistemic Value of Understanding

    Open Access•Henk W De Regt•Philosophy of Science•2009

  • In Defense of Proper Functions

    Open Access•Ruth Garrett Millikan•Philosophy of Science•1989

  • Explanatory Unification

    Open Access•Philip Kitcher•Philosophy of Science•1981

  • Functions

    Larry Wright•The Philosophical Review•1973

  • Psychologism and Behaviorism

    Ned Block•The Philosophical Review•1981

  • Is Understanding A Species Of Knowledge

    Stephen R Grimm•The British Journal for the…•2006

  • Understanding from Machine Learning Models

    Open Access•Emily Sullivan•The British Journal for the…•2022

  • Grasping in Understanding

    Miloud Belkoniene•The British Journal for the…•2023

  • Anti-Normativism Evaluated

    Ulf Hlobil•International Journal of…•2015

  • Still no lie detector for language models

    Open Access•Benjamin A Levinstein, Daniel A Herrmann•Philosophical Studies•2025

Unique citing works1
Citations per year1
Citation span2026 - 2026 (1)
Citation velocitycurrent
Highly citedNo

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae