Mechanistic indicators of understanding in large language models
Bibliographic Data
| ID | 21370950 |
|---|---|
| Authors | Pierre Beckmann (0000-0001-9247-4841, Ecole polytechnique fédérale de Lausanne, corresponding author), Matthieu Queloz (0000-0001-6644-9992, University of Bern) |
| Year | 2026 |
| Volume | 183 |
| Issue | 6 |
| Pages | 1747-1792 |
| Publication date | 2026-06-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Philosophical Studies (JOURNAL) |
| Journal identifiers | ISSN: 0031-8116 • E-ISSN: 1573-0883 |
| Publisher | Springer Science and Business Media LLC (PUBLISHER) |
| DOI | 10.1007/s11098-026-02513-1 |
| OpenAlex | W4417275656 |
| Language | EN |
| Citations received | 1 |
| References cited | 57 |
Large language models are often portrayed as merely imitating linguistic patterns without genuine understanding. We argue that recent findings in mechanistic interpretability, the emerging field probing the inner workings of LLMs, render this picture increasingly untenable—but only once those findings are integrated within a theoretical account of understanding. We propose a tiered framework for thinking about understanding in LLMs and use it to synthesize the most relevant findings to date. The framework distinguishes three hierarchical varieties of understanding, each tied to a corresponding level of computational organization: conceptual understanding emerges when a model forms “features” as directions in latent space, learning connections between diverse manifestations of a single entity or property; state-of-the-world understanding emerges when a model learns contingent factual connections between features and dynamically tracks changes in the world; principled understanding emerges when a model ceases to rely on memorized facts and discovers a compact “circuit” connecting these facts. Across these tiers, MI uncovers internal organizations that can underwrite understanding-like unification. However, these also diverge from human cognition in their parallel exploitation of heterogeneous mechanisms. Fusing philosophical theory with mechanistic evidence thus allows us to transcend binary debates over whether AI understands, paving the way for a comparative, mechanistically grounded epistemology that explores how AI understanding aligns with—and diverges from—our own
Cognition · Comprehension · Computational model · Field (mathematics) · Interpretability · Language Understanding · Computational and Text Analysis Methods · Explainable Artificial Intelligence (XAI · Language and cultural evolution
The Value of Knowledge and the Pursuit of Understanding
Understanding, Explanation, and Scientific Knowledge
Understanding Scientific Understanding
Representation Learning
Explanation and Scientific Understanding
Climbing towards NLU
The debate over understanding in AI’s large language models
Logic, Meaning, and Conceptual Role
Beyond Concepts
A Mark of the Mental
A Study of Concepts
Can structural correspondences ground real‐world representational content in large language models
Moving beyond content‐specific computation in artificial neural networks
Simulacra as conscious exotica
Understanding phenomena
No understanding without explanation
The Epistemic Value of Understanding
In Defense of Proper Functions
Explanatory Unification
Functions
Psychologism and Behaviorism
Is Understanding A Species Of Knowledge
Understanding from Machine Learning Models
Grasping in Understanding
Anti-Normativism Evaluated
Still no lie detector for language models
| Unique citing works | 1 |
|---|---|
| Citations per year | 1 |
| Citation span | 2026 - 2026 (1) |
| Citation velocity | current |
| Highly cited | No |