How Is a “Kitchen Chair” like a “Farm Horse”? Exploring the Representation of Noun-Noun Compound Semantics in Transformer-based Language Models
Bibliographic Data
| ID | 12156136 |
|---|---|
| Authors | Mark Ormerod (0000-0003-3454-2863, Queen’s University Belfast. [email protected], corresponding author), Jesús Martínez del Rincón (0000-0002-9574-4138, Queen's University Belfast), Barry Devereux (0000-0003-2128-8632, Queen's University Belfast) |
| Year | 2023 |
| Volume | 50 |
| Issue | 1 |
| Pages | 49-81 |
| Publication date | 2023-11-15 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Computational Linguistics (JOURNAL) |
| Journal identifiers | ISSN: 0891-2017 • E-ISSN: 1530-9312 |
| Publisher | Association for Computational Linguistics (PUBLISHER • US) |
| DOI | 10.1162/coli_a_00495 |
| OpenAlex | W4388691779 |
| Language | EN |
| Citations received | 1 |
| References cited | 41 |
Despite the success of Transformer-based language models in a wide variety of natural language processing tasks, our understanding of how these models process a given input in order to represent task-relevant information remains incomplete. In this work, we focus on semantic composition and examine how Transformer-based language models represent semantic information related to the meaning of English noun-noun compounds. We probe Transformer-based language models for their knowledge of the thematic relations that link the head nouns and modifier words of compounds (e.g., KITCHEN CHAIR: a chair located in a kitchen). Firstly, using a dataset featuring groups of compounds with shared lexical or semantic features, we find that token representations of six Transformer-based language models distinguish between pairs of compounds based on whether they use the same thematic relation. Secondly, we utilize fine-grained vector representations of compound semantics derived from human annotations, and find that token vectors from several models elicit a strong signal of the semantic relations used in the compounds. In a novel “compositional probe” setting, where we compare the semantic relation signal in mean-pooled token vectors of compounds to mean-pooled token vectors when the two constituent words appear in separate sentences, we find that the Transformer-based language models that best represent the semantics of noun-noun compounds also do so substantially better than in the control condition where the two constituent works are processed separately. Overall, our results shed light on the ability of Transformer-based language models to support compositional semantic processes in representing the meaning of noun-noun compounds
Natural language processing · Noun · Noun phrase · Question answering · Security token · Transformer · Computer Science · Natural Language Processing Techniques · Text Readability and Simplification · Topic Modeling · Artificial Intelligence
The Big Book of Concepts
Transformers
Composition in Distributional Models of Semantics
Degrees of Freedom in Planning, Running, Analyzing, and Reporting Psychological Studies
A Primer in Bertology
What can linguistics and deep learning contribute to each other? Response to Pater
Frequency effects in the processing of lexicalized and novel nominal compounds
The Syntax and Semantics of Complex Nominals
On the Creation and Use of English Compound Nouns
| Unique citing works | 1 |
|---|---|
| Citations per year | 1 |
| Citation span | 2025 - 2025 (1) |
| Citation velocity | recent |
| Highly cited | No |
| Citation types | Neutral: 1 |