A Primer in Bertology
What We Know About How Bert Works
Bibliographic Data
| ID | 23330874 |
|---|---|
| Authors | Anna Rogers (0000-0002-4845-4023, University of Copenhagen), Olga Kovaleva (0000-0001-7880-2781, University of Massachusetts Lowell), Anna Rumshisky (University of Massachusetts Lowell) |
| Year | 2020 |
| Volume | 8 |
| Pages | 842-866 |
| Publication date | 2020-12-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Transactions of the Association for Computational Linguistics (JOURNAL) |
| Journal identifiers | ISSN: 2307-387X • E-ISSN: 2307-387X |
| Publisher | MIT Press (PUBLISHER • US) |
| DOI | 10.1162/tacl_a_00349 |
| OpenAlex | W3006881356 |
| Language | EN |
| Citations received | 65 |
| References cited | 35 |
Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue, and approaches to compression. We then outline directions for future research.
Art · Data science · Electrical engineering · Programming language · State (computer science) · Transformer · Visual arts · Architecture · Artificial Intelligence · Computer Science · Engineering · Multimodal Machine Learning Applications · Natural Language Processing Techniques · Topic Modeling
A Brief Overview of ChatGPT
From metacognitive monitoring to control
Public opinion on the clean energy transition during times of crisis
What neural networks know about linguistic complexity
Identifying the underlying psychological constructs from self-expressed anti-vaccination argumentation
RICo
Transformers, Contextualism, and Polysemy
Fictionalism about Chatbots
CGraphNet
Sentiment and Objectivity in Iranian State-Sponsored Propaganda on Twitter
Computational Measures of Deceptive Language
Closing the Gap
A Novel Deep Learning Approach Using Contextual Embeddings for Toponym Resolution
Response to critique of the paper
Classifying Genetic Essentialist Biases using Large Language Models
Piercing a Methodological Bubble
Polysemy—Evidence from Linguistics, Behavioral Science, and Contextualized Language Models
Recollective features in the natural language used to justify memory decisions
Structural priming in humans and large language models
Incremental alternative sampling as a lens into the temporal and representational resolution of linguistic prediction
Large-scale analysis of online social data on the long-term sentiment and content dynamics of online (mis)information
Deptweet
Towards reliable generative AI-driven scaffolding
Can structural correspondences ground real‐world representational content in large language models
Moving beyond content‐specific computation in artificial neural networks
Coding energy knowledge in constructed responses with explainable NLP models
SMSF
Sharing Our Concepts with Machines
Modelling language using large language models
A chimpanzee by any other name
La Revolución en la Creación Visual
Beyond Computational Formalism or, Architecture Matters
Unmasking Machine Learning With Tensor Decomposition
Evaluating Natural Language Processing and Named Entity Recognition for Bioarchaeological Data Reuse
Syntactic Structure from Deep Learning
Keeping Humans in the Loop
Differentiating translational English from original English using large language models
Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question Answering
How Is a “Kitchen Chair” like a “Farm Horse”? Exploring the Representation of Noun-Noun Compound Semantics in Transformer-based Language Models
The Taxonomy of Writing Systems
Investigating Idiomaticity in Word Representations
Probing Classifiers
The Analysis of Synonymy and Antonymy in Discourse Relations
Language Model Behavior
Humans Learn Language from Situated Communicative Interactions. What about Machines
Categorizing political campaign messages on social media using supervised machine learning
Something to Do with Paying Attention
On the present-future impact of AI technologies on personnel selection and the exponential increase in meta-algorithmic judgments
Towards a Definition of Generative Artificial Intelligence
Multi-Hazard assessment methods for planning
Bert-deep CNN
Data augmentation using instruction-tuned models improves emotion analysis in tweets
Collaborative Growth
Automated Detection of Media Bias Using Artificial Intelligence and Natural Language Processing
Uncovering patterns of semantic predictability in sentence processing
Dissociable frequency effects attenuate as large language model surprisal predictors improve
The More Similar, the Better? Associations Between Latent Semantic Similarity and Emotional Experiences Differ Across Conversation Contexts
Too dense to comprehend—or perfectly packed? The divergent effects of propositional density on levels of text memory
Large language models and linguistic intentionality
The ambiguity of Bertology
Unveiling ideological extremes in parliamentary debates using transformer-based language models
Co‐textual dopes
ViTHSD
Classification of Poverty Condition Using Natural Language Processing
Conversing' With Qualitative Data
| Unique citing works | 65 |
|---|---|
| Citations per year | 10,83 |
| Citation span | 2020 - 2026 (7) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 65 |