James L Mcclelland
Biographic Data
| ID | 608633 |
|---|---|
| NAME | James L Mcclelland |
| GIVEN NAMES | James L |
| FAMILY NAME | Mcclelland |
| SIGNATURE | MCCLELLAND J L |
| AFFILIATIONS | Stanford University |
| ORCID | 0000-0002-8217-405X |
| VERIFIED | Yes |
| TOTAL WORKS | 21 |
| TOTAL CITATIONS | 19 |
| AUTHOR COUNT | 21 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1981 |
| LATEST PUBLICATION YEAR | 2025 |
| H-INDEX | 3 |
Learning to Decompose: Human-Like Subgoal Preferences Emerge in Neural Networks Learning Graph Traversal
Cognitive scientists have discovered normative and heuristic principles that capture human subgoal preferences when partitioning problems into smaller ones. However, it remains unclear where such preferences come from and why they tend to be both effective and efficient. In this work, we study the processes through which these preferences may be implicitly encoded over learning as learners improve towards optimal traversals. We build on the graph…
Systematic Human Learning and Generalization From a Brief Tutorial With Explanatory Feedback
We investigate human adults’ ability to learn an abstract reasoning task quickly and to generalize outside of the range of training examples. Using a task based on a solution strategy in Sudoku, we provide Sudoku-naive participants with a brief instructional tutorial with explanatory feedback using a narrow range of training examples. We find that most participants who master the task do so within 10 practice trials and generalize well to puzzles…
Do estimates of numerosity really adhere to Weber’s law? A reexamination of two case studies
Both humans and nonhuman animals can exhibit sensitivity to the approximate number of items in a visual array or events in a sequence, and across various paradigms, uncertainty in numerosity judgments increases with the number estimated or produced. The pattern of increase is usually described as exhibiting approximate adherence to Weber’s law, such that uncertainty increases proportionally to the mean estimate, resulting in a constant coefficien…
Exemplar models are useful and deep neural networks overcome their limitations: A commentary on Ambridge (2020)
Humans are sensitive to the properties of individual items, and exemplar models are useful for capturing this sensitivity. I am a proponent of an extension of exemplar-based architectures that I briefly describe. However, exemplar models are very shallow architectures in which it is necessary to stipulate a set of primitive elements that make up each example, and such architectures have not been as successful as deep neural networks in capturing …
Modelling the N400 brain potential as change in a probabilistic representation of meaning
Cognitive Neuroscience
Connectionism and the Emergence of Mind
Predicting native English-like performance by native Japanese speakers
Language is not Just for Talking: Redundant Labels Facilitate Learning of Novel Categories
In addition to having communicative functions, verbal labels may play a role in shaping concepts. Two experiments assessed whether the presence of labels affected category formation. Subjects learned to categorize “aliens” as those to be approached or those to be avoided. After accuracy feedback on each response was provided, a nonsense label was either presented or not. Providing nonsense category labels facilitated category learning even though…
Gradience of Gradience: A reply to Jackendoff
Jackendoff and other linguists have acknowledged that there is gradience in language but have tended to treat gradient phenomena as separate from the core of language, which is viewed as fully productive and compositional. This perspective suffuses Jackendoff's (2007) response to our position paper (Bybee and McClelland 2005). We argue that gradience is an inherent feature of language representation, processing, and learning, and that natural lan…
Alternatives to the combinatorial paradigm of linguistic theory based on domain general principles of human cognition
It is argued that the principles needed to explain linguistic behavior are domain-general and based on the impact that specific experiences have on the mental organization and representation of language. This organization must be sensitive to both specific information and generalized patterns. In addition, knowledge of language is highly sensitive to frequency of use: frequently-used linguistic sequences become more frequent, more accessible and …
The time course of perceptual choice: The leaky, competing accumulator model.
The time course of perceptual choice is discussed in a model of gradual, leaky, stochastic, and competitive information accumulation in nonlinear decision units. Special cases of the model match a classical diffusion process, but leakage and competition work together to address several challenges to existing diffusion, random walk, and accumulator models. The model accounts for data from choice tasks using both time-controlled (e.g., response sig…
Understanding normal and impaired word reading: Computational principles in quasi-regular domains.
A connectionist approach to processing in quasi-regular domains, as exemplified by English word reading, is developed. Networks using appropriately structured orthographic and phonological representations were trained to read both regular and exception words, and yet were also able to read pronounceable nonwords as well as skilled readers. A mathematical analysis of a simplified system clarifies the close relationship of word frequency and spelli…
Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory.
Damage to the hippocampal system disrupts recent memory but leaves remote memory intact. The account presented here suggests that memories are first stored via synaptic changes in the hippocampal system, that these changes support reinstatement of recent memories in the neocortex, that neocortical synapses change a little on each reinstatement, and that remote memory is based on accumulated neocortical changes. Models that learn via changes to co…
On the control of automatic processes: A parallel distributed processing account of the Stroop effect.
Traditional views of automaticity are in need of revision. For example, automaticity often has been treated as an all-or-none phenomenon, and traditional theories have held that automatic processes are independent of attention. Yet recent empirical data suggest that automatic processes are continuous, and furthermore are subject to attentional control. A model of attention is presented to address these issues. Within a parallel distributed proces…
A distributed, developmental model of word recognition and naming.
A parallel distributed processing model of visual word recognition and pronunciation is described. The model consists of sets of orthographic and phonological units and an interlevel of hidden units. Weights on connections between units were modified during a training phase using the back-propagation learning algorithm. The model simulates many aspects of human performance, including (a) differences between words in terms of processing difficulty…
Une nouvelle approche de la cognition: Le Connexionnisme
Parallel Distributed Processing: Explorations in the Microstructures of Cognition
1. Very rarely, a book is published which not only advances our knowledge of a particular topic, but fundamentally recasts our methods of investigating and thinking about large tracts of the map of learning. Linguists remember 1957 as the publication year of Noam Chomsky's Syntactic structures-a book whose ostensible subjects were the structure of English grammatical rules and the goals of grammatical description, but which can be seen with hinds…
Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Foundations
What makes people smarter than computers? These volumes by a pioneering neurocomputing group suggest that the answer lies in the massively parallel architecture of the human mind. They describe a new theory of cognition called connectionism that is challenging the idea of symbolic computation that has traditionally been at the center of debate in theoretical discussions about the mind. The authors' theory assumes the mind is composed of a great n…
The TRACE model of speech perception
An interactive activation model of context effects in letter perception: I. An account of basic findings.
Predicting native English-like performance by native Japanese speakers
Modelling the N400 brain potential as change in a probabilistic representation of meaning
Parallel Distributed Processing: Explorations in the Microstructures of Cognition
1. Very rarely, a book is published which not only advances our knowledge of a particular topic, but fundamentally recasts our methods of investigating and thinking about large tracts of the map of learning. Linguists remember 1957 as the publication year of Noam Chomsky's Syntactic structures-a book whose ostensible subjects were the structure of English grammatical rules and the goals of grammatical description, but which can be seen with hinds…
Une nouvelle approche de la cognition: Le Connexionnisme
An interactive activation model of context effects in letter perception: I. An account of basic findings.
Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Foundations
What makes people smarter than computers? These volumes by a pioneering neurocomputing group suggest that the answer lies in the massively parallel architecture of the human mind. They describe a new theory of cognition called connectionism that is challenging the idea of symbolic computation that has traditionally been at the center of debate in theoretical discussions about the mind. The authors' theory assumes the mind is composed of a great n…
The TRACE model of speech perception
Une nouvelle approche de la cognition: Le Connexionnisme
Parallel Distributed Processing: Explorations in the Microstructures of Cognition
1. Very rarely, a book is published which not only advances our knowledge of a particular topic, but fundamentally recasts our methods of investigating and thinking about large tracts of the map of learning. Linguists remember 1957 as the publication year of Noam Chomsky's Syntactic structures-a book whose ostensible subjects were the structure of English grammatical rules and the goals of grammatical description, but which can be seen with hinds…
A distributed, developmental model of word recognition and naming.
A parallel distributed processing model of visual word recognition and pronunciation is described. The model consists of sets of orthographic and phonological units and an interlevel of hidden units. Weights on connections between units were modified during a training phase using the back-propagation learning algorithm. The model simulates many aspects of human performance, including (a) differences between words in terms of processing difficulty…
On the control of automatic processes: A parallel distributed processing account of the Stroop effect.
Traditional views of automaticity are in need of revision. For example, automaticity often has been treated as an all-or-none phenomenon, and traditional theories have held that automatic processes are independent of attention. Yet recent empirical data suggest that automatic processes are continuous, and furthermore are subject to attentional control. A model of attention is presented to address these issues. Within a parallel distributed proces…
Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory.
Damage to the hippocampal system disrupts recent memory but leaves remote memory intact. The account presented here suggests that memories are first stored via synaptic changes in the hippocampal system, that these changes support reinstatement of recent memories in the neocortex, that neocortical synapses change a little on each reinstatement, and that remote memory is based on accumulated neocortical changes. Models that learn via changes to co…
Understanding normal and impaired word reading: Computational principles in quasi-regular domains.
A connectionist approach to processing in quasi-regular domains, as exemplified by English word reading, is developed. Networks using appropriately structured orthographic and phonological representations were trained to read both regular and exception words, and yet were also able to read pronounceable nonwords as well as skilled readers. A mathematical analysis of a simplified system clarifies the close relationship of word frequency and spelli…
The time course of perceptual choice: The leaky, competing accumulator model.
The time course of perceptual choice is discussed in a model of gradual, leaky, stochastic, and competitive information accumulation in nonlinear decision units. Special cases of the model match a classical diffusion process, but leakage and competition work together to address several challenges to existing diffusion, random walk, and accumulator models. The model accounts for data from choice tasks using both time-controlled (e.g., response sig…
Alternatives to the combinatorial paradigm of linguistic theory based on domain general principles of human cognition
It is argued that the principles needed to explain linguistic behavior are domain-general and based on the impact that specific experiences have on the mental organization and representation of language. This organization must be sensitive to both specific information and generalized patterns. In addition, knowledge of language is highly sensitive to frequency of use: frequently-used linguistic sequences become more frequent, more accessible and …
Language is not Just for Talking: Redundant Labels Facilitate Learning of Novel Categories
In addition to having communicative functions, verbal labels may play a role in shaping concepts. Two experiments assessed whether the presence of labels affected category formation. Subjects learned to categorize “aliens” as those to be approached or those to be avoided. After accuracy feedback on each response was provided, a nonsense label was either presented or not. Providing nonsense category labels facilitated category learning even though…
Gradience of Gradience: A reply to Jackendoff
Jackendoff and other linguists have acknowledged that there is gradience in language but have tended to treat gradient phenomena as separate from the core of language, which is viewed as fully productive and compositional. This perspective suffuses Jackendoff's (2007) response to our position paper (Bybee and McClelland 2005). We argue that gradience is an inherent feature of language representation, processing, and learning, and that natural lan…
Predicting native English-like performance by native Japanese speakers
Connectionism and the Emergence of Mind
Cognitive Neuroscience
Modelling the N400 brain potential as change in a probabilistic representation of meaning
Exemplar models are useful and deep neural networks overcome their limitations: A commentary on Ambridge (2020)
Humans are sensitive to the properties of individual items, and exemplar models are useful for capturing this sensitivity. I am a proponent of an extension of exemplar-based architectures that I briefly describe. However, exemplar models are very shallow architectures in which it is necessary to stipulate a set of primitive elements that make up each example, and such architectures have not been as successful as deep neural networks in capturing …
Do estimates of numerosity really adhere to Weber’s law? A reexamination of two case studies
Both humans and nonhuman animals can exhibit sensitivity to the approximate number of items in a visual array or events in a sequence, and across various paradigms, uncertainty in numerosity judgments increases with the number estimated or produced. The pattern of increase is usually described as exhibiting approximate adherence to Weber’s law, such that uncertainty increases proportionally to the mean estimate, resulting in a constant coefficien…
Systematic Human Learning and Generalization From a Brief Tutorial With Explanatory Feedback
We investigate human adults’ ability to learn an abstract reasoning task quickly and to generalize outside of the range of training examples. Using a task based on a solution strategy in Sudoku, we provide Sudoku-naive participants with a brief instructional tutorial with explanatory feedback using a narrow range of training examples. We find that most participants who master the task do so within 10 practice trials and generalize well to puzzles…
Learning to Decompose: Human-Like Subgoal Preferences Emerge in Neural Networks Learning Graph Traversal
Cognitive scientists have discovered normative and heuristic principles that capture human subgoal preferences when partitioning problems into smaller ones. However, it remains unclear where such preferences come from and why they tend to be both effective and efficient. In this work, we study the processes through which these preferences may be implicitly encoded over learning as learners improve towards optimal traversals. We build on the graph…
Psychology (18 works) · Computer Science (16 works) · Cognition (12 works) · Cognitive psychology (11 works) · Cognitive science (10 works) · Neuroscience (10 works) · Artificial neural network (8 works) · Connectionism (8 works) · Artificial Intelligence (7 works) · Linguistics (7 works)