John Nerbonne
Biographic Data
| ID | 737811 |
|---|---|
| NAME | John Nerbonne |
| GIVEN NAMES | John |
| FAMILY NAME | Nerbonne |
| SIGNATURE | NERBONNE J |
| AFFILIATIONS | University of Groningen |
| ORCID | 0000-0002-3432-675X |
| VERIFIED | Yes |
| TOTAL WORKS | 25 |
| TOTAL CITATIONS | 66 |
| AUTHOR COUNT | 24 |
| EDITOR COUNT | 1 |
| FIRST PUBLICATION YEAR | 1986 |
| LATEST PUBLICATION YEAR | 2026 |
| H-INDEX | 6 |
Honoring Past Successes and Embracing New Opportunities in Linguistic Research: Languages Broadens Its Scope
Just as Anthony entered his second year as Co-Editor-in-Chief of Languages, he extended a warm welcome to John Nerbonne as Co-Editor-in-Chief beginning in 2026 [...]
Dialectal Dynamics-An Introduction
The study of dialects leads very naturally to the study of their geographic distribution and the nature of the distribution, e.g., by examining whether the distribution is based simply on geographic distance or on relatively distinct dialect regions. Dialectal dynamics poses the further question of why the distribution takes the form it does. Does variation arise through migration, i.e., due to the relative lack of communication among people who …
Extracting Tuscan phonetic correspondences from dialect pronunciations automatically
We present a novel approach to identifying individual pairs of phonetic correspondences in a dataset of dialect pronunciations. This continues work identifying shibboleths (i.e., characteristic features of a given dialect), a category that has interested dialectology and that dialectometrical research has examined mostly in the form of categorical data or entire phonetic transcriptions. This article reaches into segmental sequences (phonetic tran…
Multivariate Analysis of Geography and Age in Dialect Vocabulary — Comprehensive Analysis of 250 Years of Language Change
In this paper, we analyze long-term changes in dialect vocabulary. By analyzing the age difference of 140 years and the regional difference of 27 survey points, we will examine historical change and geographical spread over the 250 years since the compilation of a dialect glossary. First, we present the distribution of eight representative words according to the traditional method of linguistic geography. After that, we present comprehensive maps…
Unifying Analyses of Multiple Responses
In dialectology we often encounter irreducible variation in its data, i.e., multiple responses to its probes about the form of a word or phrase. Dialectometry seeks to measure the differences between dialects and has developed several ways to measure the difference between responses when one or both of them is non-unique. We introduce here BILBAO DISTANCE, where the cardinality of response is unimportant, which may be combined with various weight…
Handbook of Dialectology
The Handbook of Dialectology
The future of dialects: Selected papers from Methods in Dialectology XV
Traditional dialects have been encroached upon by the increasing mobility of their speakers and by the onslaught of national languages in education and mass media. Typically, older dialects are “leveling” to become more like national languages. This is regrettable when the last articulate traces of a culture are lost, but it also promotes a complex dynamics of interaction as speakers shift from dialect to standard and to intermediate compromises …
A Central Asian Language Survey: Collecting Data, Measuring Relatedness and Detecting Loans
We have documented language varieties (either Turkic or Indo-European) spoken in 23 test sites by 88 informants belonging to the major ethnic groups of Kyrgyzstan, Tajikistan and Uzbekistan (Karakalpaks, Kazakhs, Kyrgyz, Tajiks, Uzbeks, Yaghnobis). The recorded linguistic material concerns 176 words of the extended Swadesh list and will be made publically available with the publication of this paper. Phonological diversity is measured by the Leve…
Using Gabmap
New Perspectives on Language Variety in the South: Historical and Contemporary Approaches
Measuring Foreign Accent Strength in English: Validating Levenshtein Distance as a Measure
With an eye toward measuring the strength of foreign accents in American English, we evaluate the suitability of a modified version of the Levenshtein distance for comparing (the phonetic transcriptions of) accented pronunciations. Although this measure has been used successfully inter alia to study the differences among dialect pronunciations, it has not been applied to studying foreign accents. Here, we use it to compare the pronunciation of no…
Advances in Dialectometry
Dialectometry applies computational and statistical analyses within dialectology, making work more easily replicable and understandable. This survey article first reviews the field briefly in order to focus on developments in the past five years. Dialectometry no longer focuses exclusively on aggregate analyses, but rather deploys various techniques to identify representative and distinctive features with respect to areal classifications. Analyse…
Lexical differences between Tuscan dialects and standard Italian: Accounting for geographic and sociodemographic variation using generalized additive mixed modeling
This study uses a generalized additive mixed-effects regression model to predict lexical differences in Tuscan dialects with respect to standard Italian. We used lexical information for 170 concepts used by 2,060 speakers in 213 locations in Tuscany. In our model, geographical position was found to be an important predictor, with locations more distant from Florence having lexical forms more likely to differ from standard Italian. In addition, th…
Inducing a measure of phonetic similarity from pronunciation variation
Data-Driven Dialectology
Most studies of language variation proceed from the geographic or social distribution of single elements (features), and find it difficult to proceed further. Data-driven dialectology, and more generally, data-driven variationist studies, begin instead from an aggregate view of language variation and reap immediate benefits in dealing with well-known exceptions in the distributions of single features and in avoiding the need to select which featu…
Associations among linguistic levels
The forests behind the trees
Do Surname Differences Mirror Dialect Variation
Our focus in this paper is the analysis of surnames, which have been proven to be reliable genetic markers because in patrilineal systems they are transmitted along generations virtually unchanged, similarly to a genetic locus on the Y chromosome. We compare the distribution of surnames to the distribution of dialect pronunciations, which are clearly culturally transmitted. Because surnames, at the time of their introduction, were words subject t…
Computational Comparison and Classification of Dialects
In this paper a range of methods for measuring the phonetic distance between dialectal variants are described. It concerns variants of the frequency method, the frequency per word method and Levenshtein distance, both simple (based on atomic characters) and complex (based on feature bundles). The measurements between feature bundles used Manhattan distance, Euclidean distance or (a measure using) Pearson’s correlation coefficient. Variants of the…
Dialect areas and dialect continua
The organizing concept behind dialect variation is still seen predominantly as the areas within which similar varieties are spoken. The opposing view—that dialects are organized in a continuum without sharp boundaries—is likewise popular. This article introduces a new element into the discussion, which is the opportunity to view dialectal differences in the aggregate. We employ a dialectometric technique that provides an additive measure of pronu…
Linguistic Databases
German in Head-Driven Phrase Structure Grammar
1. Complement inheritance as subcategorization Dale Gerdemann 2. Argument structure and case assignment in German Wolfgang Heinz and Johannes Matiasek 3. Linearizing AUXs in German verbal complexes Erhard Hinrichs and Tsuneko Nakazawa 4. Adjuncts in the Mittelfeld Robert Kasper 5. Passives without lexical rules Andreas Kathol 6. Obligatory coherence: the structure of German modal verb constructions Tibor Kiss 7. Idioms and support verb constructi…
Phantoms’ and German fronting: Poltergeist constituents
In categorial grammar (CG), required complements such as dative and accusative objects, prepositional phrases, predicatives, and adverbials are added one at a time to lexical verbs. This leads to a question about the significance of phrases generated as intermediate steps in CG derivations. That is, while verbs (with NO included complements) and verb phrases (with ALL included complements) are clearly significant units, what about partial verb ph…
A phrase-structure grammar for German passives
Personal and impersonal variants of the German werden passives are examined and argued to be (1) subjectless in the impersonal case and (2) lexically formed. A rule introducing these is formulated in GPSG and shown to account for (1) the evidence that indicates that impersonal passives are subjectless, in particular, the behavior of matrix-initial zs; and (2) the evidence that indicates a lexical rule, in particular (a) the various constituent st…
Lexical differences between Tuscan dialects and standard Italian: Accounting for geographic and sociodemographic variation using generalized additive mixed modeling
This study uses a generalized additive mixed-effects regression model to predict lexical differences in Tuscan dialects with respect to standard Italian. We used lexical information for 170 concepts used by 2,060 speakers in 213 locations in Tuscany. In our model, geographical position was found to be an important predictor, with locations more distant from Florence having lexical forms more likely to differ from standard Italian. In addition, th…
Inducing a measure of phonetic similarity from pronunciation variation
Advances in Dialectometry
Dialectometry applies computational and statistical analyses within dialectology, making work more easily replicable and understandable. This survey article first reviews the field briefly in order to focus on developments in the past five years. Dialectometry no longer focuses exclusively on aggregate analyses, but rather deploys various techniques to identify representative and distinctive features with respect to areal classifications. Analyse…
Dialect areas and dialect continua
The organizing concept behind dialect variation is still seen predominantly as the areas within which similar varieties are spoken. The opposing view—that dialects are organized in a continuum without sharp boundaries—is likewise popular. This article introduces a new element into the discussion, which is the opportunity to view dialectal differences in the aggregate. We employ a dialectometric technique that provides an additive measure of pronu…
Data-Driven Dialectology
Most studies of language variation proceed from the geographic or social distribution of single elements (features), and find it difficult to proceed further. Data-driven dialectology, and more generally, data-driven variationist studies, begin instead from an aggregate view of language variation and reap immediate benefits in dealing with well-known exceptions in the distributions of single features and in avoiding the need to select which featu…
Associations among linguistic levels
Using Gabmap
The forests behind the trees
Phantoms’ and German fronting: Poltergeist constituents
In categorial grammar (CG), required complements such as dative and accusative objects, prepositional phrases, predicatives, and adverbials are added one at a time to lexical verbs. This leads to a question about the significance of phrases generated as intermediate steps in CG derivations. That is, while verbs (with NO included complements) and verb phrases (with ALL included complements) are clearly significant units, what about partial verb ph…
A phrase-structure grammar for German passives
Personal and impersonal variants of the German werden passives are examined and argued to be (1) subjectless in the impersonal case and (2) lexically formed. A rule introducing these is formulated in GPSG and shown to account for (1) the evidence that indicates that impersonal passives are subjectless, in particular, the behavior of matrix-initial zs; and (2) the evidence that indicates a lexical rule, in particular (a) the various constituent st…
German in Head-Driven Phrase Structure Grammar
1. Complement inheritance as subcategorization Dale Gerdemann 2. Argument structure and case assignment in German Wolfgang Heinz and Johannes Matiasek 3. Linearizing AUXs in German verbal complexes Erhard Hinrichs and Tsuneko Nakazawa 4. Adjuncts in the Mittelfeld Robert Kasper 5. Passives without lexical rules Andreas Kathol 6. Obligatory coherence: the structure of German modal verb constructions Tibor Kiss 7. Idioms and support verb constructi…
Linguistic Databases
Computational Comparison and Classification of Dialects
In this paper a range of methods for measuring the phonetic distance between dialectal variants are described. It concerns variants of the frequency method, the frequency per word method and Levenshtein distance, both simple (based on atomic characters) and complex (based on feature bundles). The measurements between feature bundles used Manhattan distance, Euclidean distance or (a measure using) Pearson’s correlation coefficient. Variants of the…
Dialect areas and dialect continua
The organizing concept behind dialect variation is still seen predominantly as the areas within which similar varieties are spoken. The opposing view—that dialects are organized in a continuum without sharp boundaries—is likewise popular. This article introduces a new element into the discussion, which is the opportunity to view dialectal differences in the aggregate. We employ a dialectometric technique that provides an additive measure of pronu…
Do Surname Differences Mirror Dialect Variation
Our focus in this paper is the analysis of surnames, which have been proven to be reliable genetic markers because in patrilineal systems they are transmitted along generations virtually unchanged, similarly to a genetic locus on the Y chromosome. We compare the distribution of surnames to the distribution of dialect pronunciations, which are clearly culturally transmitted. Because surnames, at the time of their introduction, were words subject t…
Data-Driven Dialectology
Most studies of language variation proceed from the geographic or social distribution of single elements (features), and find it difficult to proceed further. Data-driven dialectology, and more generally, data-driven variationist studies, begin instead from an aggregate view of language variation and reap immediate benefits in dealing with well-known exceptions in the distributions of single features and in avoiding the need to select which featu…
Associations among linguistic levels
The forests behind the trees
Inducing a measure of phonetic similarity from pronunciation variation
Measuring Foreign Accent Strength in English: Validating Levenshtein Distance as a Measure
With an eye toward measuring the strength of foreign accents in American English, we evaluate the suitability of a modified version of the Levenshtein distance for comparing (the phonetic transcriptions of) accented pronunciations. Although this measure has been used successfully inter alia to study the differences among dialect pronunciations, it has not been applied to studying foreign accents. Here, we use it to compare the pronunciation of no…
Advances in Dialectometry
Dialectometry applies computational and statistical analyses within dialectology, making work more easily replicable and understandable. This survey article first reviews the field briefly in order to focus on developments in the past five years. Dialectometry no longer focuses exclusively on aggregate analyses, but rather deploys various techniques to identify representative and distinctive features with respect to areal classifications. Analyse…
Lexical differences between Tuscan dialects and standard Italian: Accounting for geographic and sociodemographic variation using generalized additive mixed modeling
This study uses a generalized additive mixed-effects regression model to predict lexical differences in Tuscan dialects with respect to standard Italian. We used lexical information for 170 concepts used by 2,060 speakers in 213 locations in Tuscany. In our model, geographical position was found to be an important predictor, with locations more distant from Florence having lexical forms more likely to differ from standard Italian. In addition, th…
New Perspectives on Language Variety in the South: Historical and Contemporary Approaches
The future of dialects: Selected papers from Methods in Dialectology XV
Traditional dialects have been encroached upon by the increasing mobility of their speakers and by the onslaught of national languages in education and mass media. Typically, older dialects are “leveling” to become more like national languages. This is regrettable when the last articulate traces of a culture are lost, but it also promotes a complex dynamics of interaction as speakers shift from dialect to standard and to intermediate compromises …
A Central Asian Language Survey: Collecting Data, Measuring Relatedness and Detecting Loans
We have documented language varieties (either Turkic or Indo-European) spoken in 23 test sites by 88 informants belonging to the major ethnic groups of Kyrgyzstan, Tajikistan and Uzbekistan (Karakalpaks, Kazakhs, Kyrgyz, Tajiks, Uzbeks, Yaghnobis). The recorded linguistic material concerns 176 words of the extended Swadesh list and will be made publically available with the publication of this paper. Phonological diversity is measured by the Leve…
Using Gabmap
The Handbook of Dialectology
Handbook of Dialectology
Unifying Analyses of Multiple Responses
In dialectology we often encounter irreducible variation in its data, i.e., multiple responses to its probes about the form of a word or phrase. Dialectometry seeks to measure the differences between dialects and has developed several ways to measure the difference between responses when one or both of them is non-unique. We introduce here BILBAO DISTANCE, where the cardinality of response is unimportant, which may be combined with various weight…
Multivariate Analysis of Geography and Age in Dialect Vocabulary — Comprehensive Analysis of 250 Years of Language Change
In this paper, we analyze long-term changes in dialect vocabulary. By analyzing the age difference of 140 years and the regional difference of 27 survey points, we will examine historical change and geographical spread over the 250 years since the compilation of a dialect glossary. First, we present the distribution of eight representative words according to the traditional method of linguistic geography. After that, we present comprehensive maps…
Extracting Tuscan phonetic correspondences from dialect pronunciations automatically
We present a novel approach to identifying individual pairs of phonetic correspondences in a dataset of dialect pronunciations. This continues work identifying shibboleths (i.e., characteristic features of a given dialect), a category that has interested dialectology and that dialectometrical research has examined mostly in the form of categorical data or entire phonetic transcriptions. This article reaches into segmental sequences (phonetic tran…
Dialectal Dynamics-An Introduction
The study of dialects leads very naturally to the study of their geographic distribution and the nature of the distribution, e.g., by examining whether the distribution is based simply on geographic distance or on relatively distinct dialect regions. Dialectal dynamics poses the further question of why the distribution takes the form it does. Does variation arise through migration, i.e., due to the relative lack of communication among people who …
Honoring Past Successes and Embracing New Opportunities in Linguistic Research: Languages Broadens Its Scope
Just as Anthony entered his second year as Co-Editor-in-Chief of Languages, he extended a warm welcome to John Nerbonne as Co-Editor-in-Chief beginning in 2026 [...]
Linguistics (21 works) · Computer Science (15 works) · Linguistic Variation and Morphology (14 works) · Philosophy (13 works) · Natural Language Processing Techniques (11 works) · Artificial Intelligence (10 works) · Mathematics (9 works) · Phonetics and Phonology Research (9 works) · Geography (8 works) · Natural language processing (7 works)