Harald Hammarström
Datos Biográficos
| ID | 202116 |
|---|---|
| NOMBRE | Harald Hammarström |
| NOMBRES | Harald |
| APELLIDO | Hammarström |
| FIRMA | HAMMARSTRÖM H |
| AFILIACIONES | Uppsala University |
| ORCID | 0000-0003-0120-6396 |
| VERIFICADO | Sí |
| TOTAL DE OBRAS | 21 |
| TOTAL DE CITAS | 44 |
| TOTAL COMO AUTOR | 21 |
| TOTAL COMO EDITOR | 0 |
| PRIMER AÑO DE PUBLICACIÓN | 1920 |
| AÑO MÁS RECIENTE DE PUBLICACIÓN | 2025 |
| ÍNDICE H | 3 |
Commentary
The authors (Becker and Guzmn Naranjo 2025) frame their study as replication (of the methods part) of four previous studies in quantitative typology.They have adopted a rather broad definition of the term replication, and what they do more specifically is to test the robustness of these studies against a different method.They characterize robustness as the stability of the result of a study "across different methods used for analysis" and "under …
Evaluating the validity of census data for tracking speaker numbers
Bibliographic bias and information-density sampling
In the present paper, we discuss the bibliographical limits for commonplace typological studies and address how to estimate the resources available for an in-depth study using a full-text corpus of grammatical descriptions, considering different metalanguages, temporal stages of description, theoretical perspectives, and quality of grammatical descriptions. In a case study on motion, we illustrate the above perspectives and show how computer-assi…
The dialect chain tree
A perennial conflict in historical linguistics centers around the theoretical and practical virtues of tree-like divergence and wave-like diffusion. This paper presents the Dialect Chain Tree, an extension of the tree model that incorporates both tree-like descent and disintegration of dialect chains in a systematic fashion. As such, it provides a formalization and sharpening of Ross’ ( 1997 : 212–228) linkage concept that allows integration into…
Likelihood calculation in a multistate model of vocabulary evolution for linguistic dating
Computational methods of language dating make inferences about the divergence times of protolanguages by evaluating the patterns of inheritance in the vocabulary of modern languages, given the specification of a model of vocabulary evolution. We consider a model that describes vocabulary evolution as the replacement of traits by new traits from an infinite state space along a tree. This model has been introduced in previous literature but so far …
Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss
While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here we outline the Grambank database. With over 400,000 data points and 2,400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on t…
Defining numeral classifiers and identifying classifier languages of the world
This paper presents a precise definition of numeral classifiers, steps to identify a numeral classifier language, and a database of 3,338 languages, of which 723 languages have been identified as having a numeral classifier system. The database, named World Atlas of Classifier Languages (WACL), has been systematically constructed over the last 10 years via a manual survey of relevant literature and also an automatic scan of digitized grammars fol…
Expansion by migration and diffusion by contact is a source to the global diversity of linguistic nominal categorization systems
Languages of diverse structures and different families tend to share common patterns if they are spoken in geographic proximity. This convergence is often explained by horizontal diffusibility, which is typically ascribed to language contact. In such a scenario, speakers of two or more languages interact and influence each other’s languages, and in this interaction, more grammaticalized features tend to be more resistant to diffusion compared to …
On computational historical linguistics in the 21st century
replicability, rigorous evaluation, separation of training and test data, and only
Obsolescencia lingüística, descripción gramatical y documentación de lenguas en el Perú
Following the methods and tools developed by Hammarstrm, Castermans, Forkel et al. (2018) for the simultaneous visualization of the vitality status and degree of documentation of the world's languages, this paper provides a quantitative and qualitative analysis of the achievements and the challenges in the documentation and description of Peruvian languages. We attempt to determine the real dimensions of our understanding of the linguistic divers…
Language documentation twenty-five years on
This discussion note reviews responses of the linguistics profession to the grave issues of language endangerment identified a quarter of a century ago in the journal Language by Krauss, Hale, England, Craig, and others (Hale et al. 1992). Two and a half decades of worldwide research not only have given us a much more accurate picture of the number, phylogeny, and typological variety of the world's languages, but they have also seen the developme…
Ethnologue 16/17/18th editions
This section lists languages which are missing from E16/E17/E18. To be more precise, a language is listed here as missing if: • Extant published literature can make a convincing case that the language exists (or existed, see below), and, • Extant published literature can make a convincing case that the language is not intel-ligible to any language already listed in E16/E17/E18, and, An important note is that we do not list languages which are mis…
Ethnologue 16/17/18th editions
Ethnologue (http://www.ethnologue.com) is the most widely consulted inventory of the world's languages used today. The present review article looks carefully at the goals and description of the content of the Ethnologue ‘s 16th, 17th, and 18th editions, and reports on a comprehensive survey of the accuracy of the inventory itself. While hundreds of spurious and missing languages can be documented for Ethnologue , it is at present still better tha…
Quantifying Geographical Determinants of Large-Scale Distributions of Linguistic Features
In the recent past the work on large-scale linguistic distributions across the globe has intensified considerably. Work on macro-areal relationships in Africa (Güldemann, 2010) suggests that the shape of convergence areas may be determined by climatic factors and geophysical features such as mountains, water bodies, coastlines, etc. Worldwide data is now available for geophysical features as well as linguistic features, including numeral systems …
Some Principles on the Use of Macro-Areas in Typological Comparison
While the notion of the ‘area’ or ‘Sprachbund’ has a long history in linguistics, with geographically-defined regions frequently cited as a useful means to explain typological distributions, the problem of delimiting areas has not been well addressed. Lists of general-purpose, largely independent ‘macro-areas’ (typically continent size) have been proposed as a step to rule out contact as an explanation for various large-scale linguistic phenomena…
Unsupervised Learning of Morphology
This article surveys work on Unsupervised Learning of Morphology. We define Unsupervised Learning of Morphology as the problem of inducing a description (of some kind, even if only morpheme-segmentation) of how orthographic words are built up given only raw text data of a language. We briefly go through the history and motivation of the this problem. Next, over 200 items of work are listed with a brief characterization, and the most important ide…
Automated Dating of the World's Language Families Based on Lexical Similarity
This paper describes a computerized alternative to glottochronology for estimating elapsed time since parent languages diverged into daughter languages. The method, developed by the Automated Similarity Judgment Program (ASJP) consortium, is different from glottochronology in four major respects: (1) it is automated and thus is more objective, (2) it applies a uniform analytical approach to a single database of worldwide languages, (3) it is base…
The Language Families of the World
A full-scale test of the language farming dispersal hypothesis
One attempt at explaining why some language families are large (while others are small) is the hypothesis that the families that are now large became large because their ancestral speakers had a technological advantage, most often agriculture. Variants of this idea are referred to as the Language Farming Dispersal Hypothesis. Previously, detailed language family studies have uncovered various supporting examples and counterexamples to this idea. …
Properties of Lower Numerals and their Explanation
In a recent article (Rutkowski 2003) Paweł Rutkowski argues that the numerals 1-4 are treated specially in their syntax in across languages. Rutkowski wishes to explain this contrast as due to the working memory’s limited capacity, which cognitivists argue is indeed four items. We challenge this claim by presentation typological data to show a decreasing tendency of lower numerals to be more idiosyncratic and that 4 is a soft and arbitrary cut-of…
The Native Languages of South America
In South America indigenous languages are extremely diverse. There are over one hundred language families in this region alone. Contributors from around the world explore the history and structure of these languages, combining insights from archaeology and genetics with innovative linguistic analysis. The book aims to uncover regional patterns and potential deeper genealogical relations between the languages. Based on a large-scale database of fe…
Automated Dating of the World's Language Families Based on Lexical Similarity
This paper describes a computerized alternative to glottochronology for estimating elapsed time since parent languages diverged into daughter languages. The method, developed by the Automated Similarity Judgment Program (ASJP) consortium, is different from glottochronology in four major respects: (1) it is automated and thus is more objective, (2) it applies a uniform analytical approach to a single database of worldwide languages, (3) it is base…
Language documentation twenty-five years on
This discussion note reviews responses of the linguistics profession to the grave issues of language endangerment identified a quarter of a century ago in the journal Language by Krauss, Hale, England, Craig, and others (Hale et al. 1992). Two and a half decades of worldwide research not only have given us a much more accurate picture of the number, phylogeny, and typological variety of the world's languages, but they have also seen the developme…
Unsupervised Learning of Morphology
This article surveys work on Unsupervised Learning of Morphology. We define Unsupervised Learning of Morphology as the problem of inducing a description (of some kind, even if only morpheme-segmentation) of how orthographic words are built up given only raw text data of a language. We briefly go through the history and motivation of the this problem. Next, over 200 items of work are listed with a brief characterization, and the most important ide…
Defining numeral classifiers and identifying classifier languages of the world
This paper presents a precise definition of numeral classifiers, steps to identify a numeral classifier language, and a database of 3,338 languages, of which 723 languages have been identified as having a numeral classifier system. The database, named World Atlas of Classifier Languages (WACL), has been systematically constructed over the last 10 years via a manual survey of relevant literature and also an automatic scan of digitized grammars fol…
The Native Languages of South America
In South America indigenous languages are extremely diverse. There are over one hundred language families in this region alone. Contributors from around the world explore the history and structure of these languages, combining insights from archaeology and genetics with innovative linguistic analysis. The book aims to uncover regional patterns and potential deeper genealogical relations between the languages. Based on a large-scale database of fe…
Properties of Lower Numerals and their Explanation
In a recent article (Rutkowski 2003) Paweł Rutkowski argues that the numerals 1-4 are treated specially in their syntax in across languages. Rutkowski wishes to explain this contrast as due to the working memory’s limited capacity, which cognitivists argue is indeed four items. We challenge this claim by presentation typological data to show a decreasing tendency of lower numerals to be more idiosyncratic and that 4 is a soft and arbitrary cut-of…
The Language Families of the World
A full-scale test of the language farming dispersal hypothesis
One attempt at explaining why some language families are large (while others are small) is the hypothesis that the families that are now large became large because their ancestral speakers had a technological advantage, most often agriculture. Variants of this idea are referred to as the Language Farming Dispersal Hypothesis. Previously, detailed language family studies have uncovered various supporting examples and counterexamples to this idea. …
Unsupervised Learning of Morphology
This article surveys work on Unsupervised Learning of Morphology. We define Unsupervised Learning of Morphology as the problem of inducing a description (of some kind, even if only morpheme-segmentation) of how orthographic words are built up given only raw text data of a language. We briefly go through the history and motivation of the this problem. Next, over 200 items of work are listed with a brief characterization, and the most important ide…
Automated Dating of the World's Language Families Based on Lexical Similarity
This paper describes a computerized alternative to glottochronology for estimating elapsed time since parent languages diverged into daughter languages. The method, developed by the Automated Similarity Judgment Program (ASJP) consortium, is different from glottochronology in four major respects: (1) it is automated and thus is more objective, (2) it applies a uniform analytical approach to a single database of worldwide languages, (3) it is base…
Quantifying Geographical Determinants of Large-Scale Distributions of Linguistic Features
In the recent past the work on large-scale linguistic distributions across the globe has intensified considerably. Work on macro-areal relationships in Africa (Güldemann, 2010) suggests that the shape of convergence areas may be determined by climatic factors and geophysical features such as mountains, water bodies, coastlines, etc. Worldwide data is now available for geophysical features as well as linguistic features, including numeral systems …
Some Principles on the Use of Macro-Areas in Typological Comparison
While the notion of the ‘area’ or ‘Sprachbund’ has a long history in linguistics, with geographically-defined regions frequently cited as a useful means to explain typological distributions, the problem of delimiting areas has not been well addressed. Lists of general-purpose, largely independent ‘macro-areas’ (typically continent size) have been proposed as a step to rule out contact as an explanation for various large-scale linguistic phenomena…
Ethnologue 16/17/18th editions
This section lists languages which are missing from E16/E17/E18. To be more precise, a language is listed here as missing if: • Extant published literature can make a convincing case that the language exists (or existed, see below), and, • Extant published literature can make a convincing case that the language is not intel-ligible to any language already listed in E16/E17/E18, and, An important note is that we do not list languages which are mis…
Ethnologue 16/17/18th editions
Ethnologue (http://www.ethnologue.com) is the most widely consulted inventory of the world's languages used today. The present review article looks carefully at the goals and description of the content of the Ethnologue ‘s 16th, 17th, and 18th editions, and reports on a comprehensive survey of the accuracy of the inventory itself. While hundreds of spurious and missing languages can be documented for Ethnologue , it is at present still better tha…
Language documentation twenty-five years on
This discussion note reviews responses of the linguistics profession to the grave issues of language endangerment identified a quarter of a century ago in the journal Language by Krauss, Hale, England, Craig, and others (Hale et al. 1992). Two and a half decades of worldwide research not only have given us a much more accurate picture of the number, phylogeny, and typological variety of the world's languages, but they have also seen the developme…
On computational historical linguistics in the 21st century
replicability, rigorous evaluation, separation of training and test data, and only
Obsolescencia lingüística, descripción gramatical y documentación de lenguas en el Perú
Following the methods and tools developed by Hammarstrm, Castermans, Forkel et al. (2018) for the simultaneous visualization of the vitality status and degree of documentation of the world's languages, this paper provides a quantitative and qualitative analysis of the achievements and the challenges in the documentation and description of Peruvian languages. We attempt to determine the real dimensions of our understanding of the linguistic divers…
Expansion by migration and diffusion by contact is a source to the global diversity of linguistic nominal categorization systems
Languages of diverse structures and different families tend to share common patterns if they are spoken in geographic proximity. This convergence is often explained by horizontal diffusibility, which is typically ascribed to language contact. In such a scenario, speakers of two or more languages interact and influence each other’s languages, and in this interaction, more grammaticalized features tend to be more resistant to diffusion compared to …
Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss
While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here we outline the Grambank database. With over 400,000 data points and 2,400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on t…
Defining numeral classifiers and identifying classifier languages of the world
This paper presents a precise definition of numeral classifiers, steps to identify a numeral classifier language, and a database of 3,338 languages, of which 723 languages have been identified as having a numeral classifier system. The database, named World Atlas of Classifier Languages (WACL), has been systematically constructed over the last 10 years via a manual survey of relevant literature and also an automatic scan of digitized grammars fol…
Bibliographic bias and information-density sampling
In the present paper, we discuss the bibliographical limits for commonplace typological studies and address how to estimate the resources available for an in-depth study using a full-text corpus of grammatical descriptions, considering different metalanguages, temporal stages of description, theoretical perspectives, and quality of grammatical descriptions. In a case study on motion, we illustrate the above perspectives and show how computer-assi…
The dialect chain tree
A perennial conflict in historical linguistics centers around the theoretical and practical virtues of tree-like divergence and wave-like diffusion. This paper presents the Dialect Chain Tree, an extension of the tree model that incorporates both tree-like descent and disintegration of dialect chains in a systematic fashion. As such, it provides a formalization and sharpening of Ross’ ( 1997 : 212–228) linkage concept that allows integration into…
Likelihood calculation in a multistate model of vocabulary evolution for linguistic dating
Computational methods of language dating make inferences about the divergence times of protolanguages by evaluating the patterns of inheritance in the vocabulary of modern languages, given the specification of a model of vocabulary evolution. We consider a model that describes vocabulary evolution as the replacement of traits by new traits from an infinite state space along a tree. This model has been introduced in previous literature but so far …
Commentary
The authors (Becker and Guzmn Naranjo 2025) frame their study as replication (of the methods part) of four previous studies in quantitative typology.They have adopted a rather broad definition of the term replication, and what they do more specifically is to test the robustness of these studies against a different method.They characterize robustness as the stability of the result of a study "across different methods used for analysis" and "under …
Evaluating the validity of census data for tracking speaker numbers
Linguistics (16 obras) · Computer Science (15 obras) · Language and cultural evolution (13 obras) · Philosophy (11 obras) · Linguistic Variation and Morphology (10 obras) · Geography (8 obras) · History (7 obras) · Natural Language Processing Techniques (7 obras) · Artificial Intelligence (5 obras) · Mathematics (5 obras)