Richard Sproat
Biographic Data
| ID | 504227 |
|---|---|
| NAME | Richard Sproat |
| GIVEN NAMES | Richard |
| FAMILY NAME | Sproat |
| SIGNATURE | SPROAT R |
| AFFILIATIONS | Google, Inc. |
| ORCID | 0000-0002-9040-5196 |
| VERIFIED | Yes |
| TOTAL WORKS | 26 |
| TOTAL CITATIONS | 140 |
| AUTHOR COUNT | 26 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1985 |
| LATEST PUBLICATION YEAR | 2022 |
| H-INDEX | 3 |
Boring Problems Are Sometimes the Most Interesting
In a recent position paper, Turing Award Winners Yoshua Bengio, Geoffrey Hinton, and Yann LeCun make the case that symbolic methods are not needed in AI and that, while there are still many issues to be resolved, AI will be solved using purely neural methods. In this piece I issue a challenge: Demonstrate that a purely neural approach to the problem of text normalization is possible. Various groups have tried, but so far nobody has eliminated the…
Statistical evidence for the Proto-Indo-European-Euskarian hypothesis: A word-list approach integrating phonotactics
Based on a new reconstruction of Proto-Basque, and regular sound correspondences between this Proto-Basque and Proto-Indo-European as standardly reconstructed, Blevins (2018) argues that Proto-Basque and Proto-Indo-European have a common ancestor that pre-dates the two proto-languages. Part of this argument is based on proposed Proto-Indo-European/Proto-Basque cognate sets that include basic vocabulary items. In this study we offer statistical su…
The Taxonomy of Writing Systems: How to Measure How Logographic a System Is
Taxonomies of writing systems since Gelb (1952) have classified systems based on what the written symbols represent: if they represent words or morphemes, they are logographic; if syllables, syllabic; if segments, alphabetic; and so forth. Sproat (2000) and Rogers (2005) broke with tradition by splitting the logographic and phonographic aspects into two dimensions, with logography being graded rather than a categorical distinction. A system could…
Zev Handel: Sinography: The Borrowing and Adaptation of the Chinese Script
Neural Models of Text Normalization for Speech Applications
Machine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 Ki…
Peter T. Daniels, 2018, An Exploration of Writing
A computational model of the discovery of writing
This paper reports on a computational simulation of the evolution of early writing systems from pre-linguistic symbol systems, something for which there is poor evidence in the archaeological record. The simulation starts with a completely concept-based set of symbols, and then spreads those symbols and combinations of these to morphemes of artificially generated languages based on semantic and phonetic similarity. While the simulation is crude, …
On misunderstandings and misrepresentations: A reply to Rao et al
On misunderstandings and misrepresentations:A reply to Rao et al. Richard Sproat When Language forwarded me the reply from Rao, Lee, and colleagues (2015; henceforth Rao et al.) to my article in Language 90.2 (Sproat 2014), I was offered an opportunity to respond, but I was also instructed to be brief. Fortunately, it is easy to be brief since I believe my article lays out the case very well, and the reader need only look there and at Rao et al.’…
Applications of Lexicographic Semirings to Problems in Speech and Language Processing
This paper explores lexicographic semirings and their application to problems in speech and language processing. Specifically, we present two instantiations of binary lexicographic semirings, one involving a pair of tropical weights, and the other a tropical weight paired with a novel string semiring we term the categorial semiring. The first of these is used to yield an exact encoding of backoff models with epsilon transitions. This lexicographi…
A statistical comparison of written language and nonlinguistic symbol systems: Online Supplementary Materials
A statistical comparison of written language and nonlinguistic symbol systems
Are statistical methods useful in distinguishing written language from nonlinguistic symbol systems? Some recent articles (Rao et al. 2009a, Lee et al. 2010a) have claimed so. Both of these previous articles use measures based at least in part on bigram conditional entropy, and subsequent work by one of the authors (Rao) has used other entropic measures. In both cases the authors have argued that the methods proposed either are useful for discrim…
A note on Unger’s “What linguistic units do Chinese characters represent?”
Unger (2011) observes that Chinese characters do not observe a Zipfian distribution, and he uses this fact as evidence that Chinese characters do not represent words. He then goes on to suggest that they do not represent morphemes either. In this note I argue that Unger’s observation is neither new, nor is it necessary; and that, at least with respect to his claim about morphemes, it does not support the conclusion he wishes to make
Reply to Rao et al. and Lee et al
In the last issue of this journal, I presented a piece that called into question some of the techniques reported in two papers in high-profile journals that purported to provide statistical evidence for the linguistic status of some ancient symbol systems (Sproat 2010a). Not surprisingly, the authors of those two papers took issue with a number of my claims, and have requested the opportunity to respond. The two responses, taken together, are rat…
Ancient Symbols, Computational Linguistics, and the Reviewing Practices of the General Science Journals
Few archaeological finds are as evocative as artifacts inscribed with symbols. Whenever an archaeologist finds a potsherd or a seal impression that seems to have symbols scratched or impressed on the surface, it is natural to want to “read” the symbols. And if the symbols come from an undeciphered or previously unknown symbol system it is common to ask what language the symbols supposedly represent and whether the system can be deciphered. Of cou…
Model for phonemic awareness in readers of Indian script
Previous studies have shown that segmental awareness tasks are usually influenced by the script. In this paper, we extend these studies further to propose a more concrete, script-centric metric for evaluating phonemic awareness in readers of Indian scripts. We propose that the ease or difficulty with which syllabic and phonemic segmental tasks are performed is directly proportional to the editing operations involved in applying the same task on t…
Brahmi-derived scripts, script layout, and segmental awareness
In earlier work (Sproat 2000), I characterized the layout of symbols in a script in terms of a calculus involving two dimensional catenation operators: I claimed that leftwards, rightwards, upwards, downwards and surrounding catenation are sufficient to describe the layout of any script. In the first half of this paper I analyze four Indic alphasyllabaries — Devanagari, Oriya, Kannada and Tamil — in terms of this model. A crucial claim is that de…
A corpus-based analysis of Mandarin nominal root compound
Complex verb Formation
Computational Morphology: Practical Mechanisms for the English Lexicon
Part 1 Introduction: the need for a computational lexicon what is the lexicon? morphology segmentation and orthography word structure the unit of storage lexical redundancy preprocessing and look-up what this book is about. Part 2 Morphographemics: generative phonology formalizing phonological rules transducers the two-level model the transducer version the two-level rule notation the lexicon and word grammar interface formal issues. Part 3 Word …
Allophonic variation in English /l/ and its implications for phonetic implementation
Andrew Carstairs-McCarthy (1992). Current morphology. ( Linguistic Theory Guides ). London: Routledge. Pp. xiii + 289
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content
The syntax of the modern Celtic languages
A Pragmatic Analysis of So-Called Anaphoric Islands
It is commonly assumed that words are grammatically prohibited from containing antecedents for anaphoric elements, and thus constitute 'anaphoric islands' (Postal 1969). In this paper, we argue that such anaphora-termed OUTBOUND ANAPHORA -is in fact fully grammatical and governed by independently motivated pragmatic principles. The felicity of outbound anaphora is shown to be a function of the accessibility of the discourse entity which is evoked…
Bracketing Paradoxes, Cliticization and Other Topics: The Mapping between Syntactic and Phonological Structure
On Anaphoric Islandhood
Allophonic variation in English /l/ and its implications for phonetic implementation
Welsh syntax and VSO structure
A Pragmatic Analysis of So-Called Anaphoric Islands
It is commonly assumed that words are grammatically prohibited from containing antecedents for anaphoric elements, and thus constitute 'anaphoric islands' (Postal 1969). In this paper, we argue that such anaphora-termed OUTBOUND ANAPHORA -is in fact fully grammatical and governed by independently motivated pragmatic principles. The felicity of outbound anaphora is shown to be a function of the accessibility of the discourse entity which is evoked…
Ancient Symbols, Computational Linguistics, and the Reviewing Practices of the General Science Journals
Few archaeological finds are as evocative as artifacts inscribed with symbols. Whenever an archaeologist finds a potsherd or a seal impression that seems to have symbols scratched or impressed on the surface, it is natural to want to “read” the symbols. And if the symbols come from an undeciphered or previously unknown symbol system it is common to ask what language the symbols supposedly represent and whether the system can be deciphered. Of cou…
The Taxonomy of Writing Systems: How to Measure How Logographic a System Is
Taxonomies of writing systems since Gelb (1952) have classified systems based on what the written symbols represent: if they represent words or morphemes, they are logographic; if syllables, syllabic; if segments, alphabetic; and so forth. Sproat (2000) and Rogers (2005) broke with tradition by splitting the logographic and phonographic aspects into two dimensions, with logography being graded rather than a categorical distinction. A system could…
Neural Models of Text Normalization for Speech Applications
Machine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 Ki…
Welsh syntax and VSO structure
Bracketing Paradoxes, Cliticization and Other Topics: The Mapping between Syntactic and Phonological Structure
On Anaphoric Islandhood
A Pragmatic Analysis of So-Called Anaphoric Islands
It is commonly assumed that words are grammatically prohibited from containing antecedents for anaphoric elements, and thus constitute 'anaphoric islands' (Postal 1969). In this paper, we argue that such anaphora-termed OUTBOUND ANAPHORA -is in fact fully grammatical and governed by independently motivated pragmatic principles. The felicity of outbound anaphora is shown to be a function of the accessibility of the discourse entity which is evoked…
Andrew Carstairs-McCarthy (1992). Current morphology. ( Linguistic Theory Guides ). London: Routledge. Pp. xiii + 289
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content
The syntax of the modern Celtic languages
Computational Morphology: Practical Mechanisms for the English Lexicon
Part 1 Introduction: the need for a computational lexicon what is the lexicon? morphology segmentation and orthography word structure the unit of storage lexical redundancy preprocessing and look-up what this book is about. Part 2 Morphographemics: generative phonology formalizing phonological rules transducers the two-level model the transducer version the two-level rule notation the lexicon and word grammar interface formal issues. Part 3 Word …
Allophonic variation in English /l/ and its implications for phonetic implementation
Complex verb Formation
A corpus-based analysis of Mandarin nominal root compound
Brahmi-derived scripts, script layout, and segmental awareness
In earlier work (Sproat 2000), I characterized the layout of symbols in a script in terms of a calculus involving two dimensional catenation operators: I claimed that leftwards, rightwards, upwards, downwards and surrounding catenation are sufficient to describe the layout of any script. In the first half of this paper I analyze four Indic alphasyllabaries — Devanagari, Oriya, Kannada and Tamil — in terms of this model. A crucial claim is that de…
Model for phonemic awareness in readers of Indian script
Previous studies have shown that segmental awareness tasks are usually influenced by the script. In this paper, we extend these studies further to propose a more concrete, script-centric metric for evaluating phonemic awareness in readers of Indian scripts. We propose that the ease or difficulty with which syllabic and phonemic segmental tasks are performed is directly proportional to the editing operations involved in applying the same task on t…
Reply to Rao et al. and Lee et al
In the last issue of this journal, I presented a piece that called into question some of the techniques reported in two papers in high-profile journals that purported to provide statistical evidence for the linguistic status of some ancient symbol systems (Sproat 2010a). Not surprisingly, the authors of those two papers took issue with a number of my claims, and have requested the opportunity to respond. The two responses, taken together, are rat…
Ancient Symbols, Computational Linguistics, and the Reviewing Practices of the General Science Journals
Few archaeological finds are as evocative as artifacts inscribed with symbols. Whenever an archaeologist finds a potsherd or a seal impression that seems to have symbols scratched or impressed on the surface, it is natural to want to “read” the symbols. And if the symbols come from an undeciphered or previously unknown symbol system it is common to ask what language the symbols supposedly represent and whether the system can be deciphered. Of cou…
A note on Unger’s “What linguistic units do Chinese characters represent?”
Unger (2011) observes that Chinese characters do not observe a Zipfian distribution, and he uses this fact as evidence that Chinese characters do not represent words. He then goes on to suggest that they do not represent morphemes either. In this note I argue that Unger’s observation is neither new, nor is it necessary; and that, at least with respect to his claim about morphemes, it does not support the conclusion he wishes to make
Applications of Lexicographic Semirings to Problems in Speech and Language Processing
This paper explores lexicographic semirings and their application to problems in speech and language processing. Specifically, we present two instantiations of binary lexicographic semirings, one involving a pair of tropical weights, and the other a tropical weight paired with a novel string semiring we term the categorial semiring. The first of these is used to yield an exact encoding of backoff models with epsilon transitions. This lexicographi…
A statistical comparison of written language and nonlinguistic symbol systems: Online Supplementary Materials
A statistical comparison of written language and nonlinguistic symbol systems
Are statistical methods useful in distinguishing written language from nonlinguistic symbol systems? Some recent articles (Rao et al. 2009a, Lee et al. 2010a) have claimed so. Both of these previous articles use measures based at least in part on bigram conditional entropy, and subsequent work by one of the authors (Rao) has used other entropic measures. In both cases the authors have argued that the methods proposed either are useful for discrim…
On misunderstandings and misrepresentations: A reply to Rao et al
On misunderstandings and misrepresentations:A reply to Rao et al. Richard Sproat When Language forwarded me the reply from Rao, Lee, and colleagues (2015; henceforth Rao et al.) to my article in Language 90.2 (Sproat 2014), I was offered an opportunity to respond, but I was also instructed to be brief. Fortunately, it is easy to be brief since I believe my article lays out the case very well, and the reader need only look there and at Rao et al.’…
A computational model of the discovery of writing
This paper reports on a computational simulation of the evolution of early writing systems from pre-linguistic symbol systems, something for which there is poor evidence in the archaeological record. The simulation starts with a completely concept-based set of symbols, and then spreads those symbols and combinations of these to morphemes of artificially generated languages based on semantic and phonetic similarity. While the simulation is crude, …
Peter T. Daniels, 2018, An Exploration of Writing
Neural Models of Text Normalization for Speech Applications
Machine learning, including neural network techniques, have been applied to virtually every domain in natural language processing. One problem that has been somewhat resistant to effective machine learning solutions is text normalization for speech applications such as text-to-speech synthesis (TTS). In this application, one must decide, for example, that 123 is verbalized as one hundred twenty three in 123 pages but as one twenty three in 123 Ki…
Zev Handel: Sinography: The Borrowing and Adaptation of the Chinese Script
Statistical evidence for the Proto-Indo-European-Euskarian hypothesis: A word-list approach integrating phonotactics
Based on a new reconstruction of Proto-Basque, and regular sound correspondences between this Proto-Basque and Proto-Indo-European as standardly reconstructed, Blevins (2018) argues that Proto-Basque and Proto-Indo-European have a common ancestor that pre-dates the two proto-languages. Part of this argument is based on proposed Proto-Indo-European/Proto-Basque cognate sets that include basic vocabulary items. In this study we offer statistical su…
The Taxonomy of Writing Systems: How to Measure How Logographic a System Is
Taxonomies of writing systems since Gelb (1952) have classified systems based on what the written symbols represent: if they represent words or morphemes, they are logographic; if syllables, syllabic; if segments, alphabetic; and so forth. Sproat (2000) and Rogers (2005) broke with tradition by splitting the logographic and phonographic aspects into two dimensions, with logography being graded rather than a categorical distinction. A system could…
Computer Science (22 works) · Linguistics (19 works) · Philosophy (17 works) · Natural Language Processing Techniques (14 works) · Natural language processing (12 works) · Artificial Intelligence (10 works) · Psychology (7 works) · Linguistic Variation and Morphology (6 works) · Mathematics (6 works) · Philosophy (6 works)