Steven Coat
Dados Biográficos
| ID | 108673 |
|---|---|
| NOME | Steven Coat |
| PRENOMES | Steven |
| SOBRENOME | Coat |
| ASSINATURA | COAT S |
| AFILIAÇÕES | University of Oulu |
| ORCID | 0000-0002-7295-3893 |
| VERIFICADO | Sim |
| TOTAL DE OBRAS | 10 |
| TOTAL DE CITAÇÕES | 8 |
| TOTAL COMO AUTOR | 10 |
| TOTAL COMO EDITOR | 0 |
| PRIMEIRO ANO DE PUBLICAÇÃO | 2019 |
| ANO MAIS RECENTE DE PUBLICAÇÃO | 2025 |
| ÍNDICE H | 2 |
What the X’ in Anglophone government meetings
The YouTube corpus of Singapore English podcasts
Recent advances in streaming protocols and automatic speech recognition (ASR) have enabled large-scale spoken language corpora, yet research on Singapore English remains constrained by small or text-based datasets. The YouTube Corpus of Singapore English Podcasts (YCSEP) addresses this gap with 620 hours of transcribed, diarized speech from over 1,300 podcast episodes by Singapore-based content creators. YCSEP supports the empirical analysis of p…
Double modals beyond the Atlantic
This study investigates the use of double modals in Australian and New Zealand English using Twitter/X data. Double modals are rare grammatical constructions long believed to be limited to regional dialects in the Northern UK and the Southern US. Utilizing a geolocated corpus of over 80 million tweets, the study identifies 314 authentic double modal instances across 51 types, primarily occurring in informal tweets. Findings reveal widespread, alb…
Building a searchable online corpus of Australian and New Zealand aligned speech
Double modals in contemporary British and Irish speech
This article reports on the use of double modals, a non-standard syntactic feature, in the contemporary speech of the UK and Ireland. Most data on the geographic extent of the feature and its combinatorial types come from surveys or acceptability ratings or from older attestations focused on northern England, Scotland or Northern Ireland, with relatively few attestations in naturalistic data and from England and Wales. Manual verification of doub…
Double modals in Australian and New Zealand English
This paper reports the first large-scale corpus study of double modal usage in Australian and New Zealand Englishes, based on a multi-million-word corpus of geolocated automatic speech recognition transcripts from YouTube. Double modals are considered rare grammatical features of English, which have long been extremely difficult to observe in natural language due to low frequencies, non-standardness, and restriction to oral speech registers. In a…
Naturalistic Double Modals in North America
Double modals are a well-known nonstandard feature of some regional varieties of English in North America, but due to their rareness in spoken language, questions remain as to the inventory of possible combinatorial types and the geographic extent of their use in contemporary naturalistic speech. This study investigates double modals in the Corpus of North American Spoken English (CoNASE), a 1.2-billion-word corpus of time-stamped and geolocated …
‘Bad language’ in the Nordics
This study looks at the relative frequency of ‘bad language’ according to gender in Nordic languages and in English in a 210-million-token corpus of messages by 18,686 Nordic Twitter users. For the Nordic languages, more than 19,000 ‘bad-language’ word forms were compiled on the basis of usage note annotations in major Nordic-language dictionaries. The most frequent terms overall are swear words, and while males use more of these items on average…
Articulation Rate in American English in a Corpus of YouTube Videos
Previous studies of the temporal organization of speech in American English have found differences in speaking or articulation rate according to speaker dialect or location, but small sample sizes and incomplete geographic coverage have limited the generalizability of the findings. In this study, articulation rates in American English are calculated from the automatic speech-to-text transcripts of more than 29,000 hours of video from local govern…
Language choice and gender in a Nordic social media corpus
This study analyzes language choice, bi- and multilingualism, and gender in a corpus of over 22 million Twitter messages by almost 36,000 authors from the Nordic countries and territories. Author location, gender, and tweet language are identified using a novel method. Three principal findings are discussed: First, gendered preference for particular languages in the Nordics can be explained in part by patterns of gendered migration. Second, a dis…
Naturalistic Double Modals in North America
Double modals are a well-known nonstandard feature of some regional varieties of English in North America, but due to their rareness in spoken language, questions remain as to the inventory of possible combinatorial types and the geographic extent of their use in contemporary naturalistic speech. This study investigates double modals in the Corpus of North American Spoken English (CoNASE), a 1.2-billion-word corpus of time-stamped and geolocated …
Double modals in Australian and New Zealand English
This paper reports the first large-scale corpus study of double modal usage in Australian and New Zealand Englishes, based on a multi-million-word corpus of geolocated automatic speech recognition transcripts from YouTube. Double modals are considered rare grammatical features of English, which have long been extremely difficult to observe in natural language due to low frequencies, non-standardness, and restriction to oral speech registers. In a…
Articulation Rate in American English in a Corpus of YouTube Videos
Previous studies of the temporal organization of speech in American English have found differences in speaking or articulation rate according to speaker dialect or location, but small sample sizes and incomplete geographic coverage have limited the generalizability of the findings. In this study, articulation rates in American English are calculated from the automatic speech-to-text transcripts of more than 29,000 hours of video from local govern…
Language choice and gender in a Nordic social media corpus
This study analyzes language choice, bi- and multilingualism, and gender in a corpus of over 22 million Twitter messages by almost 36,000 authors from the Nordic countries and territories. Author location, gender, and tweet language are identified using a novel method. Three principal findings are discussed: First, gendered preference for particular languages in the Nordics can be explained in part by patterns of gendered migration. Second, a dis…
Articulation Rate in American English in a Corpus of YouTube Videos
Previous studies of the temporal organization of speech in American English have found differences in speaking or articulation rate according to speaker dialect or location, but small sample sizes and incomplete geographic coverage have limited the generalizability of the findings. In this study, articulation rates in American English are calculated from the automatic speech-to-text transcripts of more than 29,000 hours of video from local govern…
‘Bad language’ in the Nordics
This study looks at the relative frequency of ‘bad language’ according to gender in Nordic languages and in English in a 210-million-token corpus of messages by 18,686 Nordic Twitter users. For the Nordic languages, more than 19,000 ‘bad-language’ word forms were compiled on the basis of usage note annotations in major Nordic-language dictionaries. The most frequent terms overall are swear words, and while males use more of these items on average…
Naturalistic Double Modals in North America
Double modals are a well-known nonstandard feature of some regional varieties of English in North America, but due to their rareness in spoken language, questions remain as to the inventory of possible combinatorial types and the geographic extent of their use in contemporary naturalistic speech. This study investigates double modals in the Corpus of North American Spoken English (CoNASE), a 1.2-billion-word corpus of time-stamped and geolocated …
Double modals in contemporary British and Irish speech
This article reports on the use of double modals, a non-standard syntactic feature, in the contemporary speech of the UK and Ireland. Most data on the geographic extent of the feature and its combinatorial types come from surveys or acceptability ratings or from older attestations focused on northern England, Scotland or Northern Ireland, with relatively few attestations in naturalistic data and from England and Wales. Manual verification of doub…
Double modals in Australian and New Zealand English
This paper reports the first large-scale corpus study of double modal usage in Australian and New Zealand Englishes, based on a multi-million-word corpus of geolocated automatic speech recognition transcripts from YouTube. Double modals are considered rare grammatical features of English, which have long been extremely difficult to observe in natural language due to low frequencies, non-standardness, and restriction to oral speech registers. In a…
Double modals beyond the Atlantic
This study investigates the use of double modals in Australian and New Zealand English using Twitter/X data. Double modals are rare grammatical constructions long believed to be limited to regional dialects in the Northern UK and the Southern US. Utilizing a geolocated corpus of over 80 million tweets, the study identifies 314 authentic double modal instances across 51 types, primarily occurring in informal tweets. Findings reveal widespread, alb…
Building a searchable online corpus of Australian and New Zealand aligned speech
What the X’ in Anglophone government meetings
The YouTube corpus of Singapore English podcasts
Recent advances in streaming protocols and automatic speech recognition (ASR) have enabled large-scale spoken language corpora, yet research on Singapore English remains constrained by small or text-based datasets. The YouTube Corpus of Singapore English Podcasts (YCSEP) addresses this gap with 620 hours of transcribed, diarized speech from over 1,300 podcast episodes by Singapore-based content creators. YCSEP supports the empirical analysis of p…
Linguistic Variation and Morphology (8 obras) · Linguistics (8 obras) · Geography (5 obras) · Language, Discourse, Communication Strategies (5 obras) · Psychology (5 obras) · Computer Science (4 obras) · History (4 obras) · Modal verb (4 obras) · Multilingual Education and Policy (4 obras) · Sociology (4 obras)