Tony Mcenery
Biographic Data
| ID | 175896 |
|---|---|
| NAME | Tony Mcenery |
| GIVEN NAMES | Tony |
| FAMILY NAME | Mcenery |
| SIGNATURE | MCENERY T |
| AFFILIATIONS | Lancaster University |
| ORCID | 0000-0002-8425-6403 |
| VERIFIED | Yes |
| TOTAL WORKS | 38 |
| TOTAL CITATIONS | 668 |
| AUTHOR COUNT | 38 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2003 |
| LATEST PUBLICATION YEAR | 2026 |
| H-INDEX | 7 |
Gendered representations in swearing: Bitch and bastard in the spoken BNC2014
This study examines how bitch and bastard construct gendered identities in contemporary British English conversation. Using corpus-assisted critical discourse analysis of the Spoken BNC2014, it examines collocational patterns and “ be + bitch / bastard ” constructions to trace how gendered meanings are enacted across speaker and target sexes. bastard predominantly targets men, representing masculinity through moral evaluation, functioning as a di…
Grammar and Corpora
The rapid development of corpus linguistics since the early 1990s has revolutionized virtually all areas of linguistic research. The corpus approach has helped redefine what a grammar is, leading grammars away from abstract, idealized, representations of language toward an approach rooted in attested language use in the form of corpora. The confrontation with corpus data has changed what is viewed as grammatical and what a grammar should account …
Corpus linguistics for language teaching and learning: A research agenda
This agenda identifies future research trajectories for the corpus revolution, proposing five specific research tasks designed to explore and advance the application of corpus linguistics in language education. These tasks focus on: (1) contrastive data-driven learning, (2) the development of corpus research for informing national language curricula, (3) the use of artificial intelligence for corpus informed language teaching and learning, (4) th…
Family, politics and media: Gladstone during the Midlothian campaign, 1879–1880
In this paper, we utilise the Nineteenth Century Newspaper Corpus to examine reporting surrounding William Gladstone’s Midlothian campaign, a key point in the democratization of British politics where a politician not only communicated with ordinary people through hustings but indirectly to a wider electorate via media reporting of those hustings. With the use of social actor analysis ( van Leeuwen 2008 ), approached through collocation, we find …
Swearing, discourse and function in conversational British English
In this paper we look at the role that macrostructures in discourse have to play in the study of swearing. While studied in isolation, such macrostructures have not yet been studied comprehensively and the range of macrostructures studied has been small. By contrast, work on microstructures is much better developed. In response to this, using spoken corpus data from the BNC2014, we take two approaches to studying discourse in this paper. In the f…
Keywords through time: Tracking changes in press discourses of Islam
This paper applies a new approach to the identification of discourses, based on Multiple Correspondence Analysis (MCA), to the study of discourse variation over time. The MCA approach to keywords deals with a major issue with the use of keywords to identify discourses: the allocation of individual keywords to multiple discourses. Yet, as this paper demonstrates, the approach also allows us to observe variation in the prevalence of discourses over…
Corpus studies of language through time: Introduction to the special issue
The study of language through time has long been an area where the corpus approach to the analysis of language has been an important method. The possibility of using other methods, such as elicitation, introspection or psycholinguistic experiments to investigate language in the past is, effectively, eliminated by a simple fact -there is no direct access possible to speakers of language beyond those generations that are living. Our only access to …
Slavery and Britain in the 19th century
This study uses a corpus of just under two billion words from one historic British newspaper, the Liverpool Mercury , to explore shifting attitudes to slavery in Britain in the nineteenth century in the context of a port city that benefitted from the trade. In doing so, we explore three methodological issues – how to explore concepts in large corpora, how to do this over time and how to deal with poor quality data. Our approach to the study of co…
The Language of Violent Jihad
How do violent jihadists use language to try to persuade people to carry out violent acts? This book analyses over two million words of texts produced by violent jihadists to identify and examine the linguistic strategies employed. Taking a mixed methods approach, the authors combine quantitative methods from corpus linguistics, which allows the identification of frequent words and phrases, alongside close reading of texts via discourse analysis.…
Identifying and describing functional discourse units in the BNC Spoken 2014
On the surface, it appears that conversational language is produced in a stream of spoken utterances. In reality conversation is composed of contiguous units that are characterized by coherent communicative purposes. A large number of important research questions about the nature of conversational discourse could be addressed if researchers could investigate linguistic variation across functional discourse units. To date, however, no corpus of co…
Remembering Geoffrey Leech
This special issue of Text and Talk is dedicated to the memory and many contributions of Geoffrey Leech.Geoff was a great colleague and a hugely influential academic who, as this special issue will show, made significant contributions to an astoundingly wide range of areas of linguistics.We will begin this guest editorial to the special issue by considering Geoff as a person and as a researcher.We will do this by outlining, based on our own memor…
The Written British National Corpus 2014 – design and comparability
The British National Corpus 2014 is a major project led by Lancaster University to create a 100-million-word corpus of present day British English. This corpus has been constructed as a comparable counterpart of the original British National Corpus (referred to as the BNC1994 in this article), which was compiled in the early 1990s. This article starts with the justification of the project answering the question of ‘Why do we need a new BNC?’. We …
Narrative evaluation in patient feedback: A study of online comments about UK healthcare services
This study examines how patients use narratives to evaluate their experiences of healthcare services online. The analysis draws on corpus linguistic techniques, specifically annotation, applying Labov and Waletzky’s (1967) framework to a sample of online comments about the NHS in England. Narratives are pervasive in this context, being present more than absent in the patients’ comments, but are particularly prominent in comments which evaluate ca…
Correlation, collocation and cohesion: A corpus-based critical analysis of violent jihadist discourse
This article explores the language of violent jihad, focussing upon lexis encoding concepts from Islam. Through the use of correlation statistics, this article demonstrates that the words encoding such concepts distribute in dependent relationships across different types of texts. The correlation between the words cannot be simply explained in terms of collocation; rather, the correlation is evidence of other forms of cohesion at work in the text…
Usage Fluctuation Analysis: A new way of analysing shifts in historical discourse
This article introduces a methodology for the diachronic analysis of large historical corpora, Usage Fluctuation Analysis (UFA). UFA looks at the fluctuation of the usage of a word as observed through collocation. It presupposes neither a commitment to a specific semantic theory, nor that the results will focus solely on semantics. We focus, rather, upon a word’s usage. UFA considers large amounts of evidence about usage, through time, as made av…
The utility of topic modelling for discourse studies: A critical evaluation
This article explores and critically evaluates the potential contribution to discourse studies of topic modelling, a group of machine learning methods which have been used with the aim of automatically discovering thematic information in large collections of texts. We critically evaluate the utility of the thematic grouping of texts into ‘topics’ emerging from a large collection of online patient comments about the National Health Service (NHS) i…
Epistemic Stance in Spoken L2 English: The Effect of Task and Speaker Style
The article discusses epistemic stance in spoken L2 production. Using a subset of the Trinity Lancaster Corpus of spoken L2 production, we analysed the speech of 132 advanced L2 speakers from different L1 and cultural backgrounds taking part in four speaking tasks: one largely monologic presentation task and three interactive tasks. The study focused on three types of epistemic forms: adverbial, adjectival, and verbal expressions. The results sho…
The public representation of homosexual men in seventeenth-century England – a corpus based view
In this article we explore public discourse around one marginalized group in early-modern English society, men who engaged in sexual relations with other males. To do this we use a large corpus of seventeenth century texts, the Early English Books Online corpus. Our exploration leads us to consider a number of methodological issues, notably low frequency data and the classical framing of some words. We consider the historical context which brings…
The Spoken BNC2014: Designing and building a spoken corpus of everyday conversations
This paper introduces the Spoken British National Corpus 2014, an 11.5-million-word corpus of orthographically transcribed conversations among L1 speakers of British English from across the UK, recorded in the years 2012–2016. After showing that a survey of the recent history of corpora of spoken British English justifies the compilation of this new corpus, we describe the main stages of the Spoken BNC2014’s creation: design, data and metadata co…
Corpus Linguistics and 17th-Century Prostitution: Computational Linguistics and History
Corpus linguistics has much to offer history, being as both disciplines engage so heavily in analysis of large amounts of textual material. This book demonstrates the opportunities for exploring corpus linguistics as a method in historiography and the humanities and social sciences more generally. Focussing on the topic of prostitution in 17th-century England, it shows how corpus methods can assist in social research, and can be used to deepen ou…
Exploring Learner Language Through Corpora: Comparing and Interpreting Corpus Frequency Information
This article contributes to the debate about the appropriate use of corpus data in language learning research. It focuses on frequencies of linguistic features in language use and their comparison across corpora. The majority of corpus-based second language acquisition studies employ a comparative design in which either one or more second language (L2) corpora are compared to a first language (L1) production corpus or two or more L2 corpora are c…
Collocations in Corpus-Based Language Learning Research: Identifying, Comparing, and Interpreting the Evidence
This article focuses on the use of collocations in language learning research (LLR). Collocations, as units of formulaic language, are becoming prominent in our understanding of language learning and use; however, while the number of corpus-based LLR studies of collocations is growing, there is still a need for a deeper understanding of factors that play a role in establishing that two words in a corpus can be considered to be collocates. In this…
Language Learning Research at the Intersection of Experimental, Computational, and Corpus-Based Approaches
Language acquisition occupies a central place in the study of human cognition, and research on how we learn language can be found across many disciplines, from developmental psychology and linguistics to education, philosophy, and neuroscience. It is a very challenging topic to investigate given that the learning target in first and second language acquisition is highly complex, and part of the challenge consists in identifying how different doma…
Collocations in context: A new perspective on collocation networks
The idea that text in a particular field of discourse is organized into lexical patterns, which can be visualized as networks of words that collocate with each other, was originally proposed by Phillips (1983). This idea has important theoretical implications for our understanding of the relationship between the lexis and the text and (ultimately) between the text and the discourse community/the mind of the speaker. Although the approaches to dat…
Grammar and Corpora
The rapid development of corpus linguistics since the early 1990s has revolutionized virtually all areas of linguistic research. Research in grammar has probably been influenced most profoundly by the corpus‐based approach. It has helped to redefine what a grammar is. Indeed, corpora have had such a strong impact on recently published reference grammars of English that “even people who have never heard of a corpus are using the product of corpus‐…
A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees and asylum seekers in the UK press
This article discusses the extent to which methods normally associated with corpus linguistics can be effectively used by critical discourse analysts. Our research is based on the analysis of a 140-million-word corpus of British news articles about refugees, asylum seekers, immigrants and migrants (collectively RASIM). We discuss how processes such as collocation and concordance analysis were able to identify common categories of representation o…
A corpus-based approach to discourses of refugees and asylum seekers in UN and newspaper texts
A corpus-based analysis of discourses of refugees and asylum seekers was carried out on data taken from a range of British newspapers and texts from the Office of the United Nations High Commissioner for Refugees website, both published in 2003. Concordances of the terms refugee(s) and asylum seeker(s) were examined and grouped along patterns which revealed linguistic traces of discourses. Discourses which framed refugees as packages, invaders, p…
Collocations in Corpus-Based Language Learning Research: Identifying, Comparing, and Interpreting the Evidence
This article focuses on the use of collocations in language learning research (LLR). Collocations, as units of formulaic language, are becoming prominent in our understanding of language learning and use; however, while the number of corpus-based LLR studies of collocations is growing, there is still a need for a deeper understanding of factors that play a role in establishing that two words in a corpus can be considered to be collocates. In this…
Collocations in context: A new perspective on collocation networks
The idea that text in a particular field of discourse is organized into lexical patterns, which can be visualized as networks of words that collocate with each other, was originally proposed by Phillips (1983). This idea has important theoretical implications for our understanding of the relationship between the lexis and the text and (ultimately) between the text and the discourse community/the mind of the speaker. Although the approaches to dat…
The Spoken BNC2014: Designing and building a spoken corpus of everyday conversations
This paper introduces the Spoken British National Corpus 2014, an 11.5-million-word corpus of orthographically transcribed conversations among L1 speakers of British English from across the UK, recorded in the years 2012–2016. After showing that a survey of the recent history of corpora of spoken British English justifies the compilation of this new corpus, we describe the main stages of the Spoken BNC2014’s creation: design, data and metadata co…
The utility of topic modelling for discourse studies: A critical evaluation
This article explores and critically evaluates the potential contribution to discourse studies of topic modelling, a group of machine learning methods which have been used with the aim of automatically discovering thematic information in large collections of texts. We critically evaluate the utility of the thematic grouping of texts into ‘topics’ emerging from a large collection of online patient comments about the National Health Service (NHS) i…
Corpus Linguistics and 17th-Century Prostitution: Computational Linguistics and History
Corpus linguistics has much to offer history, being as both disciplines engage so heavily in analysis of large amounts of textual material. This book demonstrates the opportunities for exploring corpus linguistics as a method in historiography and the humanities and social sciences more generally. Focussing on the topic of prostitution in 17th-century England, it shows how corpus methods can assist in social research, and can be used to deepen ou…
Correlation, collocation and cohesion: A corpus-based critical analysis of violent jihadist discourse
This article explores the language of violent jihad, focussing upon lexis encoding concepts from Islam. Through the use of correlation statistics, this article demonstrates that the words encoding such concepts distribute in dependent relationships across different types of texts. The correlation between the words cannot be simply explained in terms of collocation; rather, the correlation is evidence of other forms of cohesion at work in the text…
The peaks and troughs of corpus-based contextual analysis
This paper focuses upon two issues. Firstly, the question of identifying diachronic trends, and more importantly significant outliers, in corpora which permit an investigation of a feature at many sampling points over time. Secondly, we consider how best to combine more qualitatively oriented approaches to corpus data with the type of trends that can be observed in a corpus using quantitative techniques. The work uses a recently completed ESRC-fu…
Exploring Learner Language Through Corpora: Comparing and Interpreting Corpus Frequency Information
This article contributes to the debate about the appropriate use of corpus data in language learning research. It focuses on frequencies of linguistic features in language use and their comparison across corpora. The majority of corpus-based second language acquisition studies employ a comparative design in which either one or more second language (L2) corpora are compared to a first language (L1) production corpus or two or more L2 corpora are c…
Language Learning Research at the Intersection of Experimental, Computational, and Corpus-Based Approaches
Language acquisition occupies a central place in the study of human cognition, and research on how we learn language can be found across many disciplines, from developmental psychology and linguistics to education, philosophy, and neuroscience. It is a very challenging topic to investigate given that the learning target in first and second language acquisition is highly complex, and part of the challenge consists in identifying how different doma…
Help or Help to: What Do Corpora Have to Say
In this paper, we will examine a range of factors that may potentially influence a language user's choice of a full or bare infinitive following HELP. The factors include language variety, language change, spoken/written distinction, semantic distinction, and syntactic conditions, namely, an intervening noun phrase or adverbial, the number of intervening words, to preceding HELP, the passive construction, inflections of HELP, and it as the subjec…
Corpus linguistics for language teaching and learning: A research agenda
This agenda identifies future research trajectories for the corpus revolution, proposing five specific research tasks designed to explore and advance the application of corpus linguistics in language education. These tasks focus on: (1) contrastive data-driven learning, (2) the development of corpus research for informing national language curricula, (3) the use of artificial intelligence for corpus informed language teaching and learning, (4) th…
Swearing, discourse and function in conversational British English
In this paper we look at the role that macrostructures in discourse have to play in the study of swearing. While studied in isolation, such macrostructures have not yet been studied comprehensively and the range of macrostructures studied has been small. By contrast, work on microstructures is much better developed. In response to this, using spoken corpus data from the BNC2014, we take two approaches to studying discourse in this paper. In the f…
Narrative evaluation in patient feedback: A study of online comments about UK healthcare services
This study examines how patients use narratives to evaluate their experiences of healthcare services online. The analysis draws on corpus linguistic techniques, specifically annotation, applying Labov and Waletzky’s (1967) framework to a sample of online comments about the NHS in England. Narratives are pervasive in this context, being present more than absent in the patients’ comments, but are particularly prominent in comments which evaluate ca…
Usage Fluctuation Analysis: A new way of analysing shifts in historical discourse
This article introduces a methodology for the diachronic analysis of large historical corpora, Usage Fluctuation Analysis (UFA). UFA looks at the fluctuation of the usage of a word as observed through collocation. It presupposes neither a commitment to a specific semantic theory, nor that the results will focus solely on semantics. We focus, rather, upon a word’s usage. UFA considers large amounts of evidence about usage, through time, as made av…
On two traditions in corpus linguistics, and what they have in common
Preview this article: On two traditions in corpus linguistics, and what they have in common, Page 1 of 1 /docserver/preview/fulltext/ijcl.15.3.09har-1.gif
The moral panic about bad language in England, 1691–1745
In this paper I use a corpus of the writings of the Society for the Reformation of Manners to look at the discursive construction of attitudes to bad language in English. Using this corpus of texts as an example of a moral panic about language I use keywords to explore moral panic rhetoric, the formation of spirals of signification and the impact of both on attitudes to bad language in English in the late seventeenth and early eighteenth centurie…
A corpus-based two-level model of situation aspect
In this paper we will extend Smith's (1997) two-component aspect theory to develop a two-level model of situation aspect in which situation aspect is modelled as verb classes at the lexical level and as situation types at the sentential level. Situation types are the composite result of the rule-based interaction between verb classes and complements, arguments, peripheral adjuncts and viewpoint aspect at the nucleus, core and clause levels. With …
Corpus Linguistics Texts
at a slightly different audience, but in combination they deliver all the essential information one would need to fully understand and utilize the history, current state, and broad applications of corpus linguistics. Kennedy's An Introduction to Corpus Linguistics provides a thorough history of the development of corpus linguistics, beginning with investigations that used corpus methods before computers were readily available. The book is divided…
Aspect in Mandarin Chinese: A corpus-based study
Chinese, as an aspect language, has played an important role in the development of aspect theory. This book is a systematic and structured exploration of the linguistic devices that Mandarin Chinese employs to express aspectual meanings. The work presented here is the first corpus-based account of aspect in Chinese, encompassing both situation aspect and viewpoint aspect. In using corpus data, the book seeks to achieve a marriage between theory-d…
A corpus-based two-level model of situation aspect
In this paper we will extend Smith's (1997) two-component aspect theory to develop a two-level model of situation aspect in which situation aspect is modelled as verb classes at the lexical level and as situation types at the sentential level. Situation types are the composite result of the rule-based interaction between verb classes and complements, arguments, peripheral adjuncts and viewpoint aspect at the nucleus, core and clause levels. With …
Help or Help to: What Do Corpora Have to Say
In this paper, we will examine a range of factors that may potentially influence a language user's choice of a full or bare infinitive following HELP. The factors include language variety, language change, spoken/written distinction, semantic distinction, and syntactic conditions, namely, an intervening noun phrase or adverbial, the number of intervening words, to preceding HELP, the passive construction, inflections of HELP, and it as the subjec…
A corpus-based approach to discourses of refugees and asylum seekers in UN and newspaper texts
A corpus-based analysis of discourses of refugees and asylum seekers was carried out on data taken from a range of British newspapers and texts from the Office of the United Nations High Commissioner for Refugees website, both published in 2003. Concordances of the terms refugee(s) and asylum seeker(s) were examined and grouped along patterns which revealed linguistic traces of discourses. Discourses which framed refugees as packages, invaders, p…
Collocation, Semantic Prosody, and Near Synonymy: A Cross-Linguistic Perspective
This paper explores the collocational behaviour and semantic prosody of near synonyms from a cross-linguistic perspective. The importance of these concepts to language learning is well recognized. Yet while collocation and semantic prosody have recently attracted much interest from researchers studying the English language, there has been little work done on collocation and semantic prosody on languages other than English. Still less work has bee…
The moral panic about bad language in England, 1691–1745
In this paper I use a corpus of the writings of the Society for the Reformation of Manners to look at the discursive construction of attitudes to bad language in English. Using this corpus of texts as an example of a moral panic about language I use keywords to explore moral panic rhetoric, the formation of spirals of signification and the impact of both on attitudes to bad language in English in the late seventeenth and early eighteenth centurie…
A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees and asylum seekers in the UK press
This article discusses the extent to which methods normally associated with corpus linguistics can be effectively used by critical discourse analysts. Our research is based on the analysis of a 140-million-word corpus of British news articles about refugees, asylum seekers, immigrants and migrants (collectively RASIM). We discuss how processes such as collocation and concordance analysis were able to identify common categories of representation o…
On two traditions in corpus linguistics, and what they have in common
Preview this article: On two traditions in corpus linguistics, and what they have in common, Page 1 of 1 /docserver/preview/fulltext/ijcl.15.3.09har-1.gif
Corpus Linguistics: Method, Theory and Practice
Corpus linguistics is the study of language data on a large scale - the computer-aided analysis of very extensive collections of transcribed utterances or written texts. This textbook outlines the basic methods of corpus linguistics, explains how the discipline of corpus linguistics developed and surveys the major approaches to the use of corpus data. It uses a broad range of examples to show how corpus data has led to methodological and theoreti…
The peaks and troughs of corpus-based contextual analysis
This paper focuses upon two issues. Firstly, the question of identifying diachronic trends, and more importantly significant outliers, in corpora which permit an investigation of a feature at many sampling points over time. Secondly, we consider how best to combine more qualitatively oriented approaches to corpus data with the type of trends that can be observed in a corpus using quantitative techniques. The work uses a recently completed ESRC-fu…
Grammar and Corpora
The rapid development of corpus linguistics since the early 1990s has revolutionized virtually all areas of linguistic research. Research in grammar has probably been influenced most profoundly by the corpus‐based approach. It has helped to redefine what a grammar is. Indeed, corpora have had such a strong impact on recently published reference grammars of English that “even people who have never heard of a corpus are using the product of corpus‐…
Sketching Muslims: A Corpus Driven Analysis of Representations Around the Word 'Muslim' in the British Press 1998-2009
This article uses methods from corpus linguistics and critical discourse analysis to examine patterns of representation around the word Muslim in a 143 million word corpus of British newspaper articles published between 1998 and 2009. Using the analysis tool Sketch Engine, an analysis of noun collocates of Muslim found that the following categories (in order of frequency) were referenced: ethnic/national identity, characterizing/differentiating a…
Discourse Analysis and Media Attitudes: The Representation of Islam in the British Press
Is the British press prejudiced against Muslims? In what ways can prejudice be explicit or subtle? This book uses a detailed analysis of over 140 million words of newspaper articles on Muslims and Islam, combining corpus linguistics and discourse analysis methods to produce an objective picture of media attitudes. The authors analyse representations around frequently cited topics such as Muslim women who wear the veil and 'hate preachers'. The an…
Collocations in context: A new perspective on collocation networks
The idea that text in a particular field of discourse is organized into lexical patterns, which can be visualized as networks of words that collocate with each other, was originally proposed by Phillips (1983). This idea has important theoretical implications for our understanding of the relationship between the lexis and the text and (ultimately) between the text and the discourse community/the mind of the speaker. Although the approaches to dat…
Epistemic Stance in Spoken L2 English: The Effect of Task and Speaker Style
The article discusses epistemic stance in spoken L2 production. Using a subset of the Trinity Lancaster Corpus of spoken L2 production, we analysed the speech of 132 advanced L2 speakers from different L1 and cultural backgrounds taking part in four speaking tasks: one largely monologic presentation task and three interactive tasks. The study focused on three types of epistemic forms: adverbial, adjectival, and verbal expressions. The results sho…
The public representation of homosexual men in seventeenth-century England – a corpus based view
In this article we explore public discourse around one marginalized group in early-modern English society, men who engaged in sexual relations with other males. To do this we use a large corpus of seventeenth century texts, the Early English Books Online corpus. Our exploration leads us to consider a number of methodological issues, notably low frequency data and the classical framing of some words. We consider the historical context which brings…
The Spoken BNC2014: Designing and building a spoken corpus of everyday conversations
This paper introduces the Spoken British National Corpus 2014, an 11.5-million-word corpus of orthographically transcribed conversations among L1 speakers of British English from across the UK, recorded in the years 2012–2016. After showing that a survey of the recent history of corpora of spoken British English justifies the compilation of this new corpus, we describe the main stages of the Spoken BNC2014’s creation: design, data and metadata co…
Corpus Linguistics and 17th-Century Prostitution: Computational Linguistics and History
Corpus linguistics has much to offer history, being as both disciplines engage so heavily in analysis of large amounts of textual material. This book demonstrates the opportunities for exploring corpus linguistics as a method in historiography and the humanities and social sciences more generally. Focussing on the topic of prostitution in 17th-century England, it shows how corpus methods can assist in social research, and can be used to deepen ou…
Exploring Learner Language Through Corpora: Comparing and Interpreting Corpus Frequency Information
This article contributes to the debate about the appropriate use of corpus data in language learning research. It focuses on frequencies of linguistic features in language use and their comparison across corpora. The majority of corpus-based second language acquisition studies employ a comparative design in which either one or more second language (L2) corpora are compared to a first language (L1) production corpus or two or more L2 corpora are c…
Collocations in Corpus-Based Language Learning Research: Identifying, Comparing, and Interpreting the Evidence
This article focuses on the use of collocations in language learning research (LLR). Collocations, as units of formulaic language, are becoming prominent in our understanding of language learning and use; however, while the number of corpus-based LLR studies of collocations is growing, there is still a need for a deeper understanding of factors that play a role in establishing that two words in a corpus can be considered to be collocates. In this…
Language Learning Research at the Intersection of Experimental, Computational, and Corpus-Based Approaches
Language acquisition occupies a central place in the study of human cognition, and research on how we learn language can be found across many disciplines, from developmental psychology and linguistics to education, philosophy, and neuroscience. It is a very challenging topic to investigate given that the learning target in first and second language acquisition is highly complex, and part of the challenge consists in identifying how different doma…
Usage Fluctuation Analysis: A new way of analysing shifts in historical discourse
This article introduces a methodology for the diachronic analysis of large historical corpora, Usage Fluctuation Analysis (UFA). UFA looks at the fluctuation of the usage of a word as observed through collocation. It presupposes neither a commitment to a specific semantic theory, nor that the results will focus solely on semantics. We focus, rather, upon a word’s usage. UFA considers large amounts of evidence about usage, through time, as made av…
The utility of topic modelling for discourse studies: A critical evaluation
This article explores and critically evaluates the potential contribution to discourse studies of topic modelling, a group of machine learning methods which have been used with the aim of automatically discovering thematic information in large collections of texts. We critically evaluate the utility of the thematic grouping of texts into ‘topics’ emerging from a large collection of online patient comments about the National Health Service (NHS) i…
Correlation, collocation and cohesion: A corpus-based critical analysis of violent jihadist discourse
This article explores the language of violent jihad, focussing upon lexis encoding concepts from Islam. Through the use of correlation statistics, this article demonstrates that the words encoding such concepts distribute in dependent relationships across different types of texts. The correlation between the words cannot be simply explained in terms of collocation; rather, the correlation is evidence of other forms of cohesion at work in the text…
Linguistics (34 works) · Computer Science (26 works) · Philosophy (20 works) · Corpus linguistics (18 works) · Sociology (18 works) · Natural language processing (15 works) · Artificial Intelligence (14 works) · Discourse Analysis in Language Studies (13 works) · Philosophy (12 works) · Psychology (11 works)