Parallel Corpus Research and Target Language Representativeness
The Contrastive, Typological, and Translation Mining Traditions
Bibliographic Data
| ID | 5900059 |
|---|---|
| Authors | Bert Le Bruyn (0000-0001-9090-7383, Utrecht University, corresponding author), Martín Fuchs (0000-0002-6862-8422, Utrecht University), Marjolijn Van Der Klis (0000-0003-0008-9028, Utrecht University), Jianan Liu (0009-0002-7179-0129, Utrecht University), Chou Mo (Utrecht University), Jos Tellings (0000-0002-4259-8202, Utrecht University), Henriette De Swart (Utrecht University) |
| Year | 2022 |
| Volume | 7 |
| Issue | 3 |
| Pages | 176 |
| Publication date | 2022-07-07 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Languages (JOURNAL) |
| Journal identifiers | ISSN: 2226-471X • E-ISSN: 2226-471X |
| Publisher | MDPI AG (PUBLISHER • IT) |
| DOI | 10.3390/languages7030176 |
| OpenAlex | W4284894748 |
| Language | EN |
| Citations received | 7 |
| References cited | 26 |
This paper surveys the strategies that the Contrastive, Typological, and Translation Mining parallel corpus traditions rely on to deal with the issue of target language representativeness of translations. On the basis of a comparison of the corpus architectures and research designs of the three traditions, we argue that they have each developed their own representativeness strategies: (i) monolingual control corpora (Contrastive tradition), (ii) limits on the scope of research questions (Typological tradition), and (iii) parallel control corpora (Translation Mining tradition). We introduce normalized pointwise mutual information (NPMI) as a bi-directional measure of cross-linguistic association, allowing for an easy comparison of the outcomes of different traditions and the impact of the monolingual and parallel control corpus representativeness strategies. We further argue that corpus size has a major impact on the reliability of the monolingual control corpus strategy and that a sequential parallel control corpus strategy is preferable for smaller corpora
Control (management · Corpus linguistics · Linguistics · Machine translation · Mutual information · Natural language processing · Pointwise mutual information · Representativeness heuristic · Scope (computer science · Text corpus · Translation (biology · Computer Science · linguistics and terminology studies · Natural Language Processing Techniques · Psychology · Translation Studies and Practices · Artificial Intelligence
A parallel corpus-based exploration of deflected agreement in Arabic varieties
Self-talk and syntactic structure
Looking for someone
The Discovery of Aspect
Nobody's Perfect
Differences between Russian and Czech in the Use of Aspect in Narrative Discourse and Factual Contexts
Perfective Marking in the Breton Tense-Aspect System
Seeing through Multilingual Corpora
Aspect in Mandarin Chinese
Cross-Linguistic Corpora for the Study of Translations
Corpora and Cross-Linguistic Research
Similarity Semantics and Building Probabilistic Semantic Maps from Parallel Texts
The Perfect in dialogue
Semantic maps of causation
Not…Until across European Languages
Translation Mining
Perfect-Perfective Variation across Spanish Dialects
The Discovery of Aspect
Tense and Aspect in a Spanish Literary Work and Its Translations
Differences between Russian and Czech in the Use of Aspect in Narrative Discourse and Factual Contexts
Perfective Marking in the Breton Tense-Aspect System
Dutch Parallel Corpus
Lexical typology through similarity semantics
| Unique citing works | 7 |
|---|---|
| Citations per year | 1,75 |
| Citation span | 2022 - 2025 (4) |
| Citation velocity | recent |
| Highly cited | No |
| Citation types | Neutral: 7 |