Andres Karjus
Datos Biográficos
| ID | 1453383 |
|---|---|
| NOMBRE | Andres Karjus |
| NOMBRES | Andres |
| APELLIDO | Karjus |
| FIRMA | KARJUS A |
| AFILIACIONES | Estonian Business School |
| ORCID | 0000-0002-2445-5072 |
| VERIFICADO | Sí |
| TOTAL DE OBRAS | 9 |
| TOTAL DE CITAS | 2 |
| TOTAL COMO AUTOR | 9 |
| TOTAL COMO EDITOR | 0 |
| PRIMER AÑO DE PUBLICACIÓN | 2017 |
| AÑO MÁS RECIENTE DE PUBLICACIÓN | 2025 |
| ÍNDICE H | 1 |
Machine-assisted quantitizing designs
The increasing capacities of large language models (LLMs) have been shown to present an unprecedented opportunity to scale up data analytics in the humanities and social sciences, by automating complex qualitative tasks otherwise typically carried out by human researchers. While numerous benchmarking studies have assessed the analytic prowess of LLMs, there is less focus on operationalizing this capacity for inference and hypothesis testing. Addr…
Socioeconomic factors of national representation in the global film festival circuit
This study analyzes how economic, demographic, and geographic factors predict the representation of different countries in the global film festival circuit. It relies on the combination of several open-access databases, including festival programming information from the Cinando platform of the Cannes Film Market. The dataset consists of over 20,000 unique films from almost 600 festivals across the world over a decade, a total of more than 30,000…
Evolving linguistic divergence on polarizing social media
Language change is influenced by many factors, but often starts from synchronic variation, where multiple linguistic patterns or forms coexist, or where different speech communities use language in increasingly different ways. Besides regional or economic reasons, communities may form and segregate based on political alignment. The latter, referred to as political polarization, is of growing societal concern across the world. Here we map and quan…
A framework for the analysis of historical newsreels
Audiovisual news is a critical cultural phenomenon that has been influencing audience worldviews for more than a hundred years. To understand historical trends in multimodal audiovisual news, we need to explore them longitudinally using large sets of data. Despite promising developments in film history, computational video analysis, and other relevant fields, current research streams have limitations related to the scope of data used, the systema…
Perceived gendered self-representation on Tinder using machine learning
This paper explores the gendered differences between men and women as perceived through the images on the online dating platform Tinder. While personal images on Instagram, Tumblr, and Facebook have been studied en masse, large-scale studies of the landscape of visual representations on online dating platforms remain rare. We apply a machine learning algorithm to 10,680 profile images collected on Tinder in Estonia to study the perceived gendered…
Challenges in detecting evolutionary forces in language change using diachronic corpora
Newberry et al. (Detecting evolutionary forces in language change, Nature 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having applied a test from population genetics (the Frequency Increment Test) to a number of relevant examples, they suggest stochasticity has a previously under-appreciated role in language evolution. We replicate their re…
Quantifying the dynamics of topical fluctuations in language
The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements in texts, with changes in frequency taken to reflect the popularity or selective fitness of an element. However, corpus frequencies may change for a wide variety of reasons, including purely random sampling ef…
Testing an agent-based model of language choice on sociolinguistic survey data
The paper outlines an agent-based model for language choice in multilingual communities and tests its performance on samples of data drawn from a large-scale sociolinguistic survey carried out in Estonia. While previous research in the field of language competition has focused on diachronic applications, utilizing rather abstract models of uniform speakers, we aim to model synchronic language competition among more realistic, data-based agents. W…
Explaining asymmetries in number marking
This paper claims that crosslinguistic tendencies of number marking asymmetries can be explained with reference to usage frequency: The kinds of nouns which, across languages, tend to show singulative coding (with special marking of the uniplex member of a pair), rather than the more usual plurative coding (with special marking of the multiplex member), are also the kinds of nouns which tend to occur more frequently in multiplex use. We provide c…
Explaining asymmetries in number marking
This paper claims that crosslinguistic tendencies of number marking asymmetries can be explained with reference to usage frequency: The kinds of nouns which, across languages, tend to show singulative coding (with special marking of the uniplex member of a pair), rather than the more usual plurative coding (with special marking of the multiplex member), are also the kinds of nouns which tend to occur more frequently in multiplex use. We provide c…
Explaining asymmetries in number marking
This paper claims that crosslinguistic tendencies of number marking asymmetries can be explained with reference to usage frequency: The kinds of nouns which, across languages, tend to show singulative coding (with special marking of the uniplex member of a pair), rather than the more usual plurative coding (with special marking of the multiplex member), are also the kinds of nouns which tend to occur more frequently in multiplex use. We provide c…
Testing an agent-based model of language choice on sociolinguistic survey data
The paper outlines an agent-based model for language choice in multilingual communities and tests its performance on samples of data drawn from a large-scale sociolinguistic survey carried out in Estonia. While previous research in the field of language competition has focused on diachronic applications, utilizing rather abstract models of uniform speakers, we aim to model synchronic language competition among more realistic, data-based agents. W…
Challenges in detecting evolutionary forces in language change using diachronic corpora
Newberry et al. (Detecting evolutionary forces in language change, Nature 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having applied a test from population genetics (the Frequency Increment Test) to a number of relevant examples, they suggest stochasticity has a previously under-appreciated role in language evolution. We replicate their re…
Quantifying the dynamics of topical fluctuations in language
The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements in texts, with changes in frequency taken to reflect the popularity or selective fitness of an element. However, corpus frequencies may change for a wide variety of reasons, including purely random sampling ef…
Evolving linguistic divergence on polarizing social media
Language change is influenced by many factors, but often starts from synchronic variation, where multiple linguistic patterns or forms coexist, or where different speech communities use language in increasingly different ways. Besides regional or economic reasons, communities may form and segregate based on political alignment. The latter, referred to as political polarization, is of growing societal concern across the world. Here we map and quan…
A framework for the analysis of historical newsreels
Audiovisual news is a critical cultural phenomenon that has been influencing audience worldviews for more than a hundred years. To understand historical trends in multimodal audiovisual news, we need to explore them longitudinally using large sets of data. Despite promising developments in film history, computational video analysis, and other relevant fields, current research streams have limitations related to the scope of data used, the systema…
Perceived gendered self-representation on Tinder using machine learning
This paper explores the gendered differences between men and women as perceived through the images on the online dating platform Tinder. While personal images on Instagram, Tumblr, and Facebook have been studied en masse, large-scale studies of the landscape of visual representations on online dating platforms remain rare. We apply a machine learning algorithm to 10,680 profile images collected on Tinder in Estonia to study the perceived gendered…
Machine-assisted quantitizing designs
The increasing capacities of large language models (LLMs) have been shown to present an unprecedented opportunity to scale up data analytics in the humanities and social sciences, by automating complex qualitative tasks otherwise typically carried out by human researchers. While numerous benchmarking studies have assessed the analytic prowess of LLMs, there is less focus on operationalizing this capacity for inference and hypothesis testing. Addr…
Socioeconomic factors of national representation in the global film festival circuit
This study analyzes how economic, demographic, and geographic factors predict the representation of different countries in the global film festival circuit. It relies on the combination of several open-access databases, including festival programming information from the Cinando platform of the Cannes Film Market. The dataset consists of over 20,000 unique films from almost 600 festivals across the world over a decade, a total of more than 30,000…
Computer Science (5 obras) · Language and cultural evolution (5 obras) · Artificial Intelligence (3 obras) · Linguistics (3 obras) · Psychology (3 obras) · Art (2 obras) · Authorship Attribution and Profiling (2 obras) · Humanities (2 obras) · Linguistic Variation and Morphology (2 obras) · Mathematics (2 obras)