Xiaoming Xi
Datos Biográficos
| ID | 3843846 |
|---|---|
| NOMBRE | Xiaoming Xi |
| NOMBRES | Xiaoming |
| APELLIDO | Xi |
| FIRMA | XI X |
| AFILIACIONES | Educational Testing Service |
| ORCID | 0000-0002-0415-3608 |
| VERIFICADO | Sí |
| TOTAL DE OBRAS | 10 |
| TOTAL DE CITAS | 1 |
| TOTAL COMO AUTOR | 10 |
| TOTAL COMO EDITOR | 0 |
| PRIMER AÑO DE PUBLICACIÓN | 2005 |
| AÑO MÁS RECIENTE DE PUBLICACIÓN | 2022 |
| ÍNDICE H | 1 |
Validation of Language Assessments
Validation is an activity that establishes the extent to which a language assessment is meaningful and useful. This activity involves the collection of different types of evidence in order to make a holistic evaluation of an assessment's fitness for a purpose. Several frameworks have been introduced to guide the collection and evaluation of this evidence for language assessments, and the argument‐based approach has become increasingly prominent. …
What does corpus linguistics have to offer to language assessment
In recent years, continuing advances in technology have increased the capacity to automate the extraction of a range of linguistic features of texts and thus have provided the impetus for the substantial growth of corpus linguistics. While corpus linguistic tools and methods have been used extensively in second language learning research, they have also been used increasingly in the design and validation of language assessments (Callies & Götz, 2…
A study on the impact of fatigue on human raters when scoring speaking responses
The scoring of constructed responses may introduce construct-irrelevant factors to a test score and affect its validity and fairness. Fatigue is one of the factors that could negatively affect human performance in general, yet little is known about its effects on a human rater’s scoring quality on constructed responses. In this study, we compared the scoring quality of 72 raters under four shift conditions differing on the shift length (total sco…
A comparison of two scoring methods for an automated speech scoring system
This paper compares two alternative scoring methods – multiple regression and classification trees – for an automated speech scoring system used in a practice environment. The two methods were evaluated on two criteria: construct representation and empirical performance in predicting human scores. The empirical performance of the two scoring models is reported in Zechner, Higgins, Xi, & Williamson (2009), which discusses the development of the en…
Using Raters From India to Score a Large-Scale Speaking Test
We investigated the scoring of the Speaking section of the Test of English as a Foreign LanguageTM Internet-based (TOEFL iBT(R)) test by speakers of English and one or more Indian languages. We explored the extent to which raters from India, after being trained and certified, were able to score the TOEFL examinees with mixed first languages accurately and consistently. The effectiveness of a special training package designed for scoring Indian ex…
Aspects of performance on line graph description tasks
Motivated by cognitive theories of graph comprehension, this study systematically manipulated characteristics of a line graph description task in a speaking test in ways to mitigate the influence of graph familiarity, a potential source of construct-irrelevant variance. It extends Xi (2005), which found that the differences in holistic scores on graph tasks with varying characteristics, although significant, were small. Using an analytic scoring …
Automated scoring and feedback systems
How do we go about investigating test fairness
Previous test fairness frameworks have greatly expanded the scope of fairness, but do not provide a means to fully integrate fairness investigations and set priorities. This article proposes an approach to guide practitioners on fairness research and practices. This approach treats fairness as an aspect of validity and conceptualizes it as comparable validity for all relevant groups. Anything that weakens fairness compromises the validity of a te…
Evaluating analytic scoring for the Toefl® Academic Speaking Test (Tast) for operational use
This study explores the utility of analytic scoring for TAST in providing useful and reliable diagnostic information for operational use in three aspects of candidates' performance: delivery, language use and topic development. One hundred and forty examinees' responses to six TAST tasks were scored analytically on these three aspects of speech. G studies were used to investigate the dependability of the analytic scores, the distinctness of the a…
Do visual chunks and planning impact performance on the graph description task in the SPEAK exam
This study examines how task characteristics (the number of visual chunks and the amount of planning time) and test-taker characteristics (graph familiarity) influence the perceptual and cognitive processes involved in graph comprehension, the strategies used in describing graphs, and the scores obtained on the graph description task in a semi-direct oral test. Specifically, it investigates whether providing planning time and reducing the number …
Using Raters From India to Score a Large-Scale Speaking Test
We investigated the scoring of the Speaking section of the Test of English as a Foreign LanguageTM Internet-based (TOEFL iBT(R)) test by speakers of English and one or more Indian languages. We explored the extent to which raters from India, after being trained and certified, were able to score the TOEFL examinees with mixed first languages accurately and consistently. The effectiveness of a special training package designed for scoring Indian ex…
Do visual chunks and planning impact performance on the graph description task in the SPEAK exam
This study examines how task characteristics (the number of visual chunks and the amount of planning time) and test-taker characteristics (graph familiarity) influence the perceptual and cognitive processes involved in graph comprehension, the strategies used in describing graphs, and the scores obtained on the graph description task in a semi-direct oral test. Specifically, it investigates whether providing planning time and reducing the number …
Evaluating analytic scoring for the Toefl® Academic Speaking Test (Tast) for operational use
This study explores the utility of analytic scoring for TAST in providing useful and reliable diagnostic information for operational use in three aspects of candidates' performance: delivery, language use and topic development. One hundred and forty examinees' responses to six TAST tasks were scored analytically on these three aspects of speech. G studies were used to investigate the dependability of the analytic scores, the distinctness of the a…
Aspects of performance on line graph description tasks
Motivated by cognitive theories of graph comprehension, this study systematically manipulated characteristics of a line graph description task in a speaking test in ways to mitigate the influence of graph familiarity, a potential source of construct-irrelevant variance. It extends Xi (2005), which found that the differences in holistic scores on graph tasks with varying characteristics, although significant, were small. Using an analytic scoring …
Automated scoring and feedback systems
How do we go about investigating test fairness
Previous test fairness frameworks have greatly expanded the scope of fairness, but do not provide a means to fully integrate fairness investigations and set priorities. This article proposes an approach to guide practitioners on fairness research and practices. This approach treats fairness as an aspect of validity and conceptualizes it as comparable validity for all relevant groups. Anything that weakens fairness compromises the validity of a te…
Using Raters From India to Score a Large-Scale Speaking Test
We investigated the scoring of the Speaking section of the Test of English as a Foreign LanguageTM Internet-based (TOEFL iBT(R)) test by speakers of English and one or more Indian languages. We explored the extent to which raters from India, after being trained and certified, were able to score the TOEFL examinees with mixed first languages accurately and consistently. The effectiveness of a special training package designed for scoring Indian ex…
A comparison of two scoring methods for an automated speech scoring system
This paper compares two alternative scoring methods – multiple regression and classification trees – for an automated speech scoring system used in a practice environment. The two methods were evaluated on two criteria: construct representation and empirical performance in predicting human scores. The empirical performance of the two scoring models is reported in Zechner, Higgins, Xi, & Williamson (2009), which discusses the development of the en…
A study on the impact of fatigue on human raters when scoring speaking responses
The scoring of constructed responses may introduce construct-irrelevant factors to a test score and affect its validity and fairness. Fatigue is one of the factors that could negatively affect human performance in general, yet little is known about its effects on a human rater’s scoring quality on constructed responses. In this study, we compared the scoring quality of 72 raters under four shift conditions differing on the shift length (total sco…
What does corpus linguistics have to offer to language assessment
In recent years, continuing advances in technology have increased the capacity to automate the extraction of a range of linguistic features of texts and thus have provided the impetus for the substantial growth of corpus linguistics. While corpus linguistic tools and methods have been used extensively in second language learning research, they have also been used increasingly in the design and validation of language assessments (Callies & Götz, 2…
Validation of Language Assessments
Validation is an activity that establishes the extent to which a language assessment is meaningful and useful. This activity involves the collection of different types of evidence in order to make a holistic evaluation of an assessment's fitness for a purpose. Several frameworks have been introduced to guide the collection and evaluation of this evidence for language assessments, and the argument‐based approach has become increasingly prominent. …
Psychology (9 obras) · Computer Science (8 obras) · Cognitive psychology (4 obras) · Language assessment (3 obras) · Mathematics (3 obras) · Natural language processing (3 obras) · Psychometrics (3 obras) · Sociology (3 obras) · Artificial Intelligence (2 obras) · Cognition (2 obras)