Stephen G Sireci
Biographic Data
| ID | 131675 |
|---|---|
| NAME | Stephen G Sireci |
| GIVEN NAMES | Stephen G |
| FAMILY NAME | Sireci |
| SIGNATURE | SIRECI S G |
| AFFILIATIONS | University of Massachusetts Amherst |
| ORCID | 0000-0002-2174-8777 |
| VERIFIED | Yes |
| TOTAL WORKS | 17 |
| TOTAL CITATIONS | 39 |
| AUTHOR COUNT | 17 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1998 |
| LATEST PUBLICATION YEAR | 2026 |
| H-INDEX | 3 |
Perceived fairness of exam accommodations for students with special educational needs
Implementing an inclusive school means that teachers should use exam accommodations to foster the participation of students with Special Educational Needs (SEN). However, due to the emphasis on merit in most school systems, this practice can create a dilemma between equality and equity that can notably influence teachers’ perceived fairness of such accommodations. Three studies conducted in the French context with teachers, students and members o…
Analyzing the Dimensionality of O*NET Cognitive Ability Ratings to Inform Assessment Design
The O*NET database is an online repository of detailed information on the knowledge and skill requirements of thousands of jobs across the United States. Thus, it is a valuable resource for test developers who want to target cognitive and other abilities relevant to the contemporary workforce. In this study, we used multidimensional scaling (MDS) to analyze the mean importance ratings of the cognitive abilities and selected skills included in the…
Exploring Relationships among Test Takers’ Behaviors and Performance Using Response Process Data
Students exhibit many behaviors when responding to items on a computer-based test, but only some of these behaviors are relevant to estimating their proficiencies. In this study, we analyzed data from computer-based math achievement tests administered to elementary school students in grades 3 (ages 8–9) and 4 (ages 9–10). We investigated students’ response process data, including the total amount of time they spent on an item, the amount of time …
Validez y Validación para Pruebas Educativas y Psicológicas: Teoría y Recomendaciones
Antecedentes: La validez es uno de los conceptos más fundamentales en el contexto de pruebas educativas y psicológicas y se refiere al grado en el que la evidencia teórica y empírica respaldan las interpretaciones de las puntaciones obtenidas a partir de una prueba utilizada para un fin determinado. En este trabajo, trazamos la historia de la teoría de la validez, centrándonos en su evolución y explicamos cómo validar el uso de una prueba para un…
Targeted Linguistic Simplification of Science Test Items for English Learners
In this experimental study, 20 multiple-choice test items from the Massachusetts Grade 5 science test were linguistically simplified, and original and simplified test items were administered to 310 English learners (ELs) and 1,580 non-ELs in four Massachusetts school districts. This study tested the hypothesis that specific linguistic features of test items contributed to construct-irrelevant variance in science test scores of ELs. Simplification…
Student Assessment Opt Out and the Impact on Value-Added Measures of Teacher Quality
Student assessment nonparticipation (or opt out) has increased substantially in K-12 schools in states across the country. This increase in opt out has the potential to impact achievement and growth (or value-added) measures used for educator and institutional accountability. In this simulation study, we investigated the extent to which value-added measures of teacher quality are affected as a result of varying degrees of opt out, as well as a re…
A New Method for Analyzing Content Validity Data Using Multidimensional Scaling
Validity evidence based on test content is of essential importance in educational testing. One source for such evidence is an alignment study, which helps evaluate the congruence between tested objectives and those specified in the curriculum. However, the results of an alignment study do not always sufficiently capture the degree to which a test adequately represents the intended content domain. In this study, we present and evaluate a method fo…
Evaluating Structural Equivalence in Psychological Questionnaires Using Weighted Multidimensional Scaling
Cross-cultural scientists evaluate constructs across different cultural and linguistic groups. However, to make valid comparisons it is necessary to assure the equivalence of measurement instruments across populations. In this paper we use multidimensional scaling (MDS) to evaluate the structural equivalence of a measure of assertiveness developed originally for Mexican culture across Mexican and Spanish samples. 316 students from the Autonomous …
On Validity Theory and Test Validation
Lissitz and Samuelsen (2007) propose a new framework for conceptualizing test validity that separates analysis of test properties from analysis of the construct measured. In response, the author of this article reviews fundamental characteristics of test validity, drawing largely from seminal writings as well as from the accepted standards. He argues that a serious validation endeavor requires integration of construct theory, subjective analysis …
Evaluating the Predictive Validity of Graduate Management Admission Test Scores
Admissions data and first-year grade point average (GPA) data from 11 graduate management schools were analyzed to evaluate the predictive validity of Graduate Management Admission Test ® (GMAT ® ) scores and the extent to which predictive validity held across sex and race/ethnicity. The results indicated GMAT verbal and quantitative scores had substantial predictive validity, accounting for about 16% of the variance in graduate GPA beyond that p…
Evaluating Guidelines For Test Adaptations: A Methodological Analysis of Translation Quality
Guidelines for translating educational and psychological assessments for use across different languages and cultures have been developed by the International Test Commission and the Joint Committee on Standards for Educational and Psychological Testing. Common themes in these guidelines and standards are when translating items both judgmental and statistical techniques should be used to ensure item comparability across languages, and rigorous qua…
Unlabeling the Disabled: A Perspective on Flagging Scores From Accommodated Test Administrations
Accommodations to standard test administrations are granted on many tests for students who have one or more disabling conditions. In some instances, students’ scores from these nonstandard administrations are “flagged” to caution those who interpret the test score that the test was not administered under typical conditions. The practice of flagging such test scores is contentious. Some argue that it essentially informs others that a student has a…
Appraising item equivalence across multiple languages and cultures
Activity in the area of language testing is expanding beyond second language acquisition. In many contexts, tests that measure language skills are being translated into several different languages so that parallel versions exist for use in multilingual contexts. To ensure that translated items are equivalent to their original versions, both statistical and qualitative analyses are necessary. In this article, we describe a statistical method for e…
A Multitrait-Multimethod Validity Investigation of Scores from a Professional Licensure Examination
The construct validity of scores from the Uniform CPA Examination was investigated through the construction of a multitrait-multimethod matrix. The four sections of the exam were treated as distinct traits, and within-section subscores were created according to item format. First, Campbell and Fiske’s four criteria were used to evaluate the matrix. After correcting the correlations for attenuation due to unreliability, evidence was moderate for t…
A Multitrait-Multimethod Validity Investigation of Scores From a Professional Licensure Examination
Using Multidimensional Scaling to Assess the Dimensionality of Dichotomous Item Data
In this study, we investigated the utility of multidimensional scaling (MDS) for assessing the dimensionality of dichotomous test data. Two MDS proximity measures were studied: one based on the PC statistic proposed by Chen and Davison (1996), the other based on inter-item Euclidean distances. Stout's (1987) test of essential unidimensionality (DIMTEST) was also used as a standard for comparison. Twenty different conditions of unidimensional and …
The Construct of Content Validity
The Construct of Content Validity
Evaluating Guidelines For Test Adaptations: A Methodological Analysis of Translation Quality
Guidelines for translating educational and psychological assessments for use across different languages and cultures have been developed by the International Test Commission and the Joint Committee on Standards for Educational and Psychological Testing. Common themes in these guidelines and standards are when translating items both judgmental and statistical techniques should be used to ensure item comparability across languages, and rigorous qua…
On Validity Theory and Test Validation
Lissitz and Samuelsen (2007) propose a new framework for conceptualizing test validity that separates analysis of test properties from analysis of the construct measured. In response, the author of this article reviews fundamental characteristics of test validity, drawing largely from seminal writings as well as from the accepted standards. He argues that a serious validation endeavor requires integration of construct theory, subjective analysis …
Unlabeling the Disabled: A Perspective on Flagging Scores From Accommodated Test Administrations
Accommodations to standard test administrations are granted on many tests for students who have one or more disabling conditions. In some instances, students’ scores from these nonstandard administrations are “flagged” to caution those who interpret the test score that the test was not administered under typical conditions. The practice of flagging such test scores is contentious. Some argue that it essentially informs others that a student has a…
The Construct of Content Validity
Using Multidimensional Scaling to Assess the Dimensionality of Dichotomous Item Data
In this study, we investigated the utility of multidimensional scaling (MDS) for assessing the dimensionality of dichotomous test data. Two MDS proximity measures were studied: one based on the PC statistic proposed by Chen and Davison (1996), the other based on inter-item Euclidean distances. Stout's (1987) test of essential unidimensionality (DIMTEST) was also used as a standard for comparison. Twenty different conditions of unidimensional and …
A Multitrait-Multimethod Validity Investigation of Scores from a Professional Licensure Examination
The construct validity of scores from the Uniform CPA Examination was investigated through the construction of a multitrait-multimethod matrix. The four sections of the exam were treated as distinct traits, and within-section subscores were created according to item format. First, Campbell and Fiske’s four criteria were used to evaluate the matrix. After correcting the correlations for attenuation due to unreliability, evidence was moderate for t…
A Multitrait-Multimethod Validity Investigation of Scores From a Professional Licensure Examination
Appraising item equivalence across multiple languages and cultures
Activity in the area of language testing is expanding beyond second language acquisition. In many contexts, tests that measure language skills are being translated into several different languages so that parallel versions exist for use in multilingual contexts. To ensure that translated items are equivalent to their original versions, both statistical and qualitative analyses are necessary. In this article, we describe a statistical method for e…
Unlabeling the Disabled: A Perspective on Flagging Scores From Accommodated Test Administrations
Accommodations to standard test administrations are granted on many tests for students who have one or more disabling conditions. In some instances, students’ scores from these nonstandard administrations are “flagged” to caution those who interpret the test score that the test was not administered under typical conditions. The practice of flagging such test scores is contentious. Some argue that it essentially informs others that a student has a…
Evaluating the Predictive Validity of Graduate Management Admission Test Scores
Admissions data and first-year grade point average (GPA) data from 11 graduate management schools were analyzed to evaluate the predictive validity of Graduate Management Admission Test ® (GMAT ® ) scores and the extent to which predictive validity held across sex and race/ethnicity. The results indicated GMAT verbal and quantitative scores had substantial predictive validity, accounting for about 16% of the variance in graduate GPA beyond that p…
Evaluating Guidelines For Test Adaptations: A Methodological Analysis of Translation Quality
Guidelines for translating educational and psychological assessments for use across different languages and cultures have been developed by the International Test Commission and the Joint Committee on Standards for Educational and Psychological Testing. Common themes in these guidelines and standards are when translating items both judgmental and statistical techniques should be used to ensure item comparability across languages, and rigorous qua…
On Validity Theory and Test Validation
Lissitz and Samuelsen (2007) propose a new framework for conceptualizing test validity that separates analysis of test properties from analysis of the construct measured. In response, the author of this article reviews fundamental characteristics of test validity, drawing largely from seminal writings as well as from the accepted standards. He argues that a serious validation endeavor requires integration of construct theory, subjective analysis …
Evaluating Structural Equivalence in Psychological Questionnaires Using Weighted Multidimensional Scaling
Cross-cultural scientists evaluate constructs across different cultural and linguistic groups. However, to make valid comparisons it is necessary to assure the equivalence of measurement instruments across populations. In this paper we use multidimensional scaling (MDS) to evaluate the structural equivalence of a measure of assertiveness developed originally for Mexican culture across Mexican and Spanish samples. 316 students from the Autonomous …
A New Method for Analyzing Content Validity Data Using Multidimensional Scaling
Validity evidence based on test content is of essential importance in educational testing. One source for such evidence is an alignment study, which helps evaluate the congruence between tested objectives and those specified in the curriculum. However, the results of an alignment study do not always sufficiently capture the degree to which a test adequately represents the intended content domain. In this study, we present and evaluate a method fo…
Student Assessment Opt Out and the Impact on Value-Added Measures of Teacher Quality
Student assessment nonparticipation (or opt out) has increased substantially in K-12 schools in states across the country. This increase in opt out has the potential to impact achievement and growth (or value-added) measures used for educator and institutional accountability. In this simulation study, we investigated the extent to which value-added measures of teacher quality are affected as a result of varying degrees of opt out, as well as a re…
Targeted Linguistic Simplification of Science Test Items for English Learners
In this experimental study, 20 multiple-choice test items from the Massachusetts Grade 5 science test were linguistically simplified, and original and simplified test items were administered to 310 English learners (ELs) and 1,580 non-ELs in four Massachusetts school districts. This study tested the hypothesis that specific linguistic features of test items contributed to construct-irrelevant variance in science test scores of ELs. Simplification…
Validez y Validación para Pruebas Educativas y Psicológicas: Teoría y Recomendaciones
Antecedentes: La validez es uno de los conceptos más fundamentales en el contexto de pruebas educativas y psicológicas y se refiere al grado en el que la evidencia teórica y empírica respaldan las interpretaciones de las puntaciones obtenidas a partir de una prueba utilizada para un fin determinado. En este trabajo, trazamos la historia de la teoría de la validez, centrándonos en su evolución y explicamos cómo validar el uso de una prueba para un…
Exploring Relationships among Test Takers’ Behaviors and Performance Using Response Process Data
Students exhibit many behaviors when responding to items on a computer-based test, but only some of these behaviors are relevant to estimating their proficiencies. In this study, we analyzed data from computer-based math achievement tests administered to elementary school students in grades 3 (ages 8–9) and 4 (ages 9–10). We investigated students’ response process data, including the total amount of time they spent on an item, the amount of time …
Analyzing the Dimensionality of O*NET Cognitive Ability Ratings to Inform Assessment Design
The O*NET database is an online repository of detailed information on the knowledge and skill requirements of thousands of jobs across the United States. Thus, it is a valuable resource for test developers who want to target cognitive and other abilities relevant to the contemporary workforce. In this study, we used multidimensional scaling (MDS) to analyze the mean importance ratings of the cognitive abilities and selected skills included in the…
Perceived fairness of exam accommodations for students with special educational needs
Implementing an inclusive school means that teachers should use exam accommodations to foster the participation of students with Special Educational Needs (SEN). However, due to the emphasis on merit in most school systems, this practice can create a dilemma between equality and equity that can notably influence teachers’ perceived fairness of such accommodations. Three studies conducted in the French context with teachers, students and members o…
Psychology (16 works) · Computer Science (11 works) · Mathematics (9 works) · Psychometric Methodologies and Testing (9 works) · Psychometrics (7 works) · Mathematics education (6 works) · Statistics (6 works) · Social Psychology (5 works) · Applied Psychology (4 works) · Construct validity (4 works)