Randall D Penfield
Biographic Data
| ID | 5167281 |
|---|---|
| NAME | Randall D Penfield |
| GIVEN NAMES | Randall D |
| FAMILY NAME | Penfield |
| SIGNATURE | PENFIELD R D |
| AFFILIATIONS | University of Miami |
| VERIFIED | No |
| TOTAL WORKS | 12 |
| TOTAL CITATIONS | 4 |
| AUTHOR COUNT | 12 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1999 |
| LATEST PUBLICATION YEAR | 2013 |
| H-INDEX | 1 |
A Comparison of Uniform DIF Effect Size Estimators Under the MIMIC and Rasch Models
The Rasch model, a member of a larger group of models within item response theory, is widely used in empirical studies. Detection of uniform differential item functioning (DIF) within the Rasch model typically employs null hypothesis testing with a concomitant consideration of effect size (e.g., signed area [SA]). Parametric equivalence between confirmatory factor analysis under the multiple indicators, multiple causes (MIMIC) model and the Rasch…
How Are the Form and Magnitude of DIF Effects in Multiple-Choice Items Determined by Distractor-Level Invariance Effects
This article explores how the magnitude and form of differential item functioning (DIF) effects in multiple-choice items are determined by the underlying differential distractor functioning (DDF) effects, as modeled under the nominal response model. The results of a numerical investigation indicated that (a) the presence of one or more nonzero DDF effects implies a nonzero DIF effect; (b) the magnitude of the DDF effects creates an upper bound to…
Confidence Intervals for Squared Semipartial Correlation Coefficients
The increase in the squared multiple correlation coefficient (ΔR 2 ) associated with a variable in a regression equation is a commonly used measure of importance in regression analysis. Algina, Keselman, and Penfield found that intervals based on asymptotic principles were typically very inaccurate, even though the sample size was quite large (i.e., larger than 200). However, they also reported that probability coverage for the confidence interva…
Test-Based Grade Retention
A growing body of research showing that grade retention serves as an educationally low-quality placement has raised increasing concerns about whether the use of standardized tests in making decisions concerning grade retention conforms to current standards for appropriate and nondiscriminatory test use. This article examines the extent to which test-based grade retention policies comply with standards for fair and appropriate test use based on no…
Science Writing Achievement Among English Language Learners
As part of our professional development intervention, this study examined third-grade ELL students' writing achievement that included “form” (i.e., conventions, organization, and style/voice) and “content” (i.e., specific knowledge and understanding of science) in expository science writing. The study included six treatment schools from a large urban school district. Data were collected from three different groups of students over 3 separate year…
Methods for Assessing Item, Step, and Threshold Invariance in Polytomous Items Following the Partial Credit Model
Measurement invariance in the partial credit model (PCM) can be conceptualized in several different but compatible ways. In this article the authors distinguish between three forms of measurement invariance in the PCM: step invariance, item invariance, and threshold invariance. Approaches for modeling these three forms of invariance are proposed, and the mathematical relationship between the three forms is established. Parametric and contingency …
Confidence Intervals for an Effect Size Measure in Multiple Linear Regression
The increase in the squared multiple correlation coefficient (ΔR 2 ) associated with a variable in a regression equation is a commonly used measure of importance in regression analysis. The coverage probability that an asymptotic and percentile bootstrap confidence interval includes Δρ 2 was investigated. As expected, coverage probability for the asymptotic confidence interval was often inadequate (outside the interval .925 to .975 for a 95% conf…
Estimating the Standard Error of the Maximum Likelihood Ability Estimator in Adaptive Testing Using the Posterior-Weighted Test Information Function
The standard error of the maximum likelihood ability estimator is commonly estimated by evaluating the test information function at an examinee's current maximum likelihood estimate (a point estimate) of ability. Because the test information function evaluated at the point estimate may differ from the test information function evaluated at an examinee's true ability value, the estimated standard error may be biased under certain conditions. This …
Confidence Interval Coverage for Cohen's Effect Size Statistic
Kelley compared three methods for setting a confidence interval (CI) around Cohen's standardized mean difference statistic: the noncentral- t-based, percentile (PERC) bootstrap, and biased-corrected and accelerated (BCA) bootstrap methods under three conditions of nonnormality, eight cases of sample size, and six cases of population effect size (ES) magnitude. Kelley recommended the BCA bootstrap method. The authors expand on his investigation by…
Effect Sizes and their Intervals
Probability coverage for eight different confidence intervals (CIs) of measures of effect size (ES) in a two-level repeated measures design was investigated. The CIs and measures of ES differed with regard to whether they used least squares or robust estimates of central tendency and variability, whether the end critical points of the interval were obtained using a theoretical or an empirical sampling distribution, and whether the ESs used a pool…
Applying the Breslow-Day Test of Trend in Odds Ratio Heterogeneity to the Analysis of Nonuniform DIF
This article applies the Breslow-Day test of trend in odds ratio heterogeneity (BD) to the detection of nonuniform DIF. A simulation study was conducted to assess the power and Type I error rate of BD, as well as a combined decision rule (CDR) whereby a decision of the existence of DIF was based on a combination of the decisions made using BD and the Mantel-Haenszel chi-square. The results indicated that CDR displayed good Type I error rate and p…
A Procedure for Detecting Student Profile Patterns in a Performance Assessment
This study investigates student score profiles of the mathematics component of the 1997 Ontario grade 3 assessment. In addition to an overall score, students are given scores on three knowledge or skill dimensions, and five scores on content strands. The purpose of this investigation was threefold: (a) to assess the extent to which student profiles contain differentially diagnostic information, (b) to examine classroom-level patterns in the stude…
Test-Based Grade Retention
A growing body of research showing that grade retention serves as an educationally low-quality placement has raised increasing concerns about whether the use of standardized tests in making decisions concerning grade retention conforms to current standards for appropriate and nondiscriminatory test use. This article examines the extent to which test-based grade retention policies comply with standards for fair and appropriate test use based on no…
A Procedure for Detecting Student Profile Patterns in a Performance Assessment
This study investigates student score profiles of the mathematics component of the 1997 Ontario grade 3 assessment. In addition to an overall score, students are given scores on three knowledge or skill dimensions, and five scores on content strands. The purpose of this investigation was threefold: (a) to assess the extent to which student profiles contain differentially diagnostic information, (b) to examine classroom-level patterns in the stude…
Applying the Breslow-Day Test of Trend in Odds Ratio Heterogeneity to the Analysis of Nonuniform DIF
This article applies the Breslow-Day test of trend in odds ratio heterogeneity (BD) to the detection of nonuniform DIF. A simulation study was conducted to assess the power and Type I error rate of BD, as well as a combined decision rule (CDR) whereby a decision of the existence of DIF was based on a combination of the decisions made using BD and the Mantel-Haenszel chi-square. The results indicated that CDR displayed good Type I error rate and p…
Effect Sizes and their Intervals
Probability coverage for eight different confidence intervals (CIs) of measures of effect size (ES) in a two-level repeated measures design was investigated. The CIs and measures of ES differed with regard to whether they used least squares or robust estimates of central tendency and variability, whether the end critical points of the interval were obtained using a theoretical or an empirical sampling distribution, and whether the ESs used a pool…
Confidence Interval Coverage for Cohen's Effect Size Statistic
Kelley compared three methods for setting a confidence interval (CI) around Cohen's standardized mean difference statistic: the noncentral- t-based, percentile (PERC) bootstrap, and biased-corrected and accelerated (BCA) bootstrap methods under three conditions of nonnormality, eight cases of sample size, and six cases of population effect size (ES) magnitude. Kelley recommended the BCA bootstrap method. The authors expand on his investigation by…
Confidence Intervals for an Effect Size Measure in Multiple Linear Regression
The increase in the squared multiple correlation coefficient (ΔR 2 ) associated with a variable in a regression equation is a commonly used measure of importance in regression analysis. The coverage probability that an asymptotic and percentile bootstrap confidence interval includes Δρ 2 was investigated. As expected, coverage probability for the asymptotic confidence interval was often inadequate (outside the interval .925 to .975 for a 95% conf…
Estimating the Standard Error of the Maximum Likelihood Ability Estimator in Adaptive Testing Using the Posterior-Weighted Test Information Function
The standard error of the maximum likelihood ability estimator is commonly estimated by evaluating the test information function at an examinee's current maximum likelihood estimate (a point estimate) of ability. Because the test information function evaluated at the point estimate may differ from the test information function evaluated at an examinee's true ability value, the estimated standard error may be biased under certain conditions. This …
Methods for Assessing Item, Step, and Threshold Invariance in Polytomous Items Following the Partial Credit Model
Measurement invariance in the partial credit model (PCM) can be conceptualized in several different but compatible ways. In this article the authors distinguish between three forms of measurement invariance in the PCM: step invariance, item invariance, and threshold invariance. Approaches for modeling these three forms of invariance are proposed, and the mathematical relationship between the three forms is established. Parametric and contingency …
Science Writing Achievement Among English Language Learners
As part of our professional development intervention, this study examined third-grade ELL students' writing achievement that included “form” (i.e., conventions, organization, and style/voice) and “content” (i.e., specific knowledge and understanding of science) in expository science writing. The study included six treatment schools from a large urban school district. Data were collected from three different groups of students over 3 separate year…
Confidence Intervals for Squared Semipartial Correlation Coefficients
The increase in the squared multiple correlation coefficient (ΔR 2 ) associated with a variable in a regression equation is a commonly used measure of importance in regression analysis. Algina, Keselman, and Penfield found that intervals based on asymptotic principles were typically very inaccurate, even though the sample size was quite large (i.e., larger than 200). However, they also reported that probability coverage for the confidence interva…
Test-Based Grade Retention
A growing body of research showing that grade retention serves as an educationally low-quality placement has raised increasing concerns about whether the use of standardized tests in making decisions concerning grade retention conforms to current standards for appropriate and nondiscriminatory test use. This article examines the extent to which test-based grade retention policies comply with standards for fair and appropriate test use based on no…
How Are the Form and Magnitude of DIF Effects in Multiple-Choice Items Determined by Distractor-Level Invariance Effects
This article explores how the magnitude and form of differential item functioning (DIF) effects in multiple-choice items are determined by the underlying differential distractor functioning (DDF) effects, as modeled under the nominal response model. The results of a numerical investigation indicated that (a) the presence of one or more nonzero DDF effects implies a nonzero DIF effect; (b) the magnitude of the DDF effects creates an upper bound to…
A Comparison of Uniform DIF Effect Size Estimators Under the MIMIC and Rasch Models
The Rasch model, a member of a larger group of models within item response theory, is widely used in empirical studies. Detection of uniform differential item functioning (DIF) within the Rasch model typically employs null hypothesis testing with a concomitant consideration of effect size (e.g., signed area [SA]). Parametric equivalence between confirmatory factor analysis under the multiple indicators, multiple causes (MIMIC) model and the Rasch…
Mathematics (9 works) · Statistics (9 works) · Psychology (6 works) · Econometrics (5 works) · Sample size determination (5 works) · Computer Science (4 works) · Confidence interval (4 works) · Coverage probability (4 works) · Psychometric Methodologies and Testing (4 works) · Psychometrics (4 works)