Frederic M Lord
Biographic Data
| ID | 1660381 |
|---|---|
| NAME | Frederic M Lord |
| GIVEN NAMES | Frederic M |
| FAMILY NAME | Lord |
| SIGNATURE | LORD F M |
| AFFILIATIONS | Educational Testing Service |
| ORCID | 0000-0002-3702-4124 |
| VERIFIED | Yes |
| TOTAL WORKS | 31 |
| TOTAL CITATIONS | 65 |
| AUTHOR COUNT | 31 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 1944 |
| LATEST PUBLICATION YEAR | 2008 |
| H-INDEX | 3 |
Statistical Theories of Mental Test Scores
Developing a Common Metric in Item Response Theory
A common problem arises when independent esti mates of item parameters from two separate data sets must be expressed in the same metric. This problem is frequently confronted in studies of horizontal and ver tical equating and in studies of item bias. This paper discusses a number of methods for finding the appro priate transformation from one metric to another met ric and presents a new method. Data are given com paring this new method with a cu…
A Prediction Interval for a Score on a Parallel Test Form
Given any observed number-right score on a test, a method is described for obtaining a prediction interval for the corresponding number-right score on a randomly parallel form of the same test. The interval can be written down directly from published tables of the hypergeometric distribution
Significance Test for a Partial Correlation Corrected for Attenuation
Correction for attenuation is important for partial correlations because not even the sign of the partial between true scores can be inferred safely from the partial between observed (fallible) scores. Methods for inferring the corrected partial are discussed. Unfortunately, the corrected partial will sometimes have an overwhelming sampling error. A significance test is developed that largely circumvents this problem in those cases where it is en…
Power Scores Estimated By Item Characteristic Curves
A method for estimating power scores is described. By way of illustration, it is applied to 21 students who were improperly timed on a standard test. Some empirical results are given in support of the estimation procedure
Testing if two measuring procedures measure the same dimension
Robbins-Monro Procedures for Tailored Testing
In tailored testing, we try to choose for administration items at a difficulty level matching the examinee's ability, which we infer from his responses to items already administered. Robbins-Monro procedures for selecting items and for estimating the examinee's ability are evaluated. Various ideas of use for tailored testing emerge
A Theoretical Study of the Measurement Effectiveness of Flexilevel Tests
A flexilevel test is found to be inferior to a peaked conventional test for measuring examinees in the middle of the ability range, superior for examinees at the extremes. Throughout the entire range of ability, a flexilevel test is much superior to any conventional test that attempts to provi,le accurate measurement at both extremes. A THEORETICAL STUDY OF THE MEASUMMENT EFFECTIVENESS OF FLEXILEVEL TESTS* A conventional test becomes a flexilevel…
A Computer Program for Estimating True-Score Distributions and Graduating Observed-Score Distributions
The program takes a frequency distribution of number-right test scores and produces (1) an estimated distribution of true scores for the group tested, computed on the assumption that the errors of measurement have a certain compound binomial distribution, (2) the corresponding smoothed distribution of actual scores, and (3) a chi-square for comparing the smoothed and the actual distributions. All instructions and background information necessary …
Statistical adjustments when comparing preexisting groups
An illustration is given showing why the analysis of covariance usually does not provide the appropriate adjustment to compensate for preexisting differences between nonexperimental groups
An Analysis of the Verbal Scholastic Aptitude Test Using Birnbaum's Three-Parameter Logistic Model
A paradox in the interpretation of group comparisons
It is common practice in behavioral research, and in other areas, to apply the analysis of covariance in the investigation of preexisting natural groups. The research worker is usually interested in some criterion variable (y) and would like to make allowances for the fact that his groups are not matched on some important independent variable or control variable (x). The situation is such that observed differences in the dependent variable might …
The Effect of Random Guessing on Test Validity
Cutting Scores and Errors of Measurement-a Second Case
The effect of errors of measurement on the shape of optimum selection regions is investigated under a second mathematical model different from that previously used. It is found that the second model leads to the same conclusions as the first
Formula Scoring and Validity
Formulas are derived for a rather restricted, situation showing the decrement in test validity that might be expected from random guessing and the increment in validity that might be expected from a priori formula scoring. Numerical illustrations are given. The reduction in validity that occurs when formula scoring is abandoned is generally found to be small in terms of the usual correlational scale, but not necessarily small when considered in t…
Estimating Norms by Item-Sampling
If a 70-item test is to be normed for a norms population of 1,000 individuals, is it better to give the entire test to a sample of 100 individuals, or to give different samples of 7 items to each of the 1,000 individuals? When tried out on actual data, the latter, “item-sampling” method was found to provide an estimate of the “true” norms distribution slightly superior to that obtainable in the majority of cases by the former method
Tests of the Same Length do Have the Same Standard Error of Measurement
An Index of the Discriminating Power of a Test at Different Parts of the Score Range
An important characteristic of any test is its ability to discriminate among examinees who differ in ability. The conventional reliability and validity coefficients are indices of discrimination for the test as a whole; however, except under certain limited conditions, these over-all indices do not apply at all points along the score scale. The purpose of the present paper is to provide such an index and illustrate its use
Inferences About True Scores from Parallel Test Forms1
An Empirical Study of the Stability of a Group Mean in Relation to the Distribution of Test Items Among Students
Further Problems in the Measurement of Growth
Do Tests of the Same Length Have the Same Standard Errors of Measurement
The Measurement of Growth
A regression formula is derived for estimating a student's true gain from his initial and final test scores. A formula is given for the reliability of estimates thus obtained. A numerical example is given, illustrating some marked inadequacies of the simple, "common-sense" procedure that estimates gain by subtracting initial from final score. A convenient graphic procedure is presented for grouping all examinees tested according to the size of th…
A Survey of Observed Test-Score Distributions With Respect to Skewness and Kurtosis1
THE purpose of the present survey of data was to check em-pirically on two hypotheses: i. Easy tests tend to yield negatively skewed score distribu-tions; hard tests, positively skewed distributions. This hypothe-sis is so prevalent that references need not be cited. 2. Symmetric score distributions are ordinarily platykurtic. This conclusion has been reached from theoretical considerations by Keats (5) (a symmetric beta function is always platyk…
Estimating Test Reliability
Formulas for parallel-form reliability coefficients are derived from two improved definitions of parallelism. One definition, based on randomly parallel test forms, leads to a new and useful derivation for the KR formula-21 coefficient. The other, based on matched test forms, leads to a formula for the least upper bound of the test reliability. A numerical example is given, showing how a standard error of measurement for each separate examinee is…
A paradox in the interpretation of group comparisons
It is common practice in behavioral research, and in other areas, to apply the analysis of covariance in the investigation of preexisting natural groups. The research worker is usually interested in some criterion variable (y) and would like to make allowances for the fact that his groups are not matched on some important independent variable or control variable (x). The situation is such that observed differences in the dependent variable might …
On the Statistical Treatment of Football Numbers
Statistical adjustments when comparing preexisting groups
An illustration is given showing why the analysis of covariance usually does not provide the appropriate adjustment to compensate for preexisting differences between nonexperimental groups
A Report On Scholarship Examinations Given in Latin American Countries for the Selection of Students To Be Trained in Meteorology
Preparation of Profile Charts on the IBM Tabulator
The Relation of Test Score to the Trait Underlying the Test
On the Statistical Treatment of Football Numbers
"Further Comment on "Football Numbers
A Survey of Observed Test-Score Distributions With Respect to Skewness and Kurtosis1
THE purpose of the present survey of data was to check em-pirically on two hypotheses: i. Easy tests tend to yield negatively skewed score distribu-tions; hard tests, positively skewed distributions. This hypothe-sis is so prevalent that references need not be cited. 2. Symmetric score distributions are ordinarily platykurtic. This conclusion has been reached from theoretical considerations by Keats (5) (a symmetric beta function is always platyk…
Estimating Test Reliability
Formulas for parallel-form reliability coefficients are derived from two improved definitions of parallelism. One definition, based on randomly parallel test forms, leads to a new and useful derivation for the KR formula-21 coefficient. The other, based on matched test forms, leads to a formula for the least upper bound of the test reliability. A numerical example is given, showing how a standard error of measurement for each separate examinee is…
"Some perspectives on "the attenuation paradox in test theory
Four points relating to Dr. Loevinger's “attenuation paradox” have “been brought forward: The usual product-moment “validity” coefficient is inadequate for any discussion of the paradox. A curvilinear correlation coefficient must be used. The “region of paradox” is still found when the correct coefficient is used, although its size is reduced. A greatly magnified notion of the extent to which the “paradox” occurs in actual achievement and aptitud…
The Measurement of Growth
A regression formula is derived for estimating a student's true gain from his initial and final test scores. A formula is given for the reliability of estimates thus obtained. A numerical example is given, illustrating some marked inadequacies of the simple, "common-sense" procedure that estimates gain by subtracting initial from final score. A convenient graphic procedure is presented for grouping all examinees tested according to the size of th…
Do Tests of the Same Length Have the Same Standard Errors of Measurement
An Empirical Study of the Stability of a Group Mean in Relation to the Distribution of Test Items Among Students
Further Problems in the Measurement of Growth
Tests of the Same Length do Have the Same Standard Error of Measurement
An Index of the Discriminating Power of a Test at Different Parts of the Score Range
An important characteristic of any test is its ability to discriminate among examinees who differ in ability. The conventional reliability and validity coefficients are indices of discrimination for the test as a whole; however, except under certain limited conditions, these over-all indices do not apply at all points along the score scale. The purpose of the present paper is to provide such an index and illustrate its use
Inferences About True Scores from Parallel Test Forms1
Estimating Norms by Item-Sampling
If a 70-item test is to be normed for a norms population of 1,000 individuals, is it better to give the entire test to a sample of 100 individuals, or to give different samples of 7 items to each of the 1,000 individuals? When tried out on actual data, the latter, “item-sampling” method was found to provide an estimate of the “true” norms distribution slightly superior to that obtainable in the majority of cases by the former method
Cutting Scores and Errors of Measurement-a Second Case
The effect of errors of measurement on the shape of optimum selection regions is investigated under a second mathematical model different from that previously used. It is found that the second model leads to the same conclusions as the first
Formula Scoring and Validity
Formulas are derived for a rather restricted, situation showing the decrement in test validity that might be expected from random guessing and the increment in validity that might be expected from a priori formula scoring. Numerical illustrations are given. The reduction in validity that occurs when formula scoring is abandoned is generally found to be small in terms of the usual correlational scale, but not necessarily small when considered in t…
The Effect of Random Guessing on Test Validity
A paradox in the interpretation of group comparisons
It is common practice in behavioral research, and in other areas, to apply the analysis of covariance in the investigation of preexisting natural groups. The research worker is usually interested in some criterion variable (y) and would like to make allowances for the fact that his groups are not matched on some important independent variable or control variable (x). The situation is such that observed differences in the dependent variable might …
An Analysis of the Verbal Scholastic Aptitude Test Using Birnbaum's Three-Parameter Logistic Model
A Computer Program for Estimating True-Score Distributions and Graduating Observed-Score Distributions
The program takes a frequency distribution of number-right test scores and produces (1) an estimated distribution of true scores for the group tested, computed on the assumption that the errors of measurement have a certain compound binomial distribution, (2) the corresponding smoothed distribution of actual scores, and (3) a chi-square for comparing the smoothed and the actual distributions. All instructions and background information necessary …
Statistical adjustments when comparing preexisting groups
An illustration is given showing why the analysis of covariance usually does not provide the appropriate adjustment to compensate for preexisting differences between nonexperimental groups
Robbins-Monro Procedures for Tailored Testing
In tailored testing, we try to choose for administration items at a difficulty level matching the examinee's ability, which we infer from his responses to items already administered. Robbins-Monro procedures for selecting items and for estimating the examinee's ability are evaluated. Various ideas of use for tailored testing emerge
A Theoretical Study of the Measurement Effectiveness of Flexilevel Tests
A flexilevel test is found to be inferior to a peaked conventional test for measuring examinees in the middle of the ability range, superior for examinees at the extremes. Throughout the entire range of ability, a flexilevel test is much superior to any conventional test that attempts to provi,le accurate measurement at both extremes. A THEORETICAL STUDY OF THE MEASUMMENT EFFECTIVENESS OF FLEXILEVEL TESTS* A conventional test becomes a flexilevel…
Mathematics (27 works) · Statistics (26 works) · Psychology (25 works) · Econometrics (14 works) · Test (biology) (14 works) · Computer Science (12 works) · Psychometrics (7 works) · Psychometric Methodologies and Testing (6 works) · Geology (4 works) · Geology (4 works)