Natalja Menold
Datos Biográficos
| ID | 124608 |
|---|---|
| NOMBRE | Natalja Menold |
| NOMBRES | Natalja |
| APELLIDO | Menold |
| FIRMA | MENOLD N |
| AFILIACIONES | GESIS - Leibniz Institute for the Social Sciences |
| ORCID | 0000-0003-1106-474X |
| VERIFICADO | Sí |
| TOTAL DE OBRAS | 32 |
| TOTAL DE CITAS | 49 |
| TOTAL COMO AUTOR | 31 |
| TOTAL COMO EDITOR | 1 |
| PRIMER AÑO DE PUBLICACIÓN | 2003 |
| AÑO MÁS RECIENTE DE PUBLICACIÓN | 2025 |
| ÍNDICE H | 3 |
Effect of Cognitive Pretests on Measurement Invariance and Reliability in Quality of Life Measures
The issue evaluated was how cross-cultural cognitive pretests affect the reliability and comparability of data when studying refugees and using cross-language comparisons. Three instruments employed to assess general and health-related quality of life were revised based on the findings of cognitive pretests. The versions before and after cognitive pretests were randomly assigned to respondents in two web survey studies. The first study recruited …
Improving Cross-Cultural Comparability of Measures on Gender and Age Stereotypes by Means of Piloting Methods
The study addresses the effects of piloting methods on the cross-cultural comparability and reliability of the measurement of gender and age stereotypes. We conducted a summative evaluation of expert reviews, cognitive pretests and web probing. We first piloted a gender role, an ageism, and a children stereotypes instrument in German and American English. We then randomly assigned the original and piloted versions to respondents in Germany and th…
Overall Model Fit Does Not Imply Linearity in Longitudinal Structural Equation Models
This article is concerned with the assumption of linear temporal development that is often advanced in structural equation modeling-based longitudinal research. The linearity hypothesis is implemented in particular in the popular intercept-and-slope model as well as in more general models containing it as a component, such as longitudinal structural models with covariates, or models for the study of predictors and correlates of change. In empiric…
Coefficient alpha and reliability of communication science measurement scales
This note intends to complement the recent discussion in Hayes and Coutts (2020) by focusing on (i) the loading equality condition for the population identity of coefficient alpha and reliability of multiple-indicator measurement scales, as well as (ii) the potential utility of alpha when this condition is not satisfied. We show that the alpha and reliability coefficients can be very close at the population level in certain cases of loading inequ…
Linking survey and Facebook data
Do Different Devices Perform Equally Well with Different Numbers of Scale Points and Response Formats? A test of measurement invariance and reliability
Research on mixed devices in web surveys is in its infancy. Using a randomized experiment, we investigated device effects (desktop PC, tablet and mobile phone) for six response formats and four different numbers of scale points. N = 5,077 members of an online access panel participated in the experiment. An exact test of measurement invariance and Composite Reliability were investigated. The results provided full data comparability for devices and…
On the Importance of Coefficient Alpha for Measurement Research
The population relationship between coefficient alpha and scale reliability is studied in the widely used setting of unidimensional multicomponent measuring instruments. It is demonstrated that for any set of component loadings on the common factor, regardless of the extent of their inequality, the discrepancy between alpha and reliability can be arbitrarily small in any considered population and hence practically ignorable. In addition, the set …
Measurement invariance in the social sciences
Verbalization of Rating Scales Taking Account of Their Polarity
While numerical bipolar rating scales may evoke positivity bias, little is known about the corresponding bias in verbal bipolar rating scales. The choice of verbalization of the middle category may lead to response bias, particularly if it is not in line with the scale polarity. Unipolar and bipolar seven-category rating scales in which the verbalizations of the middle categories matched or did not match the implemented polarity were investigated…
Evaluation of Second- and Third-Level Variance Proportions in Multilevel Designs With Completely Observed Populations
Two- and three-level designs in educational and psychological research can involve entire populations of Level-3 and possibly Level-2 units, such as schools and educational districts nested within a given state, or neighborhoods and counties in a state. Such a design is of increasing relevance in empirical research owing to the growing popularity of large-scale studies in these and cognate disciplines. The present note discusses a readily applica…
On the Relationship Between Item Stem Formulation and Criterion Validity of Multiple-Component Measuring Instruments
The possible dependency of criterion validity on item formulation in a multicomponent measuring instrument is examined. The discussion is concerned with evaluation of the differences in criterion validity between two or more groups (populations/subpopulations) that have been administered instruments with items having differently formulated item stems. The case of complex item stems involving two stimuli description sentences (double-barreled ques…
How Do Reverse-keyed Items in Inventories Affect Measurement Quality and Information Processing
In randomized experiments, inventories with reverse-keyed items are compared with inventories in which all the items are either positively or negatively associated with the underlying concept. The results show that with reverse keying, a control of the potential bias was not sufficient; likewise, the factorial structure, reliability, and validity were negatively affected. An eye-tracking study revealed that respondents did not process information…
Rating-Scale Labeling in Online Surveys
Unlike other data collection modes, the effect of labeling rating scales on reliability and validity, as relevant aspects of measurement quality, has seldom been addressed in online surveys. In this study, verbal and numeric rating scales were compared in split-ballot online survey experiments. In the first experiment, respondents' cognitive processes were observed by means of eye tracking, that is, determining the respondent's fixations in diffe…
Qualitätssicherung sozialwissenschaftlicher Erhebungsinstrumente
Qualitätssicherung sozialwissenschaftlicher Erhebungsinstrumente
Multiple-Component Measurement Instruments in Heterogeneous Populations
This note confronts the common use of a single coefficient alpha as an index informing about reliability of a multicomponent measurement instrument in a heterogeneous population. Two or more alpha coefficients could instead be meaningfully associated with a given instrument in finite mixture settings, and this may be increasingly more likely the case in empirical educational and psychological research. It is argued that in such situations explici…
Gütekriterien quantitativer Sozialforschung
Questionnaire Design and Translation for Refugee Populations
Surveying the refugee population poses particular challenges: what measurement and culture effects need to be taken into account? Are some of the constructs related to refugees unique or can constructs used in other surveys be adapted? Due to considerable variation in educational background, in trauma history or in perception of ethnicity or gender roles in refugee populations, one needs to raise the question whether a one-size-fits-all approach …
Response Bias and Reliability in Verbal Agreement Rating Scales
Verbal rating scale polarity and verbalizations of the middle category that do not match the polarity in agreement rating scales were investigated. Two randomized web survey experiments were conducted using a probability panel of German Internet users. The classical bipolar “disagree/agree” verbalization was compared with the unipolar “do not agree/agree” alternative. In both experiments, attitudes on gender roles were measured using a different …
Studying Latent Criterion Validity for Complex Structure Measuring Instruments Using Latent Variable Modeling
Validity coefficients for multicomponent measuring instruments are known to be affected by measurement error that attenuates them, affects associated standard errors, and influences results of statistical tests with respect to population parameter values. To account for measurement error, a latent variable modeling approach is discussed that allows point and interval estimation of the relationship of an underlying latent factor to a criterion var…
Reliability of Scales With Second-Order Structure
A readily applicable procedure is discussed that allows evaluation of the discrepancy between the popular coefficient alpha and the reliability coefficient of a scale with second-order factorial structure that is frequently of relevance in empirical educational and psychological research. The approach is developed within the framework of the widely used latent variable modeling methodology and permits point and interval estimation of the slippage…
Examining Measurement Invariance and Differential Item Functioning With Discrete Latent Construct Indicators
A latent variable modeling method for studying measurement invariance when evaluating latent constructs with multiple binary or binary scored items with no guessing is outlined. The approach extends the continuous indicator procedure described by Raykov and colleagues, utilizes similarly the false discovery rate approach to multiple testing, and permits one to locate violations of measurement invariance in loading or threshold parameters. The dis…
The Impact of Payment and Respondents' Participation on Interviewers' Accuracy in Face-to-face Surveys
In face-to-face interviews, accurate work by interviewers is crucial for ensuring high-quality survey data. In a field experiment, payment of interviewers, legitimation of falsification behavior, and respondents' willingness to participate were experimentally varied. The impact of these factors on interviewers' accuracy during fieldwork was investigated. Low accuracy was operationalized, for instance, as noncompliance with the instructions concer…
Can Reliability of Multiple Component Measuring Instruments Depend on Response Option Presentation Mode
This article examines the possible dependency of composite reliability on presentation format of the elements of a multi-item measuring instrument. Using empirical data and a recent method for interval estimation of group differences in reliability, we demonstrate that the reliability of an instrument need not be the same when polarity of the response options for its individual components differs across administrations of the instrument. Implicat…
Methodological Aspects of Focus Groups in Health Research
Although focus groups are commonly used in health research to explore the perspectives of patients or health care professionals, few studies consider methodological aspects in this specific context. For this reason, we interviewed nine researchers who had conducted focus groups in the context of a project devoted to the development of an electronic personal health record. We performed qualitative content analysis on the interview data relating to…
Measurement invariance in the social sciences
How Do Respondents Attend to Verbal Labels in Rating Scales
Two formats of labeling in rating scales are commonly used in questionnaires: verbal labels for end categories only (END form) and verbal labels for each of the categories (ALL form). We examine attention processes and respondents' burden in using verbal labels in rating scales. Attention was tracked in a laboratory setting employing eye-tracking technology. The results of the two experiments are presented: One applied seven and the other applied…
The Influence of the Answer Box Size on Item Nonresponse to Open-Ended Questions in a Web Survey
This article investigates item nonresponse in open-ended survey questions because such item nonresponse is much higher than in closed questions. The difference is a result of the higher cognitive burden placed on the respondent. To study item nonresponse, we manipulate different questionnaire design characteristics, such as the size of the answer box and the inclusion of motivation texts, as well as respondent-specific characteristics, in a rando…
Rating-Scale Labeling in Online Surveys
Unlike other data collection modes, the effect of labeling rating scales on reliability and validity, as relevant aspects of measurement quality, has seldom been addressed in online surveys. In this study, verbal and numeric rating scales were compared in split-ballot online survey experiments. In the first experiment, respondents' cognitive processes were observed by means of eye tracking, that is, determining the respondent's fixations in diffe…
Measurement of Latent Variables With Different Rating Scales
Effects of rating scale forms on cross-sectional reliability and measurement equivalence were investigated. A randomized experimental design was implemented, varying category labels and number of categories. The participants were 800 students at two German universities. In contrast to previous research, reliability assessment method was used, which relies on the congeneric measurement model. The experimental manipulation had differential effects …
Questionnaire Design and Translation for Refugee Populations
Surveying the refugee population poses particular challenges: what measurement and culture effects need to be taken into account? Are some of the constructs related to refugees unique or can constructs used in other surveys be adapted? Due to considerable variation in educational background, in trauma history or in perception of ethnicity or gender roles in refugee populations, one needs to raise the question whether a one-size-fits-all approach …
Response Bias and Reliability in Verbal Agreement Rating Scales
Verbal rating scale polarity and verbalizations of the middle category that do not match the polarity in agreement rating scales were investigated. Two randomized web survey experiments were conducted using a probability panel of German Internet users. The classical bipolar “disagree/agree” verbalization was compared with the unipolar “do not agree/agree” alternative. In both experiments, attitudes on gender roles were measured using a different …
How Do Real and Falsified Data Differ? Psychology of Survey Response as a Source of Falsification Indicators in Face-to-Face Surveys
peer reviewed
Coefficient alpha and reliability of communication science measurement scales
This note intends to complement the recent discussion in Hayes and Coutts (2020) by focusing on (i) the loading equality condition for the population identity of coefficient alpha and reliability of multiple-indicator measurement scales, as well as (ii) the potential utility of alpha when this condition is not satisfied. We show that the alpha and reliability coefficients can be very close at the population level in certain cases of loading inequ…
Concepts for usable patterns of groupware applications[39] (abstract only)
Patterns, which are based on in-depth practical experience, can be instructing for the design of groupware applications as socio-technical systems. On the basis of a summary of the concept of patterns - as elaborated by the architect Christopher Alexander - its adoptions within computer science are retraced and relationships to the area of groupware are described. General principles for patterns within this domain are formulated and supported by …
How to Use Information Technology for Cooperative Work
How Do Real and Falsified Data Differ? Psychology of Survey Response as a Source of Falsification Indicators in Face-to-Face Surveys
peer reviewed
Qualitätsstandards zur Entwicklung, Anwendung und Bewertung von Messinstrumenten in der sozialwissenschaftlichen Umfrageforschung
Der vom Bundesministerium für Bildung und Forschung (BMBF) geförderte Rat für Sozial- und Wirtschaftsdaten (RatSWD) berät seit 2004 die Bundesregierung und die Regierungen der Länder in Fragen der Erweiterung und Verbesserung der Forschungsinfrastruktur für die empirischen Sozial-, Wirtschafts- und Verhaltenswissenschaften (SWV). Ende 2010 hat sich der RatSWD der Fragestellung gewidmet, wie sich die Qualität von Erhebungsinstrumenten in den Sozia…
The Influence of the Answer Box Size on Item Nonresponse to Open-Ended Questions in a Web Survey
This article investigates item nonresponse in open-ended survey questions because such item nonresponse is much higher than in closed questions. The difference is a result of the higher cognitive burden placed on the respondent. To study item nonresponse, we manipulate different questionnaire design characteristics, such as the size of the answer box and the inclusion of motivation texts, as well as respondent-specific characteristics, in a rando…
How Do Respondents Attend to Verbal Labels in Rating Scales
Two formats of labeling in rating scales are commonly used in questionnaires: verbal labels for end categories only (END form) and verbal labels for each of the categories (ALL form). We examine attention processes and respondents' burden in using verbal labels in rating scales. Attention was tracked in a laboratory setting employing eye-tracking technology. The results of the two experiments are presented: One applied seven and the other applied…
Can Reliability of Multiple Component Measuring Instruments Depend on Response Option Presentation Mode
This article examines the possible dependency of composite reliability on presentation format of the elements of a multi-item measuring instrument. Using empirical data and a recent method for interval estimation of group differences in reliability, we demonstrate that the reliability of an instrument need not be the same when polarity of the response options for its individual components differs across administrations of the instrument. Implicat…
Methodological Aspects of Focus Groups in Health Research
Although focus groups are commonly used in health research to explore the perspectives of patients or health care professionals, few studies consider methodological aspects in this specific context. For this reason, we interviewed nine researchers who had conducted focus groups in the context of a project devoted to the development of an electronic personal health record. We performed qualitative content analysis on the interview data relating to…
Measurement of Latent Variables With Different Rating Scales
Effects of rating scale forms on cross-sectional reliability and measurement equivalence were investigated. A randomized experimental design was implemented, varying category labels and number of categories. The participants were 800 students at two German universities. In contrast to previous research, reliability assessment method was used, which relies on the congeneric measurement model. The experimental manipulation had differential effects …
Studying Latent Criterion Validity for Complex Structure Measuring Instruments Using Latent Variable Modeling
Validity coefficients for multicomponent measuring instruments are known to be affected by measurement error that attenuates them, affects associated standard errors, and influences results of statistical tests with respect to population parameter values. To account for measurement error, a latent variable modeling approach is discussed that allows point and interval estimation of the relationship of an underlying latent factor to a criterion var…
Reliability of Scales With Second-Order Structure
A readily applicable procedure is discussed that allows evaluation of the discrepancy between the popular coefficient alpha and the reliability coefficient of a scale with second-order factorial structure that is frequently of relevance in empirical educational and psychological research. The approach is developed within the framework of the widely used latent variable modeling methodology and permits point and interval estimation of the slippage…
Examining Measurement Invariance and Differential Item Functioning With Discrete Latent Construct Indicators
A latent variable modeling method for studying measurement invariance when evaluating latent constructs with multiple binary or binary scored items with no guessing is outlined. The approach extends the continuous indicator procedure described by Raykov and colleagues, utilizes similarly the false discovery rate approach to multiple testing, and permits one to locate violations of measurement invariance in loading or threshold parameters. The dis…
The Impact of Payment and Respondents' Participation on Interviewers' Accuracy in Face-to-face Surveys
In face-to-face interviews, accurate work by interviewers is crucial for ensuring high-quality survey data. In a field experiment, payment of interviewers, legitimation of falsification behavior, and respondents' willingness to participate were experimentally varied. The impact of these factors on interviewers' accuracy during fieldwork was investigated. Low accuracy was operationalized, for instance, as noncompliance with the instructions concer…
Qualitätssicherung sozialwissenschaftlicher Erhebungsinstrumente
Qualitätssicherung sozialwissenschaftlicher Erhebungsinstrumente
Multiple-Component Measurement Instruments in Heterogeneous Populations
This note confronts the common use of a single coefficient alpha as an index informing about reliability of a multicomponent measurement instrument in a heterogeneous population. Two or more alpha coefficients could instead be meaningfully associated with a given instrument in finite mixture settings, and this may be increasingly more likely the case in empirical educational and psychological research. It is argued that in such situations explici…
Gütekriterien quantitativer Sozialforschung
Questionnaire Design and Translation for Refugee Populations
Surveying the refugee population poses particular challenges: what measurement and culture effects need to be taken into account? Are some of the constructs related to refugees unique or can constructs used in other surveys be adapted? Due to considerable variation in educational background, in trauma history or in perception of ethnicity or gender roles in refugee populations, one needs to raise the question whether a one-size-fits-all approach …
Response Bias and Reliability in Verbal Agreement Rating Scales
Verbal rating scale polarity and verbalizations of the middle category that do not match the polarity in agreement rating scales were investigated. Two randomized web survey experiments were conducted using a probability panel of German Internet users. The classical bipolar “disagree/agree” verbalization was compared with the unipolar “do not agree/agree” alternative. In both experiments, attitudes on gender roles were measured using a different …
How Do Reverse-keyed Items in Inventories Affect Measurement Quality and Information Processing
In randomized experiments, inventories with reverse-keyed items are compared with inventories in which all the items are either positively or negatively associated with the underlying concept. The results show that with reverse keying, a control of the potential bias was not sufficient; likewise, the factorial structure, reliability, and validity were negatively affected. An eye-tracking study revealed that respondents did not process information…
Rating-Scale Labeling in Online Surveys
Unlike other data collection modes, the effect of labeling rating scales on reliability and validity, as relevant aspects of measurement quality, has seldom been addressed in online surveys. In this study, verbal and numeric rating scales were compared in split-ballot online survey experiments. In the first experiment, respondents' cognitive processes were observed by means of eye tracking, that is, determining the respondent's fixations in diffe…
Evaluation of Second- and Third-Level Variance Proportions in Multilevel Designs With Completely Observed Populations
Two- and three-level designs in educational and psychological research can involve entire populations of Level-3 and possibly Level-2 units, such as schools and educational districts nested within a given state, or neighborhoods and counties in a state. Such a design is of increasing relevance in empirical research owing to the growing popularity of large-scale studies in these and cognate disciplines. The present note discusses a readily applica…
On the Relationship Between Item Stem Formulation and Criterion Validity of Multiple-Component Measuring Instruments
The possible dependency of criterion validity on item formulation in a multicomponent measuring instrument is examined. The discussion is concerned with evaluation of the differences in criterion validity between two or more groups (populations/subpopulations) that have been administered instruments with items having differently formulated item stems. The case of complex item stems involving two stimuli description sentences (double-barreled ques…
On the Importance of Coefficient Alpha for Measurement Research
The population relationship between coefficient alpha and scale reliability is studied in the widely used setting of unidimensional multicomponent measuring instruments. It is demonstrated that for any set of component loadings on the common factor, regardless of the extent of their inequality, the discrepancy between alpha and reliability can be arbitrarily small in any considered population and hence practically ignorable. In addition, the set …
Measurement invariance in the social sciences
Mathematics (19 obras) · Psychology (19 obras) · Statistics (19 obras) · Computer Science (18 obras) · Econometrics (10 obras) · Social Psychology (10 obras) · Survey Methodology and Nonresponse (10 obras) · Sociology (9 obras) · Psychometrics (8 obras) · Social and Intergroup Psychology (8 obras)