Deborah G Mayo
Datos Biográficos
| ID | 768271 |
|---|---|
| NOMBRE | Deborah G Mayo |
| NOMBRES | Deborah G |
| APELLIDO | Mayo |
| FIRMA | MAYO D G |
| AFILIACIONES | Virginia Tech |
| ORCID | 0000-0001-8252-9968 |
| VERIFICADO | Sí |
| TOTAL DE OBRAS | 35 |
| TOTAL DE CITAS | 174 |
| TOTAL COMO AUTOR | 34 |
| TOTAL COMO EDITOR | 1 |
| PRIMER AÑO DE PUBLICACIÓN | 1981 |
| AÑO MÁS RECIENTE DE PUBLICACIÓN | 2025 |
| ÍNDICE H | 7 |
Severe Testing
Introduction to recent issues in philosophy of statistics
Statistical significance and its critics
While the common procedure of statistical significance testing and its accompanying concept of p-values have long been surrounded by controversy, renewed concern has been triggered by the replication crisis in science. Many blame statistical significance tests themselves, and some regard them as sufficiently damaging to scientific practice as to warrant being abandoned. We take a contrary position, arguing that the central criticisms arise from m…
Significance Tests
Five ways to ensure that models serve society
Statistical Inference as Severe Testing
Mounting failures of replication in social and biological sciences give a new urgency to critically appraising proposed reforms. This book pulls back the cover on disagreements between experts charged with restoring integrity to science. It denies two pervasive views of the role of probability in inference: to assign degrees of belief, and to control error rates in a long run. If statistical consumers are unaware of assumptions behind rival evide…
Justify your alpha
Ontology & methodology
Error statistical modeling and inference
The error statistical philosopher as normative naturalist
Some Methodological Issues in Experimental Economics
The growing acceptance and success of experimental economics has increased the interest of researchers in tackling philosophical and methodological challenges to which their work increasingly gives rise. I sketch some general issues that call for the combined expertise of experimental economists and philosophers of science, of experiment, and of inductive-statistical inference and modeling
How to Discount Double-Counting When It Counts
The issues of double-counting, use-constructing, and selection effects have long been the subject of debate in the philosophical as well as statistical literature. I have argued that it is the severity, stringency, or probativeness of the test—or lack of it—that should determine if a double-use of data is admissible. Hitchcock and Sober ([2004]) question whether this ‘severity criterion' can perform its intended job. I argue that their criticisms…
Philosophical Scrutiny of Evidence of Risks
We argue that a responsible analysis of today's evidence-based risk assessments and risk debates in biology demands a critical or metascientific scrutiny of the uncertainties, assumptions, and threats of error along the manifold steps in risk analysis. Without an accompanying methodological critique, neither sensitivity to social and ethical values, nor conceptual clarification alone, suffices. In this view, restricting the invitation for philoso…
Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction
Despite the widespread use of key concepts of the Neyman–Pearson (N–P) statistical paradigm—type I and II errors, significance levels, power, confidence levels—they have been the subject of philosophical controversy and debate for over 60 years. Both current and long-standing problems of N–P tests stem from unclarity and confusion, even among N–P adherents, as to how a test's (pre-data) error probabilities are to be used for (post-data) inductive…
Methodology in Practice
The growing availability of computer power and statistical software has greatly increased the ease with which practitioners apply statistical methods, but this has not been accompanied by attention to checking the assumptions on which these methods are based. At the same time, disagreements about inferences based on statistical research frequently revolve around whether the assumptions are actually met in the studies available, e.g., in psycholog…
Novel work on problems of novelty? Comments on Hudson
What is this thing called philosophy of science
Experimental Practice and an Error Statistical Account of Evidence
In seeking general accounts of evidence, confirmation, or inference, philosophers have looked to logical relationships between evidence and hypotheses. Such logics of evidential relationship , whether hypothetico-deductive, Bayesian, or instantiationist fail to capture or be relevant to scientific practice. They require information that scientists do not generally have (e.g., an exhaustive set of hypotheses), while lacking slots within which to i…
Error and the Growth of Experimental Knowledge
Error Statistics and Learning From Error
The error statistical account of testing uses statistical considerations, not to provide a measure of probability of hypotheses, but to model patterns of irregularity that are useful for controlling, distinguishing, and learning from errors. The aim of this paper is (1) to explain the main points of contrast between the error statistical and the subjective Bayesian approach and (2) to elucidate the key errors that underlie the central objection r…
Duhem's Problem, the Bayesian Way, and Error Statistics, or “What's Belief Got to Do with It?”
I argue that the Bayesian Way of reconstructing Duhem's problem fails to advance a solution to the problem of which of a group of hypotheses ought to be rejected or “blamed” when experiment disagrees with prediction. But scientists do regularly tackle and often enough solve Duhemian problems. When they do, they employ a logic and methodology which may be called error statistics. I discuss the key properties of this approach which enable it to spl…
Response to Howson and Laudan
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content
Error and the Growth of Experimental Knowledge
We may learn from our mistakes, but this work argues that, where experimental knowledge is concerned, we haven't begun to learn enough. It provides a critique of the subjective Bayesian view of statistical inference, and proposes the author's own error-statistical approach as a more robust framework for the epistemology of experiment. Deborah Mayo seeks to address the needs of researchers who work with statistical analysis, and simultaneously eng…
Ducks, Rabbits, and Normal Science
Kuhn maintains that what marks the transition to a science is the ability to carry out ‘normal’ science—a practice he characterizes as abandoning the kind of testing that Popper lauds as the hallmark of science. Examining Kuhn's own contrast with Popper, I propose to recast Kuhnian normal science. Thus recast, it is seen to consist of severe and reliable tests of low-level experimental hypotheses (normal tests) and is, indeed, the place to look t…
Acceptable Evidence
Discussions of science and values in risk management have largely focused on how values enter into arguments about risks, that is, issues of acceptable risk. Instead this volume concentrates on how values enter into collecting, interpreting, communicating, and evaluating the evidence of risks, that is, issues of the acceptability of evidence of risk. By focusing on acceptable evidence, this volume avoids two barriers to progress. One barrier assu…
Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction
Despite the widespread use of key concepts of the Neyman–Pearson (N–P) statistical paradigm—type I and II errors, significance levels, power, confidence levels—they have been the subject of philosophical controversy and debate for over 60 years. Both current and long-standing problems of N–P tests stem from unclarity and confusion, even among N–P adherents, as to how a test's (pre-data) error probabilities are to be used for (post-data) inductive…
Novel Evidence and Severe Tests
While many philosophers of science have accorded special evidential significance to tests whose results are “novel facts”, there continues to be disagreement over both the definition of novelty and why it should matter. The view of novelty favored by Giere, Lakatos, Worrall and many others is that of use-novelty : An accordance between evidence e and hypothesis h provides a genuine test of h only if e is not used in h 's construction. I argue tha…
Justify your alpha
Methodology in Practice
The growing availability of computer power and statistical software has greatly increased the ease with which practitioners apply statistical methods, but this has not been accompanied by attention to checking the assumptions on which these methods are based. At the same time, disagreements about inferences based on statistical research frequently revolve around whether the assumptions are actually met in the studies available, e.g., in psycholog…
Duhem's Problem, the Bayesian Way, and Error Statistics, or “What's Belief Got to Do with It?”
I argue that the Bayesian Way of reconstructing Duhem's problem fails to advance a solution to the problem of which of a group of hypotheses ought to be rejected or “blamed” when experiment disagrees with prediction. But scientists do regularly tackle and often enough solve Duhemian problems. When they do, they employ a logic and methodology which may be called error statistics. I discuss the key properties of this approach which enable it to spl…
Experimental Practice and an Error Statistical Account of Evidence
In seeking general accounts of evidence, confirmation, or inference, philosophers have looked to logical relationships between evidence and hypotheses. Such logics of evidential relationship , whether hypothetico-deductive, Bayesian, or instantiationist fail to capture or be relevant to scientific practice. They require information that scientists do not generally have (e.g., an exhaustive set of hypotheses), while lacking slots within which to i…
How to Discount Double-Counting When It Counts
The issues of double-counting, use-constructing, and selection effects have long been the subject of debate in the philosophical as well as statistical literature. I have argued that it is the severity, stringency, or probativeness of the test—or lack of it—that should determine if a double-use of data is admissible. Hitchcock and Sober ([2004]) question whether this ‘severity criterion' can perform its intended job. I argue that their criticisms…
Behavioristic, Evidentialist, and Learning Models of Statistical Testing
While orthodox (Neyman-Pearson) statistical tests enjoy widespread use in science, the philosophical controversy over their appropriateness for obtaining scientific knowledge remains unresolved. I shall suggest an explanation and a resolution of this controversy. The source of the controversy, I argue, is that orthodox tests are typically interpreted as rules for making optimal decisions as to how to behave –-where optimality is measured by the f…
Error statistical modeling and inference
Error Statistics and Learning From Error
The error statistical account of testing uses statistical considerations, not to provide a measure of probability of hypotheses, but to model patterns of irregularity that are useful for controlling, distinguishing, and learning from errors. The aim of this paper is (1) to explain the main points of contrast between the error statistical and the subjective Bayesian approach and (2) to elucidate the key errors that underlie the central objection r…
Response to Howson and Laudan
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content
Did Pearson reject the Neyman-Pearson philosophy of statistics
An objective theory of statistical testing
In Defense of the Neyman-Pearson Theory of Confidence Intervals
In Philosophical Problems of Statistical Inference , Seidenfeld argues that the Neyman-Pearson (NP) theory of confidence intervals is inadequate for a theory of inductive inference because, for a given situation, the ‘best’ NP confidence interval, [CI λ ], sometimes yields intervals which are trivial (i.e., tautologous). I argue that (1) Seidenfeld's criticism of trivial intervals is based upon illegitimately interpreting confidence levels as mea…
Ducks, Rabbits, and Normal Science
Kuhn maintains that what marks the transition to a science is the ability to carry out ‘normal’ science—a practice he characterizes as abandoning the kind of testing that Popper lauds as the hallmark of science. Examining Kuhn's own contrast with Popper, I propose to recast Kuhnian normal science. Thus recast, it is seen to consist of severe and reliable tests of low-level experimental hypotheses (normal tests) and is, indeed, the place to look t…
Models of Group Selection
The key problem in the controversy over group selection is that of defining a criterion of group selection that identifies a distinct causal process that is irreducible to the causal process of individual selection. We aim to clarify this problem and to formulate an adequate model of irreducible group selection. We distinguish two types of group selection models, labeling them type I and type II models. Type I models are invoked to explain differ…
Statistical significance and its critics
While the common procedure of statistical significance testing and its accompanying concept of p-values have long been surrounded by controversy, renewed concern has been triggered by the replication crisis in science. Many blame statistical significance tests themselves, and some regard them as sufficiently damaging to scientific practice as to warrant being abandoned. We take a contrary position, arguing that the central criticisms arise from m…
Philosophical Scrutiny of Evidence of Risks
We argue that a responsible analysis of today's evidence-based risk assessments and risk debates in biology demands a critical or metascientific scrutiny of the uncertainties, assumptions, and threats of error along the manifold steps in risk analysis. Without an accompanying methodological critique, neither sensitivity to social and ethical values, nor conceptual clarification alone, suffices. In this view, restricting the invitation for philoso…
Increasing Public Participation in Controversies Involving Hazards
Despite increased public concern over the social consequences of policies regarding hazardous substances and practices (e.g., nuclear technology, toxic wastes, carcinogenic substances), there has not been adequate public representation in the controversial decisions upon which these policies are based. The problem of inadequate public participation in controversies is therefore often raised in interdisciplinary studies of science, technology, and…
The error statistical philosopher as normative naturalist
In Defense of the Neyman-Pearson Theory of Confidence Intervals
In Philosophical Problems of Statistical Inference , Seidenfeld argues that the Neyman-Pearson (NP) theory of confidence intervals is inadequate for a theory of inductive inference because, for a given situation, the ‘best’ NP confidence interval, [CI λ ], sometimes yields intervals which are trivial (i.e., tautologous). I argue that (1) Seidenfeld's criticism of trivial intervals is based upon illegitimately interpreting confidence levels as mea…
An objective theory of statistical testing
Behavioristic, Evidentialist, and Learning Models of Statistical Testing
While orthodox (Neyman-Pearson) statistical tests enjoy widespread use in science, the philosophical controversy over their appropriateness for obtaining scientific knowledge remains unresolved. I shall suggest an explanation and a resolution of this controversy. The source of the controversy, I argue, is that orthodox tests are typically interpreted as rules for making optimal decisions as to how to behave –-where optimality is measured by the f…
Increasing Public Participation in Controversies Involving Hazards
Despite increased public concern over the social consequences of policies regarding hazardous substances and practices (e.g., nuclear technology, toxic wastes, carcinogenic substances), there has not been adequate public representation in the controversial decisions upon which these policies are based. The problem of inadequate public participation in controversies is therefore often raised in interdisciplinary studies of science, technology, and…
How Everyone Can Have a Rare Property
In a recent discussion note Sober (1985) elaborates on the argument given in Sober (1982) to show the inadequacy of Ronald Giere's (1979, 1980) causal model for cases of frequency-dependent causation , and denies that Giere's (1984) response avoids the problem he raises. I argue that frequency-dependent effects do not pose a problem for Giere's original causal model, and that all parties in this dispute have been guity of misinterpreting the coun…
Models of Group Selection
The key problem in the controversy over group selection is that of defining a criterion of group selection that identifies a distinct causal process that is irreducible to the causal process of individual selection. We aim to clarify this problem and to formulate an adequate model of irreducible group selection. We distinguish two types of group selection models, labeling them type I and type II models. Type I models are invoked to explain differ…
Novel Evidence and Severe Tests
While many philosophers of science have accorded special evidential significance to tests whose results are “novel facts”, there continues to be disagreement over both the definition of novelty and why it should matter. The view of novelty favored by Giere, Lakatos, Worrall and many others is that of use-novelty : An accordance between evidence e and hypothesis h provides a genuine test of h only if e is not used in h 's construction. I argue tha…
Scientific Reasoning
Did Pearson reject the Neyman-Pearson philosophy of statistics
Acceptable Evidence
Discussions of science and values in risk management have largely focused on how values enter into arguments about risks, that is, issues of acceptable risk. Instead this volume concentrates on how values enter into collecting, interpreting, communicating, and evaluating the evidence of risks, that is, issues of the acceptability of evidence of risk. By focusing on acceptable evidence, this volume avoids two barriers to progress. One barrier assu…
Acceptable Evidence
Discussions of science and values in risk management have largely focussed on the entry of values in judging risks, that is, issues of acceptable risk. This volume instead concentrates on the entry of values in collecting, interpreting, communicating, and evaluating the evidence of risks, that is, issues of the acceptability of evidence of risk. By focusing on acceptable evidence, this volume avoids two barriers to progress: views that assume tha…
Error and the Growth of Experimental Knowledge
We may learn from our mistakes, but this work argues that, where experimental knowledge is concerned, we haven't begun to learn enough. It provides a critique of the subjective Bayesian view of statistical inference, and proposes the author's own error-statistical approach as a more robust framework for the epistemology of experiment. Deborah Mayo seeks to address the needs of researchers who work with statistical analysis, and simultaneously eng…
Ducks, Rabbits, and Normal Science
Kuhn maintains that what marks the transition to a science is the ability to carry out ‘normal’ science—a practice he characterizes as abandoning the kind of testing that Popper lauds as the hallmark of science. Examining Kuhn's own contrast with Popper, I propose to recast Kuhnian normal science. Thus recast, it is seen to consist of severe and reliable tests of low-level experimental hypotheses (normal tests) and is, indeed, the place to look t…
Error Statistics and Learning From Error
The error statistical account of testing uses statistical considerations, not to provide a measure of probability of hypotheses, but to model patterns of irregularity that are useful for controlling, distinguishing, and learning from errors. The aim of this paper is (1) to explain the main points of contrast between the error statistical and the subjective Bayesian approach and (2) to elucidate the key errors that underlie the central objection r…
Duhem's Problem, the Bayesian Way, and Error Statistics, or “What's Belief Got to Do with It?”
I argue that the Bayesian Way of reconstructing Duhem's problem fails to advance a solution to the problem of which of a group of hypotheses ought to be rejected or “blamed” when experiment disagrees with prediction. But scientists do regularly tackle and often enough solve Duhemian problems. When they do, they employ a logic and methodology which may be called error statistics. I discuss the key properties of this approach which enable it to spl…
Response to Howson and Laudan
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content
Error and the Growth of Experimental Knowledge
What is this thing called philosophy of science
Experimental Practice and an Error Statistical Account of Evidence
In seeking general accounts of evidence, confirmation, or inference, philosophers have looked to logical relationships between evidence and hypotheses. Such logics of evidential relationship , whether hypothetico-deductive, Bayesian, or instantiationist fail to capture or be relevant to scientific practice. They require information that scientists do not generally have (e.g., an exhaustive set of hypotheses), while lacking slots within which to i…
Novel work on problems of novelty? Comments on Hudson
Methodology in Practice
The growing availability of computer power and statistical software has greatly increased the ease with which practitioners apply statistical methods, but this has not been accompanied by attention to checking the assumptions on which these methods are based. At the same time, disagreements about inferences based on statistical research frequently revolve around whether the assumptions are actually met in the studies available, e.g., in psycholog…
Philosophical Scrutiny of Evidence of Risks
We argue that a responsible analysis of today's evidence-based risk assessments and risk debates in biology demands a critical or metascientific scrutiny of the uncertainties, assumptions, and threats of error along the manifold steps in risk analysis. Without an accompanying methodological critique, neither sensitivity to social and ethical values, nor conceptual clarification alone, suffices. In this view, restricting the invitation for philoso…
Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction
Despite the widespread use of key concepts of the Neyman–Pearson (N–P) statistical paradigm—type I and II errors, significance levels, power, confidence levels—they have been the subject of philosophical controversy and debate for over 60 years. Both current and long-standing problems of N–P tests stem from unclarity and confusion, even among N–P adherents, as to how a test's (pre-data) error probabilities are to be used for (post-data) inductive…
The error statistical philosopher as normative naturalist
Some Methodological Issues in Experimental Economics
The growing acceptance and success of experimental economics has increased the interest of researchers in tackling philosophical and methodological challenges to which their work increasingly gives rise. I sketch some general issues that call for the combined expertise of experimental economists and philosophers of science, of experiment, and of inductive-statistical inference and modeling
Philosophy and History of Science (26 obras) · Computer Science (24 obras) · Epistemology (24 obras) · Philosophy (24 obras) · Mathematics (20 obras) · Statistics (20 obras) · Psychology (14 obras) · Philosophy of science (13 obras) · Statistical hypothesis testing (13 obras) · Artificial Intelligence (12 obras)