Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction
Bibliographic Data
| ID | 8399303 |
|---|---|
| Authors | Deborah G Mayo (0000-0001-8252-9968, Virginia Tech), Aris Spanos (0000-0002-9229-424X, Virginia Tech) |
| Year | 2006 |
| Volume | 57 |
| Issue | 2 |
| Pages | 323-357 |
| Publication date | 2006-06-01 |
| Peer Reviewed | Yes |
| Open Access | No |
| Type | ARTICLE |
| Venue | The British Journal for the Philosophy of Science (JOURNAL) |
| Journal identifiers | ISSN: 0007-0882 • E-ISSN: 1464-3537 |
| Publisher | Oxford University Press (PUBLISHER • GB) |
| DOI | 10.1093/bjps/axl003 |
| OpenAlex | W2141673338 |
| Language | EN |
| Citations received | 54 |
| References cited | 44 |
Despite the widespread use of key concepts of the Neyman–Pearson (N–P) statistical paradigm—type I and II errors, significance levels, power, confidence levels—they have been the subject of philosophical controversy and debate for over 60 years. Both current and long-standing problems of N–P tests stem from unclarity and confusion, even among N–P adherents, as to how a test's (pre-data) error probabilities are to be used for (post-data) inductive inference as opposed to inductive behavior. We argue that the relevance of error probabilities is to ensure that only statistical hypotheses that have passed severe or probative tests are inferred from the data. The severity criterion supplies a meta-statistical principle for evaluating proposed statistical inferences, avoiding classic fallacies from tests that are overly sensitive, as well as those not sensitive enough to particular errors and discrepancies. 1. Introduction and overview1.1Behavioristic and inferential rationales for Neyman–Pearson (N–P) tests1.2Severity rationale: induction as severe testing1.3Severity as a meta-statistical concept: three required restrictions on the N–P paradigm2. Error statistical tests from the severity perspective2.1N–P test T(α): type I, II error probabilities and power2.2Specifying test T(α) using p-values3. Neyman's post-data use of power3.1Neyman: does failure to reject H warrant confirming H?4. Severe testing as a basic concept for an adequate post-data inference4.1The severity interpretation of acceptance (SIA) for test T(α)4.2The fallacy of acceptance (i.e., an insignificant difference): Ms Rosy4.3Severity and power5. Fallacy of rejection: statistical vs. substantive significance5.1Taking a rejection of H0 as evidence for a substantive claim or theory5.2A statistically significant difference from H0 may fail to indicate a substantively important magnitude5.3Principle for the severity interpretation of a rejection (SIR)5.4Comparing significant results with different sample sizes in T(α): large n problem5.5General testing rules for T(α), using the severe testing concept6. The severe testing concept and confidence intervals6.1Dualities between one and two-sided intervals and tests6.2Avoiding shortcomings of confidence intervals7. Beyond the N–P paradigm: pure significance, and misspecification tests8. Concluding comments: have we shown severity to be a basic concept in a N–P philosophy of induction?
Causal inference · Econometrics · Epistemology · Fallacy · Inductive reasoning · Inference · Relevance (law) · Statistical hypothesis testing · Statistical inference · Statistical power · Statistical significance · Statistical theory · Statistics · Test (biology) · Type I and type II errors · Artificial Intelligence · Bayesian Modeling and Causal Inference · Computer Science · Law · Mathematics · Philosophy · Philosophy and History of Science · Statistical Mechanics and Entropy
What Is Wrong with Rerandomization in Randomized Field Experiments
Philosophy and the Precautionary Principle
Evidence and Experimental Design in Sequential Trials
Prediction in Selectionist Evolutionary Theory
Empirical Evidence
Statistical Inference as Severe Testing
Philosophy and the practice of Bayesian statistics
Similarity-based interference in sentence comprehension
The fallacy of placing confidence in confidence intervals
Statistical tests, P values, confidence intervals, and power
On Falsifiable Statistical Hypotheses
Significance Tests
Neyman-Pearson Hypothesis Testing, Epistemic Reliability and Pragmatic Value-Laden Asymmetric Error Risks
Prior Information in Frequentist Research Designs
Error Statistics Using the Akaike and Bayesian Information Criteria
Inflated effect sizes and underpowered tests
Reflections on the LSE Tradition in Econometrics
What Foundations for Statistical Modeling and Inference
Strong-Form Frequentist Testing In Communication Science
Transforming structural econometrics
Is engaged pluralism the best way ahead for economic geography? Commentary on Barnes and Sheppard (2009)
Tests of Statistical Significance Made Sound
On Some Assumptions of the Null Hypothesis Statistical Testing
Internalist and externalist aspects of justification in scientific inquiry
A frequentist interpretation of probability for model-based inductive inference
Conceptual challenges for interpretable machine learning
Bayesian perspectives on the discovery of the Higgs particle
Model change and reliability in scientific inference
Early stopping of RCTs
How experimental algorithmics can benefit from Mayo’s extensions to Neyman–Pearson theory of testing
Preregistration does not improve the transparent evaluation of severity in Popper’s philosophy of science or when deviations are allowed
The Jeffreys–Lindley paradox and discovery criteria in high energy physics
Perspectival realism and frequentist statistics
Error statistical modeling and inference
What type of Type I error? Contrasting the Neyman–Pearson and Fisherian approaches in the context of exact and direct replications
Bernoulli’s golden theorem in retrospect
Statistical significance and its critics
Some surprising facts about (the problem of) surprising facts
Pursuit and inquisitive reasons
Classical versus Bayesian Statistics
Philosophical Scrutiny of Evidence of Risks
Higgs Discovery and the Look Elsewhere Effect
The Discovery of Argon
Who Should Be Afraid of the Jeffreys-Lindley Paradox
Severity and Trustworthy Evidence
Is Frequentist Testing Vulnerable to the Base-Rate Fallacy
Math approach training changes implicit identification with math
Philosophy in Science
Severe Testing
Objectivity and Underdetermination in Statistical Model Selection
How to Discount Double-Counting When It Counts
Testing What Matters (If You Must Test at All)
The logical status of applied geographical reasoning
Method against method
The Significance Test Controversy
Bayes or bust?
Probability and Evidence
Logic of Statistical Inference
Bayesian statistical inference for psychological research.
The Logic of Scientific Discovery
The Myth of the Framework
An objective theory of statistical testing
The Neyman-Pearson theory as decision theory, and as inference theory; with a criticism of the Lindley-savage argument for Bayesian theory
Frequentist probability and frequentist statistics
Did Pearson reject the Neyman-Pearson philosophy of statistics
A note on metastatistics or ?an essay toward stating a problem in the doctrine of chances
Theory-Testing in Psychology and Physics
Novel Evidence and Severe Tests
Methodology in Practice
A Logic of Induction
Behavioristic, Evidentialist, and Learning Models of Statistical Testing
Philosophical Problems of Statistical Inference
The Enterprise of Knowledge
An Objective Theory of Probability
Bayes or Bust? A Critical Examination of Bayesian Confirmation Theory
A Refutation of the Neyman-Pearson Theory of Testing
Theories of Probability
Probability and evidence
Error and the growth of experimental knowledge
| Unique citing works | 54 |
|---|---|
| Citations per year | 2,7 |
| Citation span | 2006 - 2026 (21) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 50 |