Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction

Bibliographic Data

ID8399303
AuthorsDeborah G Mayo (0000-0001-8252-9968, Virginia Tech), Aris Spanos (0000-0002-9229-424X, Virginia Tech)
Year2006
Volume57
Issue2
Pages323-357
Publication date2006-06-01
Peer ReviewedYes
Open AccessNo
TypeARTICLE
VenueThe British Journal for the Philosophy of Science (JOURNAL)
Journal identifiersISSN: 0007-0882 • E-ISSN: 1464-3537
PublisherOxford University Press (PUBLISHER • GB)
DOI10.1093/bjps/axl003
OpenAlexW2141673338
LanguageEN
Citations received54
References cited44

Despite the widespread use of key concepts of the Neyman–Pearson (N–P) statistical paradigm—type I and II errors, significance levels, power, confidence levels—they have been the subject of philosophical controversy and debate for over 60 years. Both current and long-standing problems of N–P tests stem from unclarity and confusion, even among N–P adherents, as to how a test's (pre-data) error probabilities are to be used for (post-data) inductive inference as opposed to inductive behavior. We argue that the relevance of error probabilities is to ensure that only statistical hypotheses that have passed severe or probative tests are inferred from the data. The severity criterion supplies a meta-statistical principle for evaluating proposed statistical inferences, avoiding classic fallacies from tests that are overly sensitive, as well as those not sensitive enough to particular errors and discrepancies. 1. Introduction and overview1.1Behavioristic and inferential rationales for Neyman–Pearson (N–P) tests1.2Severity rationale: induction as severe testing1.3Severity as a meta-statistical concept: three required restrictions on the N–P paradigm2. Error statistical tests from the severity perspective2.1N–P test T(α): type I, II error probabilities and power2.2Specifying test T(α) using p-values3. Neyman's post-data use of power3.1Neyman: does failure to reject H warrant confirming H?4. Severe testing as a basic concept for an adequate post-data inference4.1The severity interpretation of acceptance (SIA) for test T(α)4.2The fallacy of acceptance (i.e., an insignificant difference): Ms Rosy4.3Severity and power5. Fallacy of rejection: statistical vs. substantive significance5.1Taking a rejection of H0 as evidence for a substantive claim or theory5.2A statistically significant difference from H0 may fail to indicate a substantively important magnitude5.3Principle for the severity interpretation of a rejection (SIR)5.4Comparing significant results with different sample sizes in T(α): large n problem5.5General testing rules for T(α), using the severe testing concept6. The severe testing concept and confidence intervals6.1Dualities between one and two-sided intervals and tests6.2Avoiding shortcomings of confidence intervals7. Beyond the N–P paradigm: pure significance, and misspecification tests8. Concluding comments: have we shown severity to be a basic concept in a N–P philosophy of induction?

Causal inference · Econometrics · Epistemology · Fallacy · Inductive reasoning · Inference · Relevance (law) · Statistical hypothesis testing · Statistical inference · Statistical power · Statistical significance · Statistical theory · Statistics · Test (biology) · Type I and type II errors · Artificial Intelligence · Bayesian Modeling and Causal Inference · Computer Science · Law · Mathematics · Philosophy · Philosophy and History of Science · Statistical Mechanics and Entropy

  • What Is Wrong with Rerandomization in Randomized Field Experiments

    Open Access•Mariusz Maziarz•Philosophy of Science•2026

  • Philosophy and the Precautionary Principle

    Open Access•Daniel Steel•Philosophy and the Precautionary…•2014

  • Evidence and Experimental Design in Sequential Trials

    Open Access•Jan Sprenger•Philosophy of Science•2009

  • Prediction in Selectionist Evolutionary Theory

    Open Access•Rasmus Grønfeldt Winther, Rasmus Gr⊘nfeldt Winther•Philosophy of Science•2009

  • Empirical Evidence

    Julian Reis, Julian Reiss•The Sage Handbook of the…•2011

  • Statistical Inference as Severe Testing

    Open Access•Deborah G Mayo•Statistical Inference As Severe…•2018

  • Philosophy and the practice of Bayesian statistics

    Open Access•Andrew Gelman, Cosma Rohilla Shalizi•British Journal of Mathematical…•2013

  • Similarity-based interference in sentence comprehension

    Open Access•Lena A Jäger, FELIX ENGELMANN et al.•Journal of Memory and Language•2017

  • The fallacy of placing confidence in confidence intervals

    Open Access•Roy D Morey, Richard D Morey et al.•Psychonomic Bulletin & Review•2016

  • Statistical tests, P values, confidence intervals, and power

    Open Access•Sander Greenland, Stephen Senn et al.•European Journal of Epidemiology•2016

  • On Falsifiable Statistical Hypotheses

    Open Access•Konstantin Genin•Philosophies•2022

  • Significance Tests

    Open Access•Deborah G Mayo•Review of Philosophy and Psychology•2021

  • Neyman-Pearson Hypothesis Testing, Epistemic Reliability and Pragmatic Value-Laden Asymmetric Error Risks

    Open Access•Adam P Kubiak, Paweł Kawalec et al.•Axiomathes•2022

  • Prior Information in Frequentist Research Designs

    Open Access•Adam P Kubiak, Paweł Kawalec•Journal for General Philosophy of…•2022

  • Error Statistics Using the Akaike and Bayesian Information Criteria

    Open Access•Henrique Cheng, H P Cheng et al.•Erkenntnis•2026

  • Inflated effect sizes and underpowered tests

    Open Access•Guillaume Rochefort-Maranda•Philosophical Studies•2021

  • Reflections on the LSE Tradition in Econometrics

    Open Access•Aris Spanos•OEconomia•2014

  • What Foundations for Statistical Modeling and Inference

    Open Access•Aris Spanos•OEconomia•2019

  • Strong-Form Frequentist Testing In Communication Science

    Open Access•Lennert Coenen, Tim Smits•Communication Methods and Measures•2022

  • Transforming structural econometrics

    Aris Spanos•Review of Political Economy•2016

  • Is engaged pluralism the best way ahead for economic geography? Commentary on Barnes and Sheppard (2009)

    Open Access•Dragos Simandan•Progress in Human Geography•2011

  • Tests of Statistical Significance Made Sound

    Open Access•Brian D Haig•Educational and Psychological…•2017

  • On Some Assumptions of the Null Hypothesis Statistical Testing

    Open Access•Alexandre G Patriota•Educational and Psychological…•2017

  • Internalist and externalist aspects of justification in scientific inquiry

    Open Access•Kent W Staley, Kent Staley et al.•Synthese•2011

  • A frequentist interpretation of probability for model-based inductive inference

    Open Access•Aris Spanos•Synthese•2013

  • Conceptual challenges for interpretable machine learning

    Open Access•David S Watson•Synthese•2022

  • Bayesian perspectives on the discovery of the Higgs particle

    Open Access•Richard Dawid•Synthese•2017

  • Model change and reliability in scientific inference

    Open Access•Erich Kummerfeld, David Danks•Synthese•2014

  • Early stopping of RCTs

    Open Access•Roger Stanev•Synthese•2015

  • How experimental algorithmics can benefit from Mayo’s extensions to Neyman–Pearson theory of testing

    Open Access•Thomas Bartz-Beielstein, Thomas Bartz–Beielstein•Synthese•2008

  • Preregistration does not improve the transparent evaluation of severity in Popper’s philosophy of science or when deviations are allowed

    Open Access•M Rubin•Synthese•2025

  • The Jeffreys–Lindley paradox and discovery criteria in high energy physics

    Open Access•Robert D Cousins•Synthese•2017

  • Perspectival realism and frequentist statistics

    Open Access•Adam P Kubiak•Synthese•2024

  • Error statistical modeling and inference

    Open Access•Aris Spanos, Deborah G Mayo•Synthese•2015

  • What type of Type I error? Contrasting the Neyman–Pearson and Fisherian approaches in the context of exact and direct replications

    Open Access•M Rubin•Synthese•2021

  • Bernoulli’s golden theorem in retrospect

    Open Access•Aris Spanos•Synthese•2021

  • Statistical significance and its critics

    Open Access•Deborah G Mayo, David Hand et al.•Synthese•2022

  • Some surprising facts about (the problem of) surprising facts

    D Mayo•Studies in History and Philosophy…•2014

  • Pursuit and inquisitive reasons

    Open Access•Wilfrid Fleisher•Studies in History and Philosophy…•2022

  • Classical versus Bayesian Statistics

    Open Access•Eric Johannesson•Philosophy of Science•2020

  • Philosophical Scrutiny of Evidence of Risks

    Open Access•Deborah G Mayo, Aris Spanos•Philosophy of Science•2006

  • Higgs Discovery and the Look Elsewhere Effect

    Open Access•Richard Dawid•Philosophy of Science•2015

  • The Discovery of Argon

    Open Access•Aris Spanos•Philosophy of Science•2010

  • Who Should Be Afraid of the Jeffreys-Lindley Paradox

    Open Access•Aris Spanos•Philosophy of Science•2013

  • Severity and Trustworthy Evidence

    Open Access•Aris Spanos•Philosophy of Science•2022

  • Is Frequentist Testing Vulnerable to the Base-Rate Fallacy

    Open Access•Aris Spanos•Philosophy of Science•2010

  • Math approach training changes implicit identification with math

    Open Access•Cédric Batailler, Dominique Muller et al.•Journal of Experimental Social…•2020

  • Philosophy in Science

    Thomas Pradeu, Maël Lemoine et al.•The British Journal for the…•2024

  • Severe Testing

    Deborah G Mayo, Deborah Mayo•The British Journal for the…•2025

  • Objectivity and Underdetermination in Statistical Model Selection

    Beckett Sterner, Scott Lidgard•The British Journal for the…•2024

  • How to Discount Double-Counting When It Counts

    Deborah G Mayo•The British Journal for the…•2008

  • Testing What Matters (If You Must Test at All)

    Open Access•Justin H Gro, Justin H Gross•American Journal of Political…•2015

  • The logical status of applied geographical reasoning

    Open Access•Dragos Simandan•Geographical Journal•2012

  • Method against method

    D E Wittkower•Social Identities•2009

  • The Significance Test Controversy

    Denton Morrison, Ramon Henkel•The Significance Test Controversy•2006

  • Bayes or bust?

    John Earman•Bayes or bust?•1992

  • Probability and Evidence

    Open Access•Paul Horwich•Probability and evidence•2016

  • Logic of Statistical Inference

    Open Access•Ian Hacking, Jan-Willem Romeijn•Logic of Statistical Inference•2016

  • Bayesian statistical inference for psychological research.

    Ward Edwards, Harold Lindman et al.•Psychological Review•1963

  • The Logic of Scientific Discovery

    Karl R Popper, Karl Popper et al.•Physics Today•1959

  • The Myth of the Framework

    Karl Popper•Rational Changes in Science•1976

  • An objective theory of statistical testing

    Open Access•Deborah G Mayo•Synthese•1983

  • The Neyman-Pearson theory as decision theory, and as inference theory; with a criticism of the Lindley-savage argument for Bayesian theory

    Open Access•Allan Birnbaum•Synthese•1977

  • Frequentist probability and frequentist statistics

    Open Access•Jerzy Neyman•Synthese•1977

  • Did Pearson reject the Neyman-Pearson philosophy of statistics

    Open Access•Deborah G Mayo•Synthese•1992

  • A note on metastatistics or ?an essay toward stating a problem in the doctrine of chances

    Open Access•Lucien Lecam•Synthese•1977

  • Theory-Testing in Psychology and Physics

    Open Access•Paul E Meehl•Philosophy of Science•1967

  • Novel Evidence and Severe Tests

    Open Access•Deborah G Mayo•Philosophy of Science•1991

  • Methodology in Practice

    Open Access•Deborah G Mayo, Aris Spanos•Philosophy of Science•2004

  • A Logic of Induction

    Open Access•Colin Howson•Philosophy of Science•1997

  • Behavioristic, Evidentialist, and Learning Models of Statistical Testing

    Open Access•Deborah G Mayo•Philosophy of Science•1985

  • Philosophical Problems of Statistical Inference

    Lawrence Sklar, Teddy Seidenfeld•The Philosophical Review•1981

  • The Enterprise of Knowledge

    Mark Kaplan, Isaac Levi•The Philosophical Review•1983

  • An Objective Theory of Probability

    Mark Pastin, Donald Gillies et al.•The Philosophical Review•1975

  • Bayes or Bust? A Critical Examination of Bayesian Confirmation Theory

    Paul Castell, John Earman•The Philosophical Quarterly•1995

  • A Refutation of the Neyman-Pearson Theory of Testing

    Stephen Spielman•The British Journal for the…•1973

  • Theories of Probability

    Colin Howson•The British Journal for the…•1995

  • Probability and evidence

    Paul Horwich•Probability and evidence•2016

  • Error and the growth of experimental knowledge

    Open Access•Ryan D Tweney•Journal of the History of the…•1998

Unique citing works54
Citations per year2,7
Citation span2006 - 2026 (21)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 50

Tools

Open DOISci-Hub
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae