Jessica M. Utts, PhD Sources:

Statistical Significance versus Effect Size in Psi Research

Jessica Utts, Professor Emerita of Statistics at the University of California, Irvine, has argued throughout her career that the field of parapsychology, and science broadly, has systematically misused the p-value as a decision criterion, obscuring the more informative question of effect size. Her contributions reframe psi research not as a question of whether results are “statistically significant” but whether the magnitude of observed effects is consistent, replicable, and comparable to effects accepted in other sciences.

Key findings

  • Utts argued that statistical significance alone is an inadequate criterion for evaluating psi research, because p-values conflate effect size with sample size and do not directly measure the strength of an effect.1
  • She evaluated the ganzfeld and remote-viewing meta-analyses using the same effect-size conventions mainstream behavioral science applies to any cumulative literature — small-to-medium effects by Cohen’s conventions — and found that the parapsychological meta-analyses produced effects in that comparable range.1
  • In a co-authored reply to critics of Bem’s precognition research, Utts and colleagues argued that Bayesian and likelihood-ratio approaches provide more informative evidence summaries than null-hypothesis significance testing.2
  • Utts co-authored a 2019 paper calling for the retirement of the phrase “statistically significant” from scientific reporting, arguing it creates a false binary that harms inference across all empirical sciences.3
  • Her 1999 JSE paper demonstrated that meta-analytic evidence from mind-matter research, when evaluated by the same standards applied to conventional medical research, meets the bar for cumulative scientific evidence.1

Overview

A persistent vulnerability in psi research is the conflation of statistical significance with scientific importance. Because psi effects, if real, are likely small in magnitude, large samples are required to detect them reliably; small studies will frequently fail to reach significance even if the underlying effect is genuine. This Type-II vulnerability means that dismissing psi research on the grounds of non-significant results in underpowered studies is a methodological error, not a scientific conclusion. Utts has consistently made this argument from the standpoint of statistical methodology, applying the same inferential standards she advocates for all empirical sciences.1

Statistical Significance vs. Effect Size, the Core Distinction

A p-value answers the question: “If the null hypothesis were true, how probable is a result at least this extreme?” It does not answer: “How large is the effect?” or “Is this effect scientifically meaningful?” In large samples, trivially small effects produce highly significant p-values; in small samples, practically important effects may not reach significance. Utts’s 1999 JSE paper explicitly draws this distinction, noting that statistical methods are designed to detect and measure relationships in situations where results cannot be identically replicated due to natural variability, and that the appropriate question is whether the cumulative effect size across studies is consistent and meaningful, not whether any individual study crossed a p < .05 threshold.1

The P-Value Problem in Psi Research

Utts has argued that the near-exclusive reliance on null-hypothesis significance testing in psychology and parapsychology creates systematic distortions in how evidence accumulates. A single non-significant result is routinely interpreted as evidence against an effect, when it may simply reflect inadequate statistical power. Conversely, a highly significant result from a large study may reflect a trivially small effect of no practical or theoretical importance. In the context of psi research, where critics often cite failed replications as decisive, this distinction carries particular weight.1

The Bem Replication Debate and Bayesian Alternatives

When Daryl Bem published precognition experiments in 2011, critics argued that the results, while nominally significant, reflected the inadequacy of frequentist significance testing rather than genuine evidence for precognition. Utts, Bem, and Johnson co-authored a reply in the Journal of Personality and Social Psychology arguing that Bayesian and likelihood-ratio approaches, which directly quantify the relative support for competing hypotheses, provide more informative summaries of the evidence than p-values alone.2 The original paper by Bem, Utts, and Johnson (2011) also addressed the question of whether psychologists should change their analytic practices, arguing that the choice of statistical framework materially affects the conclusions drawn from the same data.4

Bayesian Belief vs. Strength of Evidence

Utts, Norris, Suess, and Johnson (2010) examined the relationship between prior beliefs and the interpretation of statistical evidence, asking whether scientists behave as Bayesians in practice even when they use frequentist methods formally. The paper explored how strong prior skepticism about psi can lead researchers to dismiss statistically compelling evidence that would be accepted in other domains, a form of motivated reasoning that the authors argued is inconsistent with the stated norms of scientific inference.5

Effect Size as the Appropriate Standard

Rather than asking whether a psi study achieves p < .05, Utts has consistently asked whether the effect size observed in psi meta-analyses is comparable to effect sizes accepted as meaningful in other sciences. Her position is that the same standard should apply regardless of the domain: if an effect of a given magnitude would be considered scientifically important in medicine or psychology, it should be evaluated by the same criterion in parapsychology.1

The Small-to-Medium Effect-Size Range and Ganzfeld / Remote Viewing Comparisons

In her 1999 JSE paper, Utts compared the effect sizes from ganzfeld and remote viewing meta-analyses against those from conventional medical research, specifically the antiplatelet therapy trials used as a benchmark for cumulative evidence in medicine. She noted that the effect sizes from mind-matter meta-analyses were in a range she described as consistent with the small-to-medium effect-size band by Cohen’s conventions, expressed equivalently as hit-rate elevations above chance for forced-choice paradigms. The comparison was methodological: Utts was not claiming psi effects are as large as drug effects, but that the inferential standard, cumulative meta-analytic evidence with consistent effect direction, should be applied uniformly.1

Model Selection and the Psychokinesis / Precognition / Chance Trichotomy

In a 1996 book chapter on model selection in parapsychology, Utts examined how statistical model-selection criteria, tools designed to choose among competing explanatory models, apply to the question of whether observed anomalous data are better described by psychokinesis, precognition, or chance variation. The chapter appeared in a volume honoring statistician Seymour Geisser and engaged with formal model-comparison frameworks, illustrating that the question of which psi model best fits the data is a tractable statistical problem, not merely a philosophical one.6

Meta-Analysis and Cumulative Evidence

Utts has argued that meta-analysis, the statistical synthesis of results across multiple independent studies, is the appropriate tool for evaluating whether a small but consistent effect exists in psi research. Individual studies, particularly those with small samples, have low power to detect small effects; meta-analysis pools evidence across studies to produce a more stable estimate of the underlying effect size. Her 1999 JSE paper presented two parallel meta-analytic examples, one from conventional medicine and one from mind-matter research, to demonstrate that the inferential logic is identical across domains.1

Antiplatelet Therapy as a Methodological Parallel

The antiplatelet therapy example in Utts’s 1999 JSE paper is methodologically instructive. The medical evidence for antiplatelet drugs reducing vascular disease was accepted by the scientific community on the basis of cumulative meta-analytic evidence showing a consistent, modest effect across multiple randomized controlled trials, not on the basis of any single decisive experiment. Utts argued that the same logic applied to ganzfeld and remote viewing meta-analyses: the question is whether the cumulative evidence, evaluated by standard meta-analytic methods, shows a consistent effect of meaningful magnitude. She concluded that it did, while explicitly noting that this statistical conclusion does not resolve the question of mechanism or causal explanation.1

Modern Context

The statistical reform arguments Utts has advanced in the context of psi research sit within a broader mainstream movement to retire or constrain the use of “statistical significance” as a binary decision criterion. Hurlbert, Levine, and Utts (2019) contributed directly to this mainstream debate, arguing in The American Statistician that the phrase “statistically significant” should be retired from scientific reporting because it encourages a false binary between “significant” and “non-significant” results that obscures effect size, confidence intervals, and the continuous nature of evidence.3 This mainstream statistical reform context is directly relevant to psi research: the same inferential errors that Utts identified in parapsychology, over-reliance on p-values, dismissal of underpowered null results, failure to report effect sizes, are now recognized as pervasive across empirical science.3

Broader Statistical Reform

Utts’s arguments about effect size and statistical significance in psi research are not isolated to parapsychology, they reflect her broader position as a statistician who has consistently advocated for more informative reporting practices across all empirical sciences. Her 2016 presidential address to the American Statistical Association, her textbooks, and her co-authored work on the retirement of “statistical significance” all reflect the same core argument: that p-values are systematically misused as a proxy for scientific importance, and that effect sizes, confidence intervals, and Bayesian evidence measures provide more informative summaries of what data actually show.3

“Coup de Grâce”, Retiring “Statistically Significant”

Hurlbert, Levine, and Utts (2019), published in The American Statistician as part of a special issue on statistical inference, argued that the phrase “statistically significant” should be permanently retired. The paper’s title, “Coup de Grâce for a Tough Old Bull”, signals its polemical intent. The authors argued that the binary significant/non-significant framing: (1) encourages researchers to treat p = .049 and p = .051 as categorically different outcomes; (2) suppresses reporting of effect sizes and confidence intervals; (3) creates incentives for p-hacking and optional stopping; and (4) leads to systematic misinterpretation of null results as evidence of no effect. These are precisely the inferential errors Utts had identified in critiques of psi research, and in defenses of psi research against dismissal based on underpowered null replications.3

Bayesian and Likelihood Frameworks as Alternatives

In the 2011 reply to critics of Bem’s precognition research, Utts and colleagues argued that Bayesian and likelihood-ratio approaches directly address the question that significance testing obscures: not “Is this result unlikely under the null?” but “How much more strongly do the data support the alternative hypothesis than the null?” The reply engaged specifically with critics who had argued that Bem’s results, while nominally significant, reflected the inadequacy of frequentist methods rather than genuine evidence. Utts and colleagues’ position was that applying Bayesian methods to the same data yielded Bayes factors that constituted meaningful evidence, though the magnitude and interpretation of those Bayes factors remained contested in the exchange.2

Skeptical Critiques and Discussion

Critique 1: Bayesian reanalysis of Bem’s data does not support the precognition hypothesis

Skeptic source: Critics of the Bem (2011) precognition studies argued that when Bayesian methods are applied correctly, with appropriately specified prior distributions, the data do not provide strong evidence for precognition, and that the choice of prior materially affects the Bayes factor. This critique was directed at the analytic approach advocated by Utts and colleagues in their reply.4

Response: Utts, Bem, and Johnson replied that the critics’ choice of prior distributions was itself contestable, and that likelihood-ratio approaches, which do not require specifying a prior, also yielded evidence favoring the alternative hypothesis over the null. They argued that the debate over priors, while legitimate, does not resolve in favor of the null hypothesis and that the critics had not demonstrated that the data were better described by chance than by a small consistent effect.2

Analysis. Bem, Utts, & Johnson (2011) reply to Bayesian-reanalysis critique with a counter-reanalysis (R002 + R004). Independent replication of the disputed parameter choices has not been published.

Critique 2: Applying conventional effect-size thresholds to psi research ignores the extraordinary-evidence standard

Skeptic source: A recurring skeptical argument is that Utts’s application of standard effect-size conventions (such as Cohen’s small-to-medium range) to psi research is inappropriate because extraordinary claims require extraordinary evidence, a higher evidential bar than is applied to conventional scientific hypotheses. On this view, the same meta-analytic effect size that would be accepted as evidence for a drug effect should not be accepted as evidence for anomalous cognition, because the prior probability of psi is far lower.5

Response: Utts and colleagues addressed this argument directly in their 2010 paper on the strength of evidence versus the power of belief, arguing that the “extraordinary evidence” standard, while intuitively appealing, is not operationalized in a way that is consistent with standard Bayesian inference. They noted that very strong prior skepticism can make it effectively impossible for any finite dataset to shift belief, a property that is epistemically problematic regardless of the domain. The appropriate response to low prior probability, they argued, is to specify that prior formally and update it by the likelihood ratio, not to apply an informal and unquantified higher threshold.5

Analysis. The debate about how prior probability should be incorporated into the evaluation of psi evidence is a genuine methodological dispute. Utts’s position is internally consistent with standard Bayesian inference; the skeptical position reflects a legitimate concern about the asymmetry between the costs of Type-I and Type-II errors when evaluating extraordinary claims. Neither position has been definitively resolved in the literature covered by this pool.

References
  1. Utts, J. (1999). The Significance of Statistics in Mind-Matter Research. Journal of Scientific Exploration, 13(4), 615–638. R001 [Utts 1999] ↩︎
  2. Bem, D. J., Utts, J., & Johnson, W. O. (2011). Reply: Must psychologists change the way they analyze their data? Journal of Personality and Social Psychology, 101, 716–719. R002 [Bem 2011] ↩︎
  3. Hurlbert, S. H., Levine, R. A., & Utts, J. (2019). Coup de Grâce for a Tough Old Bull: “Statistically Significant” Expires. The American Statistician, 73(sup1), 352–357. https://doi.org/10.1080/00031305.2018.1543616 R003 [Hurlbert 2019] ↩︎
  4. Bem, D. J., Utts, J., & Johnson, W. O. (2011). Must psychologists change the way they analyze their data? Journal of Personality and Social Psychology, 101(4), 716-719. https://doi.org/10.1037/a0024777 R004 [Bem 2011] ↩︎
  5. Utts, J., Norris, M., Suess, E., & Johnson, W. (2010). The strength of evidence versus the power of belief: Are we all Bayesians? In C. Reading (Ed.), Data and context in statistics education: Towards an evidence-based society. Proceedings of the Eighth International Conference on Teaching Statistics (ICOTS-8), Ljubljana, Slovenia. International Statistical Institute. https://iase-web.org/documents/papers/icots8/ICOTS8_8H1_UTTS.pdf R005 [Utts 2010] ↩︎
  6. Utts, J. (1996). Model selection in parapsychology: Psychokinesis, precognition or chance? In J. C. Lee, W. O. Johnson, & A. Zellner (Eds.), Modelling and prediction honoring Seymour Geisser (pp. 315–322). Springer New York. https://doi.org/10.1007/978-1-4612-2414-3_20 R006 [Utts 1996] ↩︎

See hub COI disclosure for subject-coauthorship transparency.