Jessica M. Utts, PhD Sources:

Replication Standards and Meta-Analysis in Parapsychology

Jessica Utts is a statistician whose 1991 paper in Statistical Science applied mainstream methodological standards, effect size estimation, statistical power analysis, and meta-analytic synthesis, to the parapsychological literature, arguing that the field’s replication controversies were largely artifacts of misapplied significance testing rather than genuine failures of experimental reproducibility. Her work does not advocate for psi as a believer; it evaluates the data using the same criteria she applies across all empirical sciences.

Key findings

  • Utts argued that defining successful replication as achieving p ≤ .05 is statistically incoherent when studies are underpowered, because a null result from an underpowered study cannot distinguish a true null from a small real effect that the study lacked the power to detect.1
  • Her 1991 Statistical Science review concluded that meta-analyses across multiple parapsychological domains, including ganzfeld, remote viewing, and random number generator studies, collectively indicated an anomalous effect in need of explanation, applying effect-size thresholds in the small-to-medium range by Cohen’s conventions, the same standards mainstream behavioral science uses for evaluating cumulative meta-analytic evidence.2
  • Utts identified vote-counting (tallying significant vs. non-significant studies) as a methodologically flawed approach to assessing replication, and advocated for effect-size estimation and formal meta-analysis as replacements.2
  • She distinguished between the data claim (anomalous effect sizes are present) and the mechanism claim (what causes them), explicitly stating that the data warranted investigation regardless of whether a psi mechanism was ultimately confirmed.2
  • Collaborative work with parapsychologists emphasized that demonstration research and meta-analysis serve different scientific functions, and that conflating them distorts evaluation of the evidence base.3

Overview

A persistent vulnerability in evaluating parapsychological research is the low statistical power of individual experiments: when the true effect size is small (as is typical in psi research), most single studies will fail to reach p ≤ .05 even if the effect is real, making apparent non-replication an artifact of sample size rather than evidence against the phenomenon. Utts addressed this Type-II vulnerability directly, arguing that the field’s critics and defenders alike had been misreading the literature by treating significance thresholds as the unit of replication rather than effect sizes.1 Her 1991 paper in Statistical Science, a peer-reviewed mainstream statistics journal, brought this argument to a broad methodological audience and invited formal published commentary from statisticians and critics.2

Statistical Power and the Replication Illusion

Utts’s 1988 Journal of Parapsychology paper laid the groundwork by tracing the historical entrenchment of p ≤ .05 as the definition of a “successful” experiment in parapsychology and psychology alike.1 She demonstrated that this convention, rooted in historical practice rather than rational design, systematically disadvantages small-effect literatures: a study with N=50 detecting a d=0.3 effect has roughly 40% power, meaning 60% of exact replications will be “non-significant” even if the effect is perfectly real. The 1991 Statistical Science paper extended this argument with a survey of parapsychological meta-analyses, noting that the National Research Council’s 1988 review had explicitly ignored non-significant studies in its tally, a vote-counting error that Utts identified as methodologically indefensible.2 She proposed that parapsychologists calculate power prior to running experiments, use confidence intervals and effect-size estimation alongside hypothesis tests, and consider Bayesian approaches as supplements to frequentist inference.1

The Replication Problem Reframed

Utts’s central methodological contribution was to reframe what counts as replication failure. Rather than asking “did this study reach p ≤ .05?”, she argued the correct question is “is the effect size consistent across studies, and is the aggregate evidence sufficient to rule out chance?”2 This reframing has direct implications for how critics and proponents should read a literature in which individual studies are typically underpowered.

Vote-Counting vs. Effect-Size Synthesis

Vote-counting, tallying the number of studies that reach significance versus those that do not, is statistically known to be a biased and inconsistent estimator of the true effect, becoming more misleading as the number of studies increases when the true effect is small. Utts documented this problem explicitly in the context of parapsychology, noting that critics who cited a low percentage of “significant” studies as evidence against psi were applying a method that would also appear to disconfirm many well-established small-effect phenomena in medicine and psychology.2 The alternative she advocated, formal meta-analysis computing a weighted mean effect size with confidence intervals, had by 1991 been applied to several parapsychological domains, and she summarized those results in the Statistical Science paper. She also cited the broader methodological casebook literature on meta-analysis as providing the framework for this approach.4

The d ≈ 0.4 Benchmark and Its Application

Utts applied a consistent effect-size benchmark across sciences: she treated effect sizes in the range of d ≈ 0.2–0.4 as small but scientifically meaningful when replicated across independent laboratories, consistent with standards applied in medical and psychological research. Her 1991 review found that several parapsychological meta-analyses reported effect sizes in this range, leading her to conclude that the aggregate evidence indicated an anomalous effect warranting explanation.2 She was explicit that this conclusion was about the data, not about the mechanism: the proposed explanation (psi, broadly construed) remained contested, but the data pattern itself met the same evidentiary threshold she would apply to any other small-effect scientific domain. Her rejoinder to published commentaries on the Statistical Science paper reinforced this data/mechanism distinction in response to critics who conflated her statistical conclusions with metaphysical endorsement.5

Meta-Analytic Survey of Parapsychology

The 1991 Statistical Science paper (published with invited discussion and a rejoinder, making it a formal peer-reviewed exchange) surveyed meta-analyses across ganzfeld ESP, remote viewing, random number generator experiments, and other paradigms, concluding that the cumulative evidence across domains showed consistent small positive effects that could not be attributed to chance alone.6 The paper also addressed the file-drawer problem, the concern that unpublished null studies inflate apparent effect sizes, by examining fail-safe N calculations and noting that the number of unpublished null studies required to nullify the observed effects was implausibly large in several domains.

File-Drawer Problem: Artifact Addressed but Not Eliminated

The file-drawer problem (publication bias: unpublished null results inflate observed effect sizes in meta-analyses) is a specific artifact Utts addressed in her 1991 review. She applied fail-safe N analysis, calculating how many null studies would need to exist in file drawers to reduce the observed aggregate effect to chance, and found that for several parapsychological domains the required number was implausibly large relative to the known research output of the field.2 This partially addressed but did not eliminate the publication-bias concern: fail-safe N is a conservative and widely criticized metric, and Utts acknowledged that more sophisticated bias-detection methods (such as funnel-plot regression) were not uniformly applied across the meta-analyses she reviewed. The artifact was therefore mitigated but not fully ruled out at the time of the 1991 paper.

Scope of the Meta-Analytic Survey

Utts’s 1991 survey covered meta-analyses from multiple parapsychological paradigms available at the time of writing, including ganzfeld telepathy studies, remote viewing experiments, and random number generator (RNG) mind-matter interaction studies. For each domain she reported the aggregate effect size, the number of studies included, and the fail-safe N. She noted that the ganzfeld literature had been the subject of an extended methodological debate, including a specific exchange between Hyman and Honorton, and that a new series of autoganzfeld experiments had been designed specifically to address the methodological criticisms raised in that debate. The autoganzfeld results, available by 1991, showed effect sizes consistent with the earlier ganzfeld meta-analysis, which Utts treated as meaningful convergent evidence given that the new studies had addressed the specific artifact concerns (sensory leakage via physical isolation of sender and receiver; randomization via computer-controlled target selection) raised by critics.2

Demonstration Research and Methodological Standards

In collaborative work with parapsychologists including Krippner, Braud, Palmer, Rao, Schlitz, and others, Utts contributed to a discussion of the distinction between demonstration research (designed to show an effect exists) and hypothesis-testing research (designed to test specific theoretical claims about the effect).3 She argued that conflating these two research modes leads to methodological confusion: demonstration studies are appropriately evaluated by whether they produce replicable effect sizes, while hypothesis-testing studies require pre-specified predictions and stricter controls against multiple comparisons and optional stopping.

Distinguishing Demonstration from Hypothesis-Testing Research

The 1993 Journal of Parapsychology collaborative paper, co-authored by Utts alongside Krippner, Braud, Child, Palmer, Rao, Schlitz, and White, addressed the role of meta-analysis in synthesizing demonstration research.3 The paper argued that parapsychology had accumulated a substantial body of demonstration-level evidence, consistent small effects across independent laboratories, but had not yet moved systematically to the hypothesis-testing phase in which specific mechanistic predictions are pre-specified and tested. This framing positioned meta-analysis as the appropriate tool for the demonstration phase, while calling for more rigorous pre-specification in future research. The paper also noted that multiple-comparisons inflation (testing many outcome variables without correction) and optional stopping (continuing data collection until significance is reached) were specific artifacts that hypothesis-testing designs needed to address explicitly, and that these artifacts were partially but not uniformly controlled in the existing demonstration literature.

Rejoinder to Statistical Science Critics

The published discussion of Utts’s 1991 Statistical Science paper included commentary from multiple statisticians and critics. Her rejoinder addressed several specific objections: that the meta-analyses she surveyed were themselves methodologically heterogeneous (she acknowledged this but argued heterogeneity does not nullify aggregate evidence); that the file-drawer problem was more severe than her fail-safe N calculations suggested (she acknowledged the limitation of fail-safe N as a metric); and that her conclusion about an anomalous effect implied endorsement of a psi mechanism (she explicitly rejected this conflation, reiterating that the data claim and the mechanism claim are separable).7 The rejoinder also noted that critics who demanded a complete mechanistic explanation before accepting the data were applying a standard not uniformly applied in other sciences, aspirin’s mechanism was not fully understood when its clinical efficacy was accepted.

Autobiographical Context: Entry into Parapsychology Statistics

In a 2022 reflective paper in the Journal of Anomalistics, Utts described how she came to work with the parapsychological community as a statistician, noting that the welcoming nature of the community was a significant factor in her continued engagement.8 She held visiting positions at Stanford and SRI International in addition to her tenured positions at UC Davis and UC Irvine, which placed her in proximity to the remote viewing research programs of the 1980s and 1990s. Her engagement with parapsychology was consistently framed as a statistician applying her expertise to a contested empirical domain, not as a researcher with prior belief in psi phenomena.

Modern Context

The meta-analytic standards Utts has applied to parapsychological literatures — formal effect-size estimation, attention to between-study heterogeneity, and explicit discussion of publication bias — align with mainstream guidance now codified across medicine and behavioral science. Borenstein, Hedges, Higgins, and Rothstein’s Introduction to Meta-Analysis (Wiley, 2009) provides the canonical treatment of fixed- vs. random-effects models, heterogeneity statistics, and bias diagnostics. Sterne et al. (2011, BMJ; 10.1136/bmj.d4002) issued the standard recommendations for examining and interpreting funnel-plot asymmetry, refining the funnel-plot diagnostic introduced by Egger et al. (1997, BMJ; 10.1136/bmj.315.7109.629). The Open Science Collaboration’s 2015 reproducibility project (10.1126/science.aac4716) brought wide attention to the replication problem in behavioral science that Utts had been arguing for decades was a general statistical-practice issue, not unique to controversial domains.

Skeptical Critiques and Discussion

Critique 1: The meta-analyses Utts surveyed were methodologically heterogeneous and their aggregation was therefore invalid

Skeptic source: Commentators in the published discussion of the 1991 Statistical Science paper argued that the parapsychological meta-analyses Utts reviewed varied substantially in methodology, target type, subject selection, and outcome measurement, making aggregate effect-size synthesis misleading, a concern about heterogeneity as an artifact inflating or distorting apparent effects.6

Response: Utts acknowledged the heterogeneity concern directly in her rejoinder but argued that methodological heterogeneity is a feature of virtually all meta-analytic literatures in psychology and medicine, and that its presence does not invalidate aggregate synthesis, it instead calls for moderator analysis to identify which study features predict larger or smaller effects. She noted that the heterogeneity argument, applied consistently, would also invalidate many accepted meta-analyses in mainstream psychology and medicine.7 The 1993 collaborative paper further addressed this by distinguishing demonstration-phase from hypothesis-testing-phase research, arguing that heterogeneity is expected and acceptable in the demonstration phase.3

Analysis. The heterogeneity concern is methodologically legitimate and partially unresolved; Utts’s response that it applies equally to mainstream meta-analyses is also methodologically defensible. No consensus resolution of this dispute exists in the pool.

Critique 2: Utts’s statistical conclusions implied endorsement of a psi mechanism, conflating data and interpretation

Skeptic source: Several commentators in the Statistical Science discussion argued that Utts’s conclusion, that the aggregate evidence indicated “an anomalous effect in need of explanation”, effectively endorsed psi as a real phenomenon, and that a statistician evaluating the data should have remained agnostic rather than drawing any positive conclusion from effect-size aggregates.6

Response: Utts explicitly and repeatedly separated the data claim from the mechanism claim, both in the original paper and in her rejoinder. She argued that concluding an anomalous effect is present in the data is a statistical judgment, not a metaphysical one, and that demanding a mechanistic explanation before accepting a data pattern is a standard not applied uniformly in other sciences. She cited the aspirin analogy: clinical efficacy was accepted before the mechanism was understood.7 Her 1991 paper’s abstract explicitly states the conclusion as “an anomalous effect in need of an explanation”, framing it as a data puzzle, not a confirmed phenomenon with a known cause.2

Analysis. Utts’s data/mechanism separation is methodologically well-grounded and consistently maintained across her publications. The critique conflates statistical inference with metaphysical commitment; her published responses address this directly and the separation is explicit in the original text.

References
  1. Utts, J. (1988). Successful replication versus statistical significance. Journal of Parapsychology, 52(4), 305–20. R001 [Utts 1988] ↩︎
  2. Utts, J. (1991). Replication and Meta-Analysis in Parapsychology. Statistical Science, 6(4), 363–403. https://doi.org/10.1214/ss/1177011577 R002 [Utts 1991] ↩︎
  3. Krippner, S., Braud, W., Child, I. L., Palmer, J., Rao, K. R., Schlitz, M., White, R. A., & Utts, J. (1993). Demonstration research and meta-analysis in parapsychology. Journal of Parapsychology, 57, 275–286. R003 [Krippner 1993] ↩︎
  4. Utts, J., Cook, T. D., Cooper, H. C., Cordray, D. S., et al. (1994). Meta-Analysis for Explanation: A Casebook. Journal of the American Statistical Association, 89(426), 715–715. https://doi.org/10.2307/2290883 R004 [Utts 1994] ↩︎
  5. Utts, J. M. (1991). Rejoinder. Statistical Science, 6, 396–403. R005 [Utts 1991] ↩︎
  6. Utts, J. (1991). Replication and meta-analysis in parapsychology. Statistical Science, 6(4), 363–403. https://doi.org/10.1214/ss/1177011577. R006 ↩︎
  7. Utts, J. (1991). [Replication and Meta-Analysis in Parapsychology]: Rejoinder. Statistical Science, 6(4). https://doi.org/10.1214/ss/1177011585 R007 [Utts 1991] ↩︎
  8. Utts, J. (2022). General and Personal Reflections on Succeeding as a Woman Science Researcher. Journal of anomalistics, 22(2), 355–399. https://doi.org/10.23793/zfa.2022.447 R008 [Utts 2022] ↩︎

See hub COI disclosure for subject-coauthorship transparency.