Jessica M. Utts, PhD Sources:
Ganzfeld ESP Methodology and Statistical Analysis
Jessica Utts brought the rigorous tools of academic statistics to bear on one of parapsychology’s most contested experimental paradigms: the ganzfeld procedure. Her contributions were methodological rather than experimental, she evaluated existing data, identified statistical weaknesses, and proposed frameworks for assessing whether the accumulated evidence met the standards applied in other sciences.
Deeper dives — Utts:
Key findings
- Utts’s 1986 statistician’s perspective on the ganzfeld debate identified specific methodological weaknesses in both pro-psi and skeptical analyses of the early ganzfeld literature, arguing that neither side had applied fully adequate statistical reasoning.1
- A 1999 Bayesian analysis of ganzfeld studies co-authored by Utts provided posterior probability estimates for effect sizes, offering an alternative to frequentist significance testing that better captured cumulative evidence across trials.2
- Exploratory analyses of the Princeton Research Laboratories (PRL) automated ganzfeld database examined whether sender–receiver sex pairing, target type, and geomagnetic activity moderated hit rates, findings intended as hypothesis-generating rather than confirmatory.3
- In a 2020 referee report on a large preregistered ganzfeld meta-analysis, Utts provided detailed methodological commentary on inclusion criteria, effect-size estimation, and the handling of heterogeneity across studies.4
- Across her ganzfeld-related work, Utts consistently applied the same effect-size threshold she applied to other sciences, arguing that dismissing a replicable effect of comparable magnitude to accepted findings in other domains requires explicit justification.1
Overview
The ganzfeld procedure, in which a receiver in a state of mild sensory attenuation attempts to identify a target viewed by a distant sender, accumulated a substantial experimental literature through the 1970s and 1980s. By the mid-1980s, the statistical adequacy of that literature had itself become contested, with critics and proponents disagreeing not only about whether effects were real but about whether the analyses used to evaluate them were sound. Utts entered this debate as a statistician, not as an experimenter, and her contributions focused on the quality of the statistical reasoning applied to ganzfeld data rather than on any single experiment’s outcome. A key Type-II vulnerability in this literature is that individual ganzfeld studies are typically underpowered to detect small true effects (hit rates modestly above the 25% chance baseline in a four-alternative forced-choice design), meaning that dismissal based on any single null result would be statistically premature without adequate power analysis.1
The Ganzfeld Procedure and Its Statistical Baseline
In the standard ganzfeld protocol, a receiver is placed in a reclining position with halved ping-pong balls over the eyes and white noise delivered through headphones, producing uniform visual and auditory fields that reduce patterned sensory input without eliminating sensory experience. A sender in a separate room views a randomly selected target (typically a video clip or image). After the session, the receiver is shown four items, the actual target plus three decoys, and asked to rank or select the best match. Chance performance on this four-alternative forced-choice task is 25%. The effect size of interest is the degree to which hit rates exceed this baseline. Utts noted that the statistical power of individual studies to detect a true hit rate of, say, 33% above a 25% baseline is modest at typical sample sizes, making cumulative meta-analytic assessment essential rather than optional.1
Early Statistical Critique of the Ganzfeld Debate
In 1986, Utts published a statistician’s perspective on the ganzfeld debate in the Journal of Parapsychology, examining the statistical arguments made by both proponents and critics of the early ganzfeld literature. Her analysis identified specific problems on both sides: proponents had not always adequately addressed multiple-comparison issues arising from exploratory analyses, while critics had sometimes applied standards of statistical rigor that were not consistently demanded of other scientific literatures.1
Specific Statistical Problems Identified in the 1986 Analysis
Utts’s 1986 paper examined the ganzfeld debate from a statistical standpoint, identifying that the dispute involved not just empirical disagreement but disagreement about appropriate statistical standards. On the proponent side, she noted that exploratory analyses of moderator variables, such as sender–receiver relationship quality or target type, were sometimes reported without adequate correction for multiple comparisons, inflating apparent significance. On the skeptic side, she observed that demands for replication under conditions of exact protocol replication were more stringent than those applied to comparable literatures in psychology and medicine, where conceptual replication across varied protocols is the norm. She also noted that the effect sizes reported in ganzfeld studies, while modest, were not smaller than effect sizes routinely accepted as meaningful in other behavioral sciences, a point she would develop more fully in later work on effect-size standards.1
Bayesian Inference for Ganzfeld Studies
In 2010, Utts presented a Bayesian analysis of the accumulated ganzfeld literature co-authored with Norris, Suess, and Johnson at the Eighth International Conference on Teaching Statistics (ICOTS-8) in Ljubljana. The paper, “The Strength of Evidence Versus the Power of Belief: Are We All Bayesians?”, analyzed 56 procedurally standard ganzfeld experiments comprising 2,124 sessions and 709 hits — an aggregate hit rate of 33.4% against a chance baseline of 25% (exact-binomial p = 2.26 × 10⁻¹⁸). Rather than relying solely on frequentist p-values, the analysis modeled three explicit Bayesian priors — a “psi-skeptic” prior centered near chance, an “open-minded” weakly informative prior, and a “psi-believer” prior — and reported how each was updated by the data into a posterior. The framing was pedagogical as well as evidentiary: presented at a statistics-education conference, the paper used ganzfeld as a worked example of how prior belief shapes the interpretation of identical data.2
Bayesian Framework Applied to Ganzfeld Hit Rates
The Utts, Norris, Suess, and Johnson (2010) Bayesian analysis modeled the ganzfeld hit rate first under a simple binomial framework (X ~ Binomial(2124, p), with conjugate Beta priors) and then under a Bayesian hierarchical model that allowed for between-study heterogeneity in true effect sizes. Under the hierarchical model with a weakly informative prior, the posterior median for the per-study hit rate sat at approximately 0.33 with a 95% credible interval of roughly 0.30 to 0.36. Under the skeptic’s tight prior centered at 0.25, the data shifted the posterior median only slightly (to ≈ 0.258) — making explicit, in Utts’s framing, why skeptics “still are not convinced by the evidence, even with a p-value of 2.26 × 10⁻¹⁸”: a sufficiently tight prior on chance overwhelms even very strong likelihood evidence. The paper used this contrast as a pedagogical illustration of how identical data update different priors differently, addressing the broader concern that disagreement over psi evidence often reflects prior belief differences rather than disagreement about the data themselves.2
Moderator Variables in the Automated Ganzfeld Series
A 1995 paper co-authored by Dalton and Utts examined the PRL automated ganzfeld database for potential moderators of hit rates, including sender–receiver sex pairing, target type, and ambient geomagnetic activity. The analyses were explicitly framed as exploratory, intended to generate hypotheses for future preregistered testing rather than to confirm effects, and the paper acknowledged that the multiple comparisons involved required cautious interpretation.3
Sex Pairing, Target Type, and Geomagnetic Moderators in the PRL Database
The Dalton and Utts (1995) analysis drew on the PRL automated ganzfeld database, which at the time represented one of the largest single-laboratory ganzfeld datasets available, spanning over a decade of trials conducted at the Princeton Research Laboratories using a standardized automated protocol. The automated system addressed sensory leakage by using computer-controlled target selection and presentation, physically isolating sender and receiver in separate rooms, eliminating direct cueing between sender and receiver, though not fully addressing all potential experimenter-mediated pathways. The paper examined three potential moderators: (1) whether same-sex versus mixed-sex sender–receiver pairs showed different hit rates; (2) whether target type (static images versus dynamic video clips) moderated performance; and (3) whether ambient geomagnetic activity on the day of testing correlated with hit rates, following earlier reports by other researchers of geomagnetic correlates in anomalous cognition tasks. The analyses were daily hit-rate and effect-value calculations using the PRL database. Because these analyses were purely exploratory and involved multiple comparisons across the three moderator variables and their interactions, the paper explicitly cautioned against treating any observed patterns as confirmed effects without independent preregistered replication.3
Modern Context
The methodological challenges Utts identified in ganzfeld moderator analyses, particularly the multiple-comparisons problem in exploratory moderator searches and the need for Bayesian rather than purely frequentist inference, sit within a broader mainstream statistical debate about exploratory versus confirmatory research practices. The distinction between exploratory and confirmatory analysis that Utts applied to ganzfeld moderator work anticipates concerns about researcher degrees of freedom that have become central to replication discussions across behavioral science more broadly.5 The integration of ethical and methodological standards in statistical practice, including transparency about analysis provenance, is a theme Utts has continued to develop in her mainstream statistical work, including contributions to guidelines for statistical education and practice.6
Registered Report Meta-Analysis: Referee Engagement
In 2020, Utts served as a referee on a Stage 1 Registered Report for a large preregistered meta-analysis of ganzfeld anomalous perception studies. Her referee report engaged with the methodological design of the meta-analysis before data collection was complete, the defining feature of the Registered Report format, which separates methodological review from outcome-dependent publication decisions.4
Referee Commentary on the Preregistered Ganzfeld Meta-Analysis
The 2020 Utts referee report (published via Faculty of 1000 Research as part of the open peer review process for the Stage 1 Registered Report on anomalous perception in a ganzfeld condition) addressed methodological questions including study inclusion criteria, the handling of heterogeneity across studies with different protocols, and the statistical approach to effect-size estimation. The Registered Report format is directly relevant to concerns Utts had raised across her career about the need to separate exploratory from confirmatory analyses: by reviewing the protocol before outcomes are known, referees can evaluate the methodological adequacy of the design without the results influencing their assessment. Utts’s engagement with this format as a referee represented a continuation of her long-standing position that the ganzfeld literature’s credibility depends on the quality of its statistical and methodological infrastructure, not merely on the accumulation of nominally significant results.4
The open questions that remain in the ganzfeld literature, including whether effect sizes are stable across independent laboratories using fully automated protocols, and whether identified moderators such as sender–receiver relationship quality replicate under preregistered conditions, are precisely the questions that a well-designed adversarial preregistered multi-site study would be positioned to address. Utts’s career-long emphasis on distinguishing exploratory from confirmatory findings points toward this as the methodological standard the literature would need to meet to move beyond the current contested state.4
Skeptical Critiques and Discussion
Critique 1: Ganzfeld meta-analyses conflate methodologically heterogeneous studies, inflating apparent effect sizes
Skeptic source: A persistent methodological critique of ganzfeld meta-analyses is that pooling studies with substantially different protocols, varying degrees of sender–receiver isolation, different target sets, different judging procedures, produces an effect-size estimate that does not correspond to any single well-controlled design. Under this view, apparent consistency across studies reflects shared methodological weaknesses rather than a replicable phenomenon. This critique applies to the statistical aggregation approach Utts and colleagues employed.1
Response: Utts’s 1986 analysis directly engaged this concern, noting that heterogeneity across studies is a feature of virtually all behavioral science meta-analyses and that the appropriate response is to model heterogeneity explicitly rather than to reject meta-analytic aggregation entirely. The 1999 Bayesian analysis by Utts, Johnson, and Suess addressed between-study heterogeneity by using a hierarchical model that allowed true effect sizes to vary across studies, rather than assuming a single fixed effect, partially mitigating (though not eliminating) the concern that pooling masks meaningful protocol differences.2
Analysis. The hierarchical-modeling approach Utts applied partitions between-study variance from within-study sampling variance, which addresses one component of the heterogeneity raised by Hyman. The reported between-study variance component remained non-zero after modeling, indicating that protocol differences across studies account for some portion of the effect-size spread. Utts has argued that the residual effect after modeling supports the cumulative-evidence interpretation; Hyman has argued that protocol variation remains a candidate explanation for the residual.
Critique 2: Exploratory moderator analyses in ganzfeld research inflate false-positive rates through multiple comparisons
Skeptic source: The search for moderators of ganzfeld hit rates, including sender–receiver sex pairing, target type, and geomagnetic activity examined in the Dalton and Utts (1995) analysis, involves testing multiple hypotheses on the same dataset. Without preregistration and correction for multiple comparisons, the probability of finding at least one nominally significant moderator by chance is substantially elevated above the nominal alpha level. Critics have argued that the moderator literature in ganzfeld research is largely a product of this multiple-comparisons inflation rather than genuine moderation of a real effect.3
Response: Utts and Dalton explicitly acknowledged this limitation in their 1995 paper, framing the analyses as purely exploratory and cautioning that observed patterns required independent preregistered replication before being treated as confirmed effects. This acknowledgment does not resolve the multiple-comparisons problem for the specific analyses reported, but it does represent an honest characterization of the evidential weight the findings should carry, consistent with Utts’s broader methodological position that exploratory and confirmatory findings must be clearly distinguished.3 Her 2020 referee engagement with a preregistered ganzfeld meta-analysis reflects the same principle applied at the meta-analytic level.4
Analysis. The multiple-comparisons concern is well-founded and was acknowledged by the authors themselves. The specific moderator findings from the 1995 exploratory analysis carry limited evidential weight pending preregistered replication. The critique is partially addressed by the authors’ own framing but not by the data.
References
- Utts, J. M. (1986). The ganzfeld debate: A statistician’s perspective. JP, 50, 393–402. R001 [Utts 1986] ↩︎
- Utts, J., Norris, M., Suess, E., & Johnson, W. (2010). The strength of evidence versus the power of belief: Are we all Bayesians? In C. Reading (Ed.), Data and context in statistics education: Towards an evidence-based society. Proceedings of the Eighth International Conference on Teaching Statistics (ICOTS-8), Ljubljana, Slovenia. International Statistical Institute. https://iase-web.org/documents/papers/icots8/ICOTS8_8H1_UTTS.pdf R002 [Utts 2010] ↩︎
- Dalton, K., & Utts, J. (1995). Sex pairings, target type, and geomagnetism in the PRL automated ganzfeld series. Proceedings of the 38th Annual Convention of the Parapsychological Association, 99–112. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/dalton-utts-1995.pdf R003 [Dalton 1995] ↩︎
- Utts, J. (2020). Referee report. For: Stage 1 Registered Report: Anomalous perception in a Ganzfeld condition – A meta-analysis of more than 40 years investigation [version 1; peer review: 1 approved]. Faculty of 1000 Research…. https://doi.org/10.5256/f1000research.27439.r68427 R004 [Utts 2020] ↩︎
- Utts, J. (2021). Statistical Practice Is Not a Spectator Sport. Harvard Data Science Review. https://doi.org/10.1162/99608f92.ff65fd7a R005 [Utts 2021] ↩︎
- Raman, R., Utts, J., Cohen, A. I., & Hayat, M. J. (2022). Integrating Ethics into the Guidelines for Assessment and Instruction in Statistics Education (Gaise). The American Statistician, 77(3), 323–330. https://doi.org/10.1080/00031305.2022.2156612 R006 [Raman 2022] ↩︎
Deeper dives — Utts:
See hub COI disclosure for subject-coauthorship transparency.