Jessica M. Utts experiments and data
Jessica M. Utts, PhD — Experiments and Data
Before drawing any conclusions from what follows, a coverage qualification is warranted: the library’s share of the published literature on Jessica Utts’s work has not been measured. What appears below summarizes the studies currently in the ESP-Nexus library, and should be read as a sample of the record rather than a complete account of the field.
Utts is a statistician and Professor Emerita at the University of California, Irvine. Her ESP-Nexus profile is at Jessica M. Utts. Her contributions to parapsychology are methodological as much as experimental — she has applied meta-analysis, power analysis, and Bayesian reasoning to evaluate evidence assembled by others, and has also co-authored original empirical research on physiological anticipation.
Experiments
Utts’s role spans three distinct lines of work in the sources retrieved for this question:
Government remote-viewing evaluation. In 1995 Utts authored a report commissioned by the U.S. government to assess two decades of remote-viewing research conducted at SRI International and SAIC. This was not original experimental work but a formal statistical evaluation of an existing research program. The same AIR review paired her assessment with an independent skeptical evaluation by Ray Hyman.
Meta-analytic surveys of parapsychology. In a 1991 paper published in Statistical Science, Utts surveyed meta-analyses of ganzfeld, forced-choice precognition, RNG, and dice experiments, arguing that the accumulated data showed small but consistently nonzero effects across studies, experimenters, and laboratories. A 1999 analysis extended this to combined ganzfeld and remote-viewing data. A 2016 Bayesian analysis applied both frequentist and Bayesian frameworks to the same ganzfeld meta-analytic corpus.
Predictive anticipatory activity (PAA) meta-analyses. Utts co-authored two meta-analyses with Julia Mossbridge and Patrizio Tressoldi — a 2012 paper in Frontiers in Psychology and a 2014 follow-up in Frontiers in Human Neuroscience — examining whether physiological measures (skin conductance, heart rate, BOLD activity, and others) show anticipatory changes before randomly selected emotional stimuli, before those stimuli are presented.
SRI International technical report (relayed figures). The evidence table also includes figures from a 1988/1989 SRI International technical report by May, Utts, Trask, Luke, Frivold, and Humphrey, covering the SRI remote-viewing program from 1973 to 1988; those figures are discussed below under Data.
Bem, Utts, and Johnson (2011). The sources also reference a co-authored paper responding to Wagenmakers et al.’s Bayesian reanalysis of Daryl Bem’s precognition experiments, in which Utts and collaborators argued that psychologists need not change their analytic approach in the way the critics proposed.
Methodology
Meta-analysis and power analysis. Utts’s central methodological argument — stated in her 1991 Statistical Science paper and elaborated in the 2016 Bayesian analysis — is that defining successful replication as achieving p ≤.05 is incoherent when studies are underpowered. For small true effect sizes, most single studies will fail to reach significance even if the effect is real, making apparent non-replication an artifact of sample size. She used power curves to demonstrate this for the ganzfeld literature specifically, showing that median sample sizes in that literature produced low power against a true hit rate modestly above chance.
Bayesian framing. The 2016 analysis explicitly compared frequentist and Bayesian evaluations of the same ganzfeld data, computing both an overall p-value and posterior intervals under a range of priors — from an open-minded prior to a strongly skeptical one — to show how prior belief shapes posterior conclusions.
PAA meta-analytic protocol. In the 2012 and 2014 PAA meta-analyses, the protocol required that physiological measures be pre-specified as primary outcomes; post-hoc experiments were excluded. To guard against the multiple-comparisons problem, one analysis was constrained solely to electrodermal data. Studies were coded for methodological quality (randomization adequacy and expectation-bias controls), and higher-quality studies were analyzed separately to test whether quality moderated effect size. Publication bias was assessed using both classical fail-safe N (Rosenthal, 1979) and Orwin’s fail-safe, plus a trim-and-fill analysis.
Data
The evidence table covers six result rows across different metrics — standardized effect sizes (Cohen’s d or h), raw hit rates, and one p-value-only entry. Because these are different, non-comparable metrics, no pooled “overall” figure across all rows is meaningful; the table below reports each study on its own terms.
| Study | Metric | ES / Hit rate | CI | z | p | k / N | Direction |
|---|---|---|---|---|---|---|---|
| Mossbridge, Tressoldi & Utts (2012) | Cohen’s d | ES = 0.21 | 6.9 | < 2.7 × 10⁻¹² | k=26 studies | Positive | |
| Schmidt (2004) | Cohen’s d | ES = 0.11 | — | .001 | k=36, N=1,015 sessions | Positive | |
| Utts (1999) | Raw hit-rate diff. | +0.09 above chance; hit rate = 0.34 | — | — | N=2,097 sessions | Positive | |
| May, Utts et al. (1988/1989) ⚠ | p only | — | — | — | < 10⁻²⁰ | k=154, N=26,000 trials, 227 participants | Positive |
| Dalton (1995) | Raw hit rate | 0.314 | — | — | .0009 | N=354 sessions | Positive |
| Utts (1991) | Cohen’s h | ES = 0.20; hit rate = 0.344 | — | 0.00005 | k=11, N=355 trials, 241 participants | Positive |
On the PAA meta-analyses. Mossbridge, Tressoldi, and Utts (2012) reported a fixed-effect overall Cohen’s d of 0.21 (95% CI 0.15–0.27, z = 6.9, p < 2.7 × 10⁻¹²) across 26 studies of physiological anticipatory activity. Higher-quality studies produced a quantitatively larger effect size and greater significance than lower-quality ones. The 2014 follow-up carried the same overall fixed-effect figures and extended the discussion to practical applications. Individual study effect sizes ranged from −0.138 to 0.67, indicating meaningful heterogeneity; this is an unsettled signal — not all contributing studies pointed in the same direction.
On the SRI figures (⚠ relayed). The row for May, Utts, Trask, Luke, Frivold, and Humphrey (1988/1989) reports figures from the SRI International Technical Report as relayed and re-derived in Utts (1996). These are the SRI program’s figures, not Utts’s own independent finding; they are attributed to the original technical report, and the holding paper’s role is to document and evaluate them.
On the DMILS row. The Schmidt (2004) row reports a DMILS (Direct Mental Interaction with Living Systems) meta-analysis of 36 studies, quality-weighted, yielding Cohen’s d = 0.11. This row is associated with Utts in the evidence table but the primary authorship is Schmidt (2004); Utts’s role in that analysis is not specified in the retrieved sources.
Unsettled signals the evidence carries:
- Substantial study-to-study variation in hit rates: the 2016 Bayesian analysis found that individual study-specific hit rates spanned a wide interval, a pattern the frequentist and Bayesian analyses agreed on.
- The PAA individual study ESs range from negative (−0.138) to strongly positive (0.67), reflecting genuine heterogeneity.
- Effect sizes across the different paradigms (PAA, ganzfeld, remote viewing, DMILS) are on different metrics and cannot be compared directly.
No null or below-chance meta-analytic result appears in the six evidence rows held by the library for this question, but this may reflect the library’s current holdings rather than the complete picture — the coverage is unmeasured.
Skeptical critiques
What critics argue. No published critique of this work appears in the sources retrieved for this question.
Regarding the PAA meta-analyses, the 2012 paper itself documented that one plausible alternative explanation is expectation bias — researchers’ awareness of which trials are emotional or calm could influence the physiological recording or analysis. The paper also acknowledged that multiple-analyses artifacts could inflate the apparent effect if researchers selected the dependent variable that produced the strongest result, which is why the constrained electrodermal-only analysis was conducted as a check.
Wagenmakers and colleagues’ Bayesian reanalysis of Bem’s precognition data — referenced in — concluded that the data did not support the psi hypothesis under Bayesian evaluation. Utts, Bem, and Johnson (2011) replied in the same journal that the Wagenmakers analysis used an inappropriately diffuse prior that biased the conclusion against any small effect.
What the experimental data show. The PAA meta-analyses reported that higher-quality studies — those that specifically addressed randomization and expectation-bias concerns — produced a quantitatively larger effect size rather than a smaller one, which cuts against the methodological-artifact explanation, though the authors noted that the difference between high- and low-quality subgroups was not itself statistically significant. Publication bias analyses (fail-safe N, Orwin’s method, trim-and-fill) did not eliminate the effect.
Analysis. Hyman and Utts reached opposite conclusions from the same body of government-sponsored data in 1995, and that disagreement has not been resolved by a jointly agreed reanalysis. The PAA literature’s individual study effect sizes range from negative to strongly positive, meaning the phenomenon is not uniformly observed. The Wagenmakers–Utts/Bem exchange illustrates a genuine methodological tension: Bayesian conclusions in this domain are sensitive to prior specification, and no neutral standard for choosing a prior in psi research has been established. Pre-registration of PAA experiments — recommended explicitly in Mossbridge, Tressoldi, and Utts (2014) — has been proposed as a way to address the multiple-analyses concern, but the sources retrieved here do not report how many subsequent studies followed that recommendation.
| Paper | Reported finding | Effect / significance | Basis |
|---|---|---|---|
| Mossbridge et al. (2012), Frontiers in Psychology [source] | Overall pooled effect – fixed-effect model. | ES 0.21, z = 6.9, p < 2.7 × 10−12 | 26 studies |
| Schmidt et al. (2004), British Journal of Psychology | DMILS Model 1 – 36 studies, effect sizes weighted by overall quality. | ES 0.11, p = .001 | 36 studies; N = 1015 sessions |
| Utts (1999), Journal of Scientific Exploration | Combined ganzfeld + remote viewing. | hit rate 0.34 | N = 2097 sessions |
| Utts (1996), Journal of Scientific Exploration ⚠ relayed figures | Figures are May, Utts, Trask, Luke, Frivold & Humphrey (1988/1989), SRI International Technical Report — as relayed/re-derived in this paper, not its own experiment. SRI overall analysis 1973-1988. | p < 1 × 10−20 | k = 154; N = 26000 trials; 227 participants |
| Dalton et al. (1995), Proceedings of the 38th Annual Convention of The Parapsychological Association [source] | Overall PRL autoganzfeld ESP hit rate. | p = 9 × 10−4, hit rate 0.314 | N = 354 sessions; 354 participants |
| Utts (1991), Statistical Science [source] | Autoganzfeld – overall direct hits. | ES 0.2, p = 5 × 10−5, hit rate 0.344 | 11 studies; N = 355 trials; 241 participants |
References
- Mossbridge, J., Tressoldi, P., & Utts, J. (2012). Predictive physiological anticipation preceding seemingly unpredictable stimuli: a meta-analysis. Frontiers in Psychology, 3. https://doi.org/10.3389/fpsyg.2012.00390
- Schmidt, S., Schneider, R., Utts, J., & Walach, H. (2004). Distant intentionality and the feeling of being stared at: Two meta-analyses. British Journal of Psychology, 95, 235–247.
- Utts, J. (1999). The Significance of Statistics in Mind-Matter Research. Journal of Scientific Exploration, 13(4), 615–638.
- Utts, J. (1996). An Assessment of the Evidence for Psychic Functioning. Journal of Scientific Exploration.
- OCR-garbled byline, affiliation line reads ‘University of Edinburgh, & University of California at SARC.’ Per the matching reference in the paper’s own reference list (and the claimed citation), the authors are reported as Dalton, K., Watt, C. & Lawrence, T. (1995) (1995). Sex pairings, target type and geomagnetism in the PRL automated ganzfeld series. Proceedings of the 38th Annual Convention of The Parapsychological Association.
- Utts, J. (1991). Replication and Meta-Analysis in Parapsychology. Statistical Science, 6(4), 363–403.
More questions answered
- Pool the effect sizes across every ganzfeld study you hold and compare your number to the published Storm and Tressoldi meta-analyses—do they agree?
- Has the effect size for PK data increased in the last 50 years?
- Tell me about the clairvoyance work of the last 50 years
- What trends can you see in precognition research?
- What trends can you see in ESP research in the last 50 years?
- What are parapsychology's current arguments to justify that the phenomena are real and should be taken seriously?