Jessica M. Utts, PhD Sources:
Remote Viewing Statistical Evaluation and Analysis
Jessica Utts is a statistician whose contributions to remote viewing research lie entirely in the domain of rigorous quantitative evaluation, applying standard statistical tools to assess whether the experimental record meets conventional scientific thresholds. Her work spans government-commissioned program reviews, methodological critiques of specific laboratory protocols, and the development of analytical frameworks for judging anomalous cognition data.
Deeper dives — Utts:
Key findings
- Utts’s 1995–1996 evaluation of the Stargate program concluded that the statistical evidence for anomalous cognition met conventional scientific standards, applying the same effect-size threshold she would apply to any scientific domain.1
- Her co-authored review of the SRI International remote viewing corpus (1973–1988) documented the cumulative statistical record across fifteen years of government-sponsored research.2
- Utts, Hansen, and Markwick identified specific methodological and statistical problems in the Princeton Engineering Anomalies Research (PEAR) remote viewing experiments, including inadequate randomization documentation and judging procedure ambiguities.3
- Fuzzy set technology was explored as an alternative analytical tool for remote viewing data, offering a way to quantify partial matches between viewer descriptions and target characteristics.4
- Decision Augmentation Theory, developed with Edwin May and colleagues, proposed a specific mechanistic framework for how anomalous cognition might operate, distinct from classical force-based models, with testable statistical predictions.5
Overview
Remote viewing research presents a distinctive statistical challenge: the outcome variable is a qualitative match between a viewer’s free-response description and a physical target, requiring a judging procedure that converts narrative similarity into a rankable or scorable quantity. Utts’s engagement with this literature is that of an applied statistician evaluating whether the scoring systems, randomization procedures, and inferential frameworks used in major remote viewing programs meet the standards she would apply to any scientific claim. A key Type-II vulnerability in this domain is that effect sizes are typically in the small-to-medium range by Cohen’s conventions, meaning that underpowered individual studies will frequently fail to detect genuine effects even if they exist, making cumulative meta-analytic evaluation essential rather than optional.1 Utts does not frame herself as a psi advocate; she frames her role as applying consistent statistical standards to an unusual empirical question.
Effect Size Threshold and the Statistician’s Role
Utts has articulated that her evaluation of the Stargate evidence applied the same criterion she would use for any scientific claim: whether the effect size is large enough, and the replication record consistent enough, to warrant taking the phenomenon seriously as a research object. She explicitly declined to treat parapsychology as requiring a higher evidential bar than other sciences.1 Her position is methodological rather than metaphysical, the data either meet conventional standards or they do not, and her assessment was that they did, while acknowledging that the mechanism remains entirely unknown.
The AIR / Stargate Evaluation
In 1995, the American Institutes for Research commissioned a formal evaluation of the U.S. government’s remote viewing program, known as Stargate, for the CIA and Department of Defense. Utts was one of two independent evaluators; the other was cognitive psychologist Ray Hyman. Her evaluation concluded that the statistical evidence for anomalous cognition in the Stargate corpus met conventional scientific standards, while Hyman’s evaluation, reviewing the same data, concluded that non-psi explanations remained viable and that the evidence did not warrant the anomalous cognition interpretation.16
Utts vs. Hyman, Where They Agreed and Disagreed
Utts and Hyman agreed on the statistical analysis itself, both acknowledged that the hit rates in the Stargate corpus exceeded chance expectation at statistically significant levels. Their disagreement was interpretive: Utts held that the effect sizes were comparable to those accepted as real in other scientific domains and that the replication record was sufficient to conclude the effect was genuine; Hyman held that methodological inadequacies in individual studies and the absence of a satisfactory non-psi explanation meant the anomalous cognition interpretation was premature.6 Utts’s response to Hyman’s report specifically addressed his methodological objections, arguing that the cumulative evidence across multiple independent studies could not be dismissed on the grounds of any single study’s flaws.6
Scope of the Stargate Corpus Evaluated
The Stargate program encompassed remote viewing research conducted across multiple government-sponsored laboratories over more than two decades. Utts’s evaluation drew on the cumulative record rather than individual studies, treating the program as a whole as the unit of analysis. Her published evaluation appeared in the Journal of Parapsychology and addressed both the laboratory experimental record and the operational applications of remote viewing within the program.1 The evaluation also appears in a later Elsevier volume co-authored with Edwin May, which situates the Stargate findings within the broader literature on non-sensory information access.7
Statistical Analysis of the SRI Remote Viewing Corpus
Before the AIR evaluation, Utts contributed to the statistical analysis of the remote viewing research conducted at SRI International between 1973 and 1988, a fifteen-year corpus representing the core of the government-sponsored experimental record. This work, co-authored with Edwin May and colleagues, documented the cumulative statistical outcomes across the SRI program and provided the quantitative foundation that later evaluations, including the AIR report, drew upon.2
The SRI Corpus, Scale and Statistical Summary
The SRI International review covered psychoenergetic research conducted from 1973 through 1988, representing a substantial body of experimental trials across multiple viewer populations and target types. The review by May, Utts, Trask, Luke, Frivold, and Humphrey documented the statistical outcomes of this corpus as an internal SRI report, providing the primary quantitative record of the program’s experimental history.2 A subsequent published analysis in the Journal of Parapsychology extended this work, examining advances in remote viewing analysis methodology and the consistency of effects across the corpus.8
Advances in Remote Viewing Analysis Methods
May, Utts, Luke, Frivold, and Trask’s 1990 paper in the Journal of Parapsychology examined methodological advances in how remote viewing data were analyzed, addressing the challenge of converting qualitative viewer descriptions into quantifiable outcomes. The paper considered the properties of different scoring and judging systems and their implications for statistical inference, a methodological concern that runs through all of Utts’s remote viewing work: that the statistical conclusions are only as valid as the measurement and judging procedures that generate the numbers.8
Critique of the PEAR Remote Viewing Experiments
Utts’s engagement with remote viewing research was not uniformly supportive of positive claims. Together with George Hansen and Betty Markwick, she published a detailed methodological critique of the remote viewing experiments conducted at Princeton Engineering Anomalies Research (PEAR), identifying specific statistical and procedural problems that undermined the evidential value of that program’s results.3 This critique illustrates that her evaluative standard was applied symmetrically, positive results from methodologically compromised studies do not support the hypothesis, consistent with the C28 doctrine that bad science yields zero conclusions in either direction.
Specific Methodological Problems Identified in the PEAR Remote Viewing Program
Hansen, Utts, and Markwick’s 1992 Journal of Parapsychology paper and the companion 1991 Research in Parapsychology proceedings paper identified several categories of problems in the PEAR remote viewing experiments.39 These included: (1) judging contamination, ambiguities in whether judges were adequately blind to target identity during the rating process, partially addressed by procedural description but not fully eliminated; (2) randomization documentation, inadequate documentation of how targets were selected, leaving open the possibility of non-random target assignment that could produce spurious matches; (3) multiple comparisons, the use of multiple dependent measures without appropriate correction, inflating the probability of false positives; and (4) response bias, the possibility that viewers’ descriptions were systematically biased toward certain target categories in ways that could interact with non-random target selection. The critique concluded that these problems were serious enough that the PEAR remote viewing results could not be taken as clean evidence for anomalous cognition, regardless of their statistical significance, an assessment that, given documented flaws, yields no conclusion in either direction per the symmetric standard Utts applies.
Analytical Frameworks: Fuzzy Sets and Decision Augmentation
Beyond evaluating existing data, Utts contributed to the development of novel analytical frameworks for remote viewing research. Two distinct approaches appear in the pool: fuzzy set technology as a scoring method for qualitative viewer descriptions, and Decision Augmentation Theory (DAT) as a mechanistic model with testable statistical implications.45 These represent Utts’s contribution to the measurement and modeling problems that sit upstream of any inferential conclusion about remote viewing.
Fuzzy Set Technology as a Remote Viewing Scoring Method
Humphries, May, and Utts’s 1988 paper presented at the 31st Annual Convention of the Parapsychological Association explored the application of fuzzy set theory to the problem of scoring remote viewing sessions. The core challenge is that viewer descriptions are rarely exact matches or complete misses, they are partial, graded correspondences between a verbal or drawn description and a physical target. Classical binary scoring (hit/miss) discards information; rank-order judging preserves ordinal information but loses the magnitude of match quality. Fuzzy set technology offers a mathematical framework for representing graded membership, a description can be 0.7 consistent with target A and 0.3 consistent with decoy B, potentially increasing the sensitivity of the scoring procedure and reducing the information loss inherent in cruder methods.4 The specific artifact this addresses is judging insensitivity: coarse scoring systems may fail to detect genuine partial correspondences, contributing to Type-II error (false negatives) in individual studies.
Decision Augmentation Theory, Model and Statistical Predictions
May, Utts, and Spottiswoode’s 1995 paper in the Journal of Parapsychology introduced Decision Augmentation Theory as a formal model of anomalous mental phenomena.5 DAT proposes that anomalous cognition operates not by physically influencing random systems (the classical psychokinesis model) but by influencing the timing of decisions about when to sample from a random process, so that the observer’s anomalous access to future information leads them to initiate measurements at moments that happen to produce favorable outcomes. This is a mechanistic interpretation, not a data claim; the researcher’s preferred interpretation is that DAT better fits the statistical signature of the existing data than force-based models. The proposed mechanism remains contested and has not achieved consensus. A companion paper applied DAT specifically to random number generator data, deriving testable predictions about the relationship between effect size and sample size that differ from what a genuine psychokinetic force model would predict.10 Alternative interpretations include conventional psychokinesis, publication bias, and methodological artifact, none of which DAT’s statistical predictions fully rule out in the existing literature.
DAT Applied to Random Number Generator Data
May, Utts, and Spottiswoode’s 1996 Journal of Scientific Exploration paper applied Decision Augmentation Theory to the random number generator (RNG) literature, testing whether the statistical signature of reported effects was more consistent with a genuine force-based influence on physical systems or with the decision-timing mechanism DAT proposes.10 The specific competing non-psi explanation tested was optional stopping, the possibility that experimenters unconsciously or consciously stopped data collection when results were favorable, producing inflated effect sizes. DAT’s predictions about the relationship between effect size and sample size were used as a discriminating criterion: if effects are genuine physical influences, effect size should be independent of sample size; if effects arise from decision timing, a specific pattern of size-sample covariation is predicted. The paper reported that the RNG data were more consistent with the DAT prediction than with a force model, though this finding is exploratory and the analysis has not been independently replicated with a preregistered protocol.
Modern Context
The statistical-evaluation tradition Utts applied to remote-viewing data — formal effect-size estimation under blinded judging, with explicit attention to multiple-comparison correction and post-hoc moderator selection — connects to mainstream signal-detection-theory frameworks for forced-choice and free-response paradigms. Canonical references on the underlying methodology include Macmillan and Creelman’s Detection Theory: A User’s Guide (2nd ed., 2005), which provides the standard treatment of d-prime estimation, bias measurement, and blinded-judging procedures applicable to closed-set target identification tasks. Subsequent metascience work by Ioannidis (2005, PLoS Medicine; 10.1371/journal.pmed.0020124) on inflated false-positive rates in underpowered exploratory analyses has reinforced the methodological concerns Utts raised in her AIR-era critiques: when many candidate moderators are tested without correction, spurious findings are guaranteed.
Skeptical Critiques and Discussion
Critique 1: Hyman’s critique: Methodological inadequacies in individual Stargate studies prevent an anomalous cognition conclusion
Skeptic source: Ray Hyman, reviewing the same Stargate corpus as Utts for the AIR evaluation, argued that the positive statistical outcomes could not be interpreted as evidence for anomalous cognition because individual studies within the corpus had methodological problems, including inadequate blinding of judges to target identity (judging contamination), insufficient documentation of randomization procedures, and the absence of a satisfactory non-psi explanation that had been rigorously tested and ruled out. Hyman’s position was that these unresolved methodological questions meant the anomalous cognition interpretation was premature, regardless of the statistical significance of the cumulative record.6
Response: Utts’s published response addressed Hyman’s methodological objections directly, arguing that the cumulative evidence across multiple independent studies, conducted at different laboratories, with different viewer populations, and with varying but overlapping methodological safeguards, could not be dismissed on the basis of any single study’s flaws. Her position was that the consistent replication of above-chance effects across methodologically varied studies was itself evidence against the hypothesis that the results were artifacts of any particular procedural weakness. She further argued that Hyman applied a higher evidential standard to parapsychology than to other scientific domains with comparable effect sizes and replication records.6 The SRI corpus review provides the underlying cumulative statistical record that Utts’s rebuttal draws upon.2
Analysis. Both Utts and Hyman reviewed the same data and reached different interpretive conclusions. Utts argued that the cumulative statistical record warranted recognition of an anomaly, while Hyman argued that methodological concerns in individual studies were sufficient to discount the cumulative record. Utts and Hyman did not contest each other’s calculations of the statistical effect; their disagreement concerns the weight to assign methodological critique relative to aggregate effect size.
Critique 2: The PEAR remote viewing results: Positive findings from a methodologically compromised program yield no evidential conclusion
Skeptic source: Hansen, Utts, and Markwick’s critique of the PEAR remote viewing experiments identified specific methodological problems, including judging contamination (ambiguities in whether judges were blind to target identity), inadequate randomization documentation (leaving open the possibility of non-random target assignment), multiple comparisons without correction (inflating false positive rates), and potential response bias (systematic viewer tendencies toward certain target categories), that collectively undermined the evidential value of the PEAR program’s positive results.39
Response: This critique is itself a methodological rebuttal from within the parapsychological research community, Utts is not defending the PEAR results but applying the same symmetric standard she applies elsewhere: methodologically compromised positive results do not support the anomalous cognition hypothesis. The critique does not claim the PEAR results demonstrate that remote viewing is absent; it claims they demonstrate nothing either way, because the documented flaws prevent any inference. This is the inconclusive tier: documented flaws mean the evidence yields no conclusion in either direction because the judging contamination, randomization gaps, and multiple-comparisons problems are unresolved in the PEAR protocol as described.3
Analysis. The PEAR remote viewing evidence carries zero weight in either direction because of the specific methodological problems identified by Hansen, Utts, and Markwick. This evidence neither supports nor refutes anomalous cognition claims.
References
- Utts, J. M. (1996). An evaluation of remote viewing: Research and applications. Journal of Parapsychology, 59(4), 289–320. R001 [Utts 1996] ↩︎
- May, E. C., Utts, J. M., Trask, V. V., Luke, W. W., Frivold, T. J., & Humphrey, B. S. (1989). Review of the psychoenergetic research conducted at SRI International (1973–1988). SRI-International. R002 [May 1989] ↩︎
- Hansen, G. P., Utts, J., & Markwick, B. (1992). Critique of the PEAR remote-viewing experiments. Journal of Parapsychology, 56(2), 97–113. R003 [Hansen 1992] ↩︎
- Humphries, B. S., May, E. C., & Utts, J. M. (1988). Fuzzy set technology in the analysis of remote viewing. Proceedings of the 31st Annual Convention of the Parapsychological Association, 378–394. R004 [Humphries 1988] ↩︎
- May, E. C., Utts, J., & Spottiswoode, S. J. P. (1995). Decision Augmentation Theory: Toward a Model of Anomalous Mental Phenomena. Journal of Parapsychology, 59(3), 195–220. R005 [May 1995] ↩︎
- Utts, J. (1995). Response to Ray Hyman’s Report ‘Evaluation of the Program on Anomalous Mental Phenomena.’ (Response to Ray Hyman’s Article in This Issue, P. 321). Journal of Parapsychology, 59(4), 353. R006 [Utts 1995] ↩︎
- Utts, J., & May, E. C. (2003). Non-sensory access to information: remote viewing. Elsevier, 59–73. https://doi.org/10.1016/b978-0-443-07237-6.50011-5 R007 [Utts 2003] ↩︎
- May, Utts, Luke, Frivold, Trask (1990). Advances in Remote Viewing Analysis. JP, 54(3), 193–228. R008 [May 1990] ↩︎
- Hansen, G. P., Utts, J., & Markwick, B. (1991). Statistical and methodological problems of the PEAR remote viewing experiments. Research in Parapsychology 1991. R009 [Hansen 1991] ↩︎
- May, E. C., Utts, J. M., & Spottiswoode, S. J. P. (1996). Decision augmentation theory: Applications to the random number generator. Journal of Scientific Exploration, 9, 453–488. R010 [May 1996] ↩︎
Deeper dives — Utts:
See hub COI disclosure for subject-coauthorship transparency.