Jessica M. Utts, PhD Sources:
Non-Sensory Access to Information: The 1995 AIR Assessment of Remote Viewing
This page covers the specific 1995 AIR (American Institutes for Research) review document in which Jessica Utts evaluated the statistical evidence from the U.S. government’s Stargate remote viewing program at SRI International and SAIC. The document, titled An Assessment of the Evidence for Psychic Functioning, was commissioned by the CIA alongside a companion review by Ray Hyman; together they constitute one of the most cited mainstream-statistical evaluations of laboratory psi data. This page focuses on the AIR document as a standalone artifact: its scope, the specific effect sizes Utts reported, and the specific conclusion she drew. The broader topic of Utts’s remote-viewing statistical work is covered on the Remote viewing statistical evaluation spoke.
Deeper dives — Utts:
Key findings
- The AIR review was commissioned by the CIA in 1995 as a retrospective evaluation of the Stargate program’s two decades of government-funded remote viewing research; Utts and Hyman were selected as the two domain reviewers familiar with parapsychological methodology.1
- Utts reported aggregate effect sizes for the laboratory remote-viewing corpus of d = 0.209 across 770 SRI sessions and d = 0.230 across 455 SAIC sessions, consistent with the small but cumulatively significant effects she had documented in her 1991 Statistical Science review.1
- Her central statistical conclusion was that “the effect sizes reported … are too large and consistent to be dismissed as statistical flukes” and that, by ordinary scientific standards applied to small-effect literatures in medicine and psychology, the laboratory remote-viewing data demonstrated an anomalous form of information access in need of explanation.1
- Utts explicitly preserved the data/mechanism separation: her conclusion concerned what the data show, not what causes it. She did not claim the AIR review established a specific psi mechanism, only that the aggregate evidence pattern met the same evidentiary threshold she applied across sciences.1
- Hyman’s companion review accepted the statistical reliability of the effect sizes but disputed the inference that the laboratory results demonstrated psychic functioning, arguing that the absence of a known mechanism and the lack of independent replication outside the SRI–SAIC programs warranted reservation.2
Overview
The 1995 AIR review was a contracted retrospective evaluation: Utts and Hyman were provided with documentation of the laboratory remote-viewing experiments conducted at SRI International (under principal investigators including Harold Puthoff and Russell Targ, and later Edwin May) and at Science Applications International Corporation (SAIC, under Edwin May). The reviewers were not asked to evaluate the operational intelligence-collection use of remote viewing, which was the subject of a separate AIR analysis, but to assess whether the laboratory data established the existence of an anomalous information-access phenomenon under controlled conditions.1
Why “Non-Sensory Access to Information” Is the Operative Phrase
Utts’s framing of the AIR question, whether the data demonstrate “non-sensory access to information”, deliberately uses a description-neutral phrasing that focuses on the data pattern rather than on any specific mechanism. The phrase encompasses what the laboratory protocols were designed to test: under conditions in which sender, target, and receiver were physically isolated, with computer-controlled target selection, with double-blinded judging of receiver-produced descriptions against target pools, would receiver descriptions correspond to targets more often than chance allows? This is a description of an information-access pattern, not a commitment to any particular causal account (precognition, clairvoyance, electromagnetic coupling, statistical artifact, residual sensory leakage). Utts’s report consistently maintained this descriptive framing.1
Commission, Scope, and Data Provided
In a memorandum dated July 25, 1995, SAIC’s Edwin May provided the AIR reviewers with a documented set of ten experiments, each accompanied by a short title, the number of trials, the computed effect size, and the overall p-value. This pre-computed data package constituted the primary empirical input to Utts’s review; she also drew on the broader published SRI and SAIC literature where it bore on methodological questions.1
What the Ten Experiments Were Designed to Test
The SAIC experiments documented in the July 25 memorandum were, per Utts’s report, “all designed to answer questions about psychic functioning” rather than merely to add additional confirmation of existence. That is, the SAIC program had moved past the demonstration-research phase Utts described in her 1993 collaborative paper into hypothesis-testing research, with each experiment designed to address a specific question: target-pool design effects, sender presence/absence, distance effects, target-feedback effects, and so forth.1 This shift toward hypothesis-testing was precisely what Utts had argued the field needed; she treated the SAIC corpus as the most methodologically mature segment of the laboratory remote-viewing literature available at the time of the review.
Effect Sizes and Statistical Conclusion
Utts reported aggregate effect sizes of d = 0.209 across 770 SRI sessions and d = 0.230 across 455 SAIC sessions, with overall p-values that combined to statistical significance at conventional thresholds many orders of magnitude beyond chance. These effect sizes fell within the d ≈ 0.2–0.4 range Utts had treated as scientifically meaningful when replicated across independent laboratories in her 1991 Statistical Science review.1
The “Too Large and Too Consistent” Conclusion
Utts’s central statistical conclusion in the AIR review was that “the effect sizes reported … are too large and consistent to be dismissed as statistical flukes.”1 The argument rested on two observations: first, that the effect sizes were of magnitude comparable to small but accepted effects in medicine and psychology (where d ≈ 0.2–0.3 is treated as scientifically meaningful when replicated); and second, that the SRI and SAIC effect sizes converged closely (0.209 vs. 0.230) despite the two programs being run by partially different personnel under partially different protocols, a convergence that would be unlikely if either set were a chance artifact. She did not conclude that a specific psi mechanism was confirmed; she concluded that the data pattern was real and required explanation. The data/mechanism distinction was made explicit in the report.
SRI and SAIC Convergence as Cross-Laboratory Replication
The methodological argument Utts emphasized was that the SRI and SAIC programs, while operating under partially shared personnel and institutional culture, were nevertheless distinct enough to constitute meaningful cross-laboratory replication. The SAIC program had been designed in part to address methodological criticisms raised against the earlier SRI protocols, including criticisms about target-pool sensory-leakage and judging procedures, and the persistence of effect sizes at the same magnitude under the improved SAIC protocols was, in Utts’s view, evidence that the effect was robust to the specific artifact concerns raised against the earlier work.1 Critics, including Hyman, accepted the convergence point but argued that SRI and SAIC were too closely related to constitute genuinely independent replication and that confirmation outside the SRI–SAIC institutional lineage remained outstanding.2
Modern Context
The methodological frame Utts applied in the 1995 AIR evaluation — formal effect-size estimation, blinded judging, attention to Type-I and Type-II error trade-offs — sits within a mainstream signal-detection literature canonically treated in Macmillan and Creelman’s Detection Theory: A User’s Guide (2nd ed., Lawrence Erlbaum, 2005; ISBN 978-0805842319). The AIR report was paired with Ray Hyman’s critical appraisal, continuing a long-running mainstream-skeptical engagement with the parapsychological literature that traces to Hyman’s earlier methodological critiques. From the Bayesian side, Wagenmakers, Wetzels, Borsboom, and van der Maas (2011, Journal of Personality and Social Psychology; 10.1037/a0022790) argued that default Bayesian analyses of psi data yield weaker evidence than frequentist analyses, prompting the response from Bem, Utts, and Johnson (2011) and a subsequent multi-cycle methodological exchange about how Bayes factors should be specified for novel-effect claims.
The Hyman Counter-Evaluation
Ray Hyman’s companion AIR review accepted the statistical reliability of the effect sizes Utts reported but disputed the inference that they established the existence of psychic functioning. His central objections were three: that the absence of a known mechanism warranted reservation about accepting the data as evidence for a novel phenomenon; that independent replication outside the SRI–SAIC institutional lineage was insufficient; and that the SAIC experiments, while methodologically improved, had not been independently audited at the data-handling level.2
Utts’s 1996 Response to Hyman’s Report
Following the AIR review, Utts published a formal response to Hyman’s report addressing each of his objections.3 On the mechanism objection, she returned to the aspirin analogy she had used in her 1991 Statistical Science rejoinder: clinical efficacy was accepted before mechanism was understood, and demanding mechanistic explanation as a prerequisite for accepting a data pattern is a standard not uniformly applied across sciences. On the independent-replication objection, she pointed to the ganzfeld and PRL meta-analyses as constituting independent laboratory programs producing comparable effect sizes, with the autoganzfeld results in particular addressing the specific artifact concerns raised against earlier ganzfeld work. On the data-handling audit objection, she acknowledged the limitation but noted that the AIR commission had not provided for such an audit and that the absence of audit-level scrutiny applied equally to many accepted small-effect literatures in mainstream science.
Skeptical Critiques and Discussion
Critique 1: The SRI and SAIC programs were too institutionally entangled to constitute independent replication
Skeptic source: Hyman’s AIR companion review and subsequent commentary argued that the SRI and SAIC remote-viewing programs shared personnel, methodology, target-pool conventions, and institutional culture to a degree that they should be treated as a single institutional lineage rather than as two independent replications, weakening the cross-laboratory convergence argument Utts emphasized.2
Response: Utts acknowledged the partial overlap but argued that the SAIC program had been specifically designed to address SRI methodological criticisms and that the persistence of comparable effect sizes under the improved SAIC protocols was meaningful evidence of effect robustness, even if SAIC was not maximally independent of SRI.3 She also pointed to converging evidence from other laboratory programs (ganzfeld at Honorton‘s PRL, RNG studies at PEAR) as constituting cross-paradigm rather than within-paradigm convergence. The critique that maximally independent replication outside the SRI–SAIC lineage was insufficient remains a methodologically legitimate concern.
Analysis. The institutional-entanglement concern is real; Utts’s response acknowledges the limitation and points to cross-paradigm convergence as partial mitigation. No consensus resolution exists; subsequent decades have produced additional ganzfeld and PAA meta-analyses but no maximally independent remote-viewing replication of the SAIC scale.
Critique 2: The conclusion that the data establish “non-sensory access” presupposes the absence of subtle sensory leakage that the protocols failed to detect
Skeptic source: A standard methodological concern raised by Hyman and others against the SRI–SAIC remote-viewing corpus is that the conclusion of “non-sensory access” requires affirmatively ruling out all conventional information channels, including subtle target-pool stereotypy, judge-rating artifacts, and statistical dependencies among target sets, and that the protocols did not perform exhaustive artifact ruling-out.2
Response: Utts argued in the AIR report and in her 1996 response that the specific artifacts named by critics, including target-pool stereotypy and judge-rating bias, had been addressed in the SAIC protocol revisions and that the persistence of the effect under the revised protocols was evidence the named artifacts did not fully account for it.13 She acknowledged that no protocol can rule out all conceivable artifacts and that the burden of identifying a specific unaddressed artifact, rather than asserting the abstract possibility of one, falls on the critic. The data/mechanism separation she maintained throughout means her conclusion does not require commitment to “non-sensory” in any metaphysical sense, only to “not explained by the conventional sensory channels the protocols controlled for.”
Analysis. Utts’s response is methodologically well-grounded; the residual concern that some unnamed subtle artifact remains is real but, absent a specific identified candidate, does not constitute a positive disconfirmation.
References
- Utts, J. (1995). An Assessment of the Evidence for Psychic Functioning. (Report prepared for the American Institutes for Research review of the U.S. Government’s remote viewing programs.) https://ics.uci.edu/~jutts/air.pdf R001 [Utts 1995] ↩︎
- Hyman, R. (1995). Evaluation of a program on anomalous mental phenomena. (Report prepared for the American Institutes for Research review of the U.S. Government’s remote viewing programs.) R002 [Hyman 1995] ↩︎
- Utts, J. (1996). Response to Ray Hyman’s report of September 11, 1995, “Evaluation of program on anomalous mental phenomena.” Journal of Scientific Exploration, 10(1), 87–92. R003 [Utts 1996] ↩︎
Deeper dives — Utts:
See hub COI disclosure for subject-coauthorship transparency.