Replication in Psi Research
Measurement and Methodology
Measurement choices in psi research carry consequences the statistical argument on the hub does not engage. Four observations the published methodological literature documents: binary scoring discards qualitatively rich data; forced-choice precognition has zero sensory-leakage opportunity by design; experimenter is a documented source of effect-size variance; and meta-analytic inclusion-criteria choices can produce diametrically opposed conclusions on overlapping data.
Deeper dives — Replication in Psi Research:
1. Binary scoring discards qualitatively rich data
The dominant scoring methods in psi research collapse qualitatively complex output to a single binary or rank-order number. In the ganzfeld paradigm, a participant in a sensory-isolation state produces a mentation transcript — sometimes paragraphs long, often including specific imagery, emotional tone, structural detail, and metaphorical content. Independent judges then rank-order the transcript against four target images. The participant’s “hit” status reduces to whether the actual target was ranked first. A vivid transcript that describes the actual target’s structural and emotional content but happened to be ranked second or third counts as a miss.
In remote-viewing protocols (CIA-era Stanford Research Institute work; later Princeton Engineering Anomalies Research) the same reduction happens. The viewer’s transcript can include drawings, written impressions, references to specific features (a structure, a body of water, a mood). A judge or computer matches against possible targets and produces a rank score. The score doesn’t capture how close the transcript was to the actual target — only whether the rank ordering placed the target high enough.
This is a known limitation in the field. Utts (1995)[1], in her statistical evaluation of the SRI/SAIC remote-viewing program for the U.S. government, explicitly notes that the published effect-size estimates use scoring methods that necessarily discard correspondence information. Viewers with strong qualitative correspondence to a target that wasn’t ranked first contribute nothing to the standard hit-rate statistics.
What this means for replication interpretation: the published effect sizes — already at the upper end of what mainstream behavioral science accepts as informative (autoganzfeld Cohen’s h = 0.20) — are computed on a reduced representation of the data. Richer scoring methods that capture partial correspondence, structural match, and degree-of-similarity would produce different — and on the available evidence, likely larger — effect estimates.
Whether richer scoring would change the proponent-vs-skeptic balance is an open empirical question. What is not open: the current scoring methods are information-lossy, and the effect-size estimates they produce are floor estimates of what richer methods might detect. Modern computational approaches — large language model judges, semantic-similarity embeddings, structural-feature comparators — make richer scoring methodologically feasible in a way that wasn’t true even ten years ago. This is an active research frontier.
2. The sensory-leakage paradox in forced-choice precognition
The most common skeptical artifact-explanation for ganzfeld results is sensory leakage: subtle cues from the experimenter, equipment, or environment that reveal target information through normal sensory channels. This is a legitimate concern for ganzfeld designs where the target image exists at the time of the viewer’s mentation and a sender (in classic ganzfeld) or the experimenter (in modified protocols) has access to it.
But the parapsychology effect base extends beyond ganzfeld. The cross-paradigm table on the replication hub includes forced-choice precognition (Honorton & Ferrari 1989[2]), a 309-study meta-analysis covering more than 2 million trials, with combined z = 11.41 and p = 6.3 × 10⁻²⁵.
Forced-choice precognition has a methodologically distinctive property: the target does not exist at the time of the participant’s response. The typical protocol asks the participant to predict, at time T, the outcome of a random event that will be generated at time T + Δt (where Δt ranges from seconds to days). The target is generated AFTER the response is recorded.
This eliminates sensory leakage as an artifact-explanation by design. There is nothing to leak. No experimenter knows the target. No equipment has the target. No environmental cue can carry information about an event that has not yet occurred. The standard sensory-channel hypothesis for psi effects — that participants are picking up subtle normal-channel information — cannot apply to forced-choice precognition by the structure of the experiment itself.
This forces critics of forced-choice precognition results onto a different artifact-explanation. The remaining possibilities: (a) selective publication of positive results (file-drawer effect); (b) optional-stopping or researcher-degrees-of-freedom bias; (c) systematic non-randomness in the random-event generators; (d) the participant somehow influencing the random-event generation prospectively (which itself would be an anomalous result of independent interest). Each of these has been engaged in the published methodological literature, but none is sensory leakage.
The implication: a critic who invokes sensory leakage as the umbrella artifact-explanation for psi must concede that explanation does not apply to forced-choice precognition, which has effect-size estimates of comparable magnitude to ganzfeld. The critique must shift to file-drawer or design-level bias arguments, which the cross-paradigm table on the hub engages with directly (fail-safe N analysis, quality-effect-size correlations).
This is not a proof of psi. It is a structural property of one paradigm that constrains the space of valid artifact-explanations.
3. Experimenter effects: same protocol, different result
Among the most-documented and methodologically uncomfortable findings in psi research is the experimenter effect. The clearest published demonstration is Wiseman & Schlitz (1997)[3]: both researchers conducted formally identical remote-staring protocols, in the same laboratory, with their own recruited participants. Schlitz, the proponent researcher, reported significant effects. Wiseman, the skeptic researcher, reported null effects. A follow-up Schlitz, Wiseman, Radin & Watt (2006)[4] 2×2 greeter/sender cross-over design — where each researcher served as greeter or sender on different trial subsets — failed to replicate the previous findings: the ANOVA showed no significant greeter effect, sender effect, or interaction, and the condition with Schlitz as both greeter and sender was also nonsignificant. The correct conclusion from the three-study sequence is that experimenter effects remain a serious open variable, not that the 1997/1999 pattern was preserved.
Two interpretations of this pattern, both consistent with the data:
The proponent reading: psi is a skill-mediated phenomenon. Some experimenters, by some combination of rapport, expectancy, belief, or unmeasured-and-perhaps-unmeasurable participant-interaction skill, are able to elicit the relevant psychological state in participants. Others are not. This is consistent with how performance-dependent psychological phenomena work in other domains (clinical-interview-elicited responses, sports-coach-elicited athletic performance, music-teacher-elicited student performance). The experimenter is not a nuisance variable to be controlled away; the experimenter is part of the eliciting condition.
The skeptical reading: experimenter expectancy effects persist under formally identical protocols. Even when procedural details are matched, experimenter belief can leak through micro-cues, subtle scoring decisions, participant-recruitment criteria, or rapport differences that affect participant cooperation. Standard double-blind protocols do not always control these; in face-to-face experimental conditions they are particularly hard to eliminate. The experimenter effect is therefore an artifact of imperfect procedural control, not evidence of psi.
The empirical observation is identical under both readings: under formally identical conditions, two researchers produce reliably different results. Both interpretations have substantial published defenders. Both have substantial published critics.
The recent parapsychology methods literature has formalized the experimenter-psi research direction as its own theoretical question: Kruth (2022)[5] in a Journal of Parapsychology editorial; Graff (2023)[6] and Drucker (2023)[7] follow-on engagement; Braud (1975)[8] as the foundational reference. The field has not converged on whether experimenter effects represent a methodological problem (artifact-mediated) or a real moderator of a real phenomenon (skill-mediated), but the documentation that experimenter is a non-trivial source of variance is consistent across decades.
What this means for replication interpretation: experimenter is not a nuisance variable. It is a documented systematic source of effect-size variance. The methodological implication is shared by both interpretations: “independent replication” requires independence at the experimenter level, not just at the laboratory level. Adversarial collaboration in the Hyman-Honorton tradition — the same skeptical and proponent researchers jointly designing and conducting the replication — is the methodological response that takes the experimenter effect seriously.
4. Meta-analytic divergence: same paradigm, opposite conclusions
A second methodological observation, distinct from the experimenter effect but related: meta-analytic inclusion-criteria choices can produce diametrically opposed conclusions on overlapping data. The clearest example in the ganzfeld database is a published exchange in Psychological Bulletin:
- Milton & Wiseman (1999)[9] meta-analyzed 30 ganzfeld studies from 1987 to 1997, using narrow post-Joint-Communique inclusion criteria. Result: Stouffer Z = 0.70, p = 0.24, mean effect size d = 0.013. Near-chance.
- Storm, Tressoldi & Di Risio (2010)[10] meta-analyzed a broader set of free-response studies from 1992 to 2008, using inclusion criteria that incorporated additional studies. Result for the homogeneous ganzfeld data set (29 studies, 1997-2008): mean effect size = 0.142, Stouffer Z = 5.48, p = 2.13 × 10⁻⁸. (Stouffer’s combined-Z method, same statistic as Milton-Wiseman 1999, so the comparison is apples-to-apples; the divergence reflects different study-inclusion sets, not different statistical methods.) Two other free-response subsets in the same paper showed positive but smaller effects (nonganzfeld noise-reduction, Stouffer Z = 3.35) or near-zero effects (standard free-response, Stouffer Z = -2.29).
- Hyman (2010)[11], in the same journal, published a critical comment titled “Meta-analysis that conceals more than it reveals,” arguing the Storm et al inclusion criteria and quality-weighting choices were too permissive to produce a defensible conclusion.
- Storm, Tressoldi & Di Risio (2010 reply)[12] in the same journal, “A meta-analysis with nothing to hide: Reply to Hyman,” defending their methodology choices and engaging Hyman’s specific objections.
Same fundamental paradigm. Overlapping data. Different inclusion criteria. Opposite conclusions. And the disagreement itself is a peer-reviewed published exchange in a top mainstream psychology journal — not a behind-closed-doors dispute, but an adversarial collaboration playing out in the open record. That is the right model for how contested evidence should be debated.
The deeper methodological question this case study raises: could both meta-analyses be correct? The answer depends on what “correct” means.
If “correct” means “internally consistent given the explicit methodological choices,” then yes, both can be correct. Milton-Wiseman applied a narrow inclusion frame consistent with post-Joint-Communique standards; their Z = 0.70 is the correct aggregate within that frame. Storm et al applied a broader free-response frame including pre- and post-Communique work; their positive Z is the correct aggregate within that frame. Both internally consistent; both producing different answers to slightly different questions.
If “correct” means “produces the truth about whether psi exists,” the answer is that at most one can be — and possibly neither, if both inclusion frames are too narrow or too broad to characterize the underlying phenomenon (or its absence) accurately.
The factor most likely to explain meta-analytic divergence on contested topics is subjective quality coding. Meta-analytic protocols typically require coders to rate each included study on quality dimensions (randomization adequacy, blinding, scoring procedure, sample-size justification, etc.). These ratings then weight or filter the included studies. Quality coding is rarely entirely objective — reasonable coders working from the same study report can produce different quality scores depending on which features they prioritize and how strictly they apply criteria. A coder predisposed to skepticism can rate flaws more severely; a coder predisposed to proponency can rate the same flaws as immaterial. The resulting aggregate — the “corrected” effect size after quality-weighting — can swing substantially based on this subjective layer.
The Hyman 2010 critical comment focused precisely on this question: Hyman argued Storm et al’s quality coding and inclusion choices systematically favored producing a positive aggregate. Storm-Tressoldi-Di Risio 2010 in their rejoinder defended their choices as defensible by published meta-analytic standards. Both made their choices explicit, in the published record, and the reader is left to evaluate which choices are more defensible given the specific studies.
For the broader question this page addresses — how should “failed replication” evidence be interpreted — the implication is: meta-analytic conclusions are inclusion-criteria-dependent and quality-coding-dependent in ways that are not transparent to readers reading either paper in isolation. A reader reading either Milton-Wiseman or Storm et al alone could reasonably conclude “the ganzfeld evidence is settled” in their direction. Reading both papers together, plus the Hyman comment and the Storm reply, the appropriate conclusion is that meta-analytic methodology is itself a moderator of meta-analytic outcome — inclusion choices shape conclusions in ways that have nothing to do with whether psi is real. The methodological implication: any meta-analytic claim about psi should be evaluated against the specific inclusion criteria and quality coding used, not against an assumed-universal “the ganzfeld evidence.”
This is not a proof that one paper is right and the other wrong. It is a constraint on how meta-analytic evidence should be read — including by readers tempted to cite either paper as decisive.
5. Scope: “psi” as umbrella vs. distinct phenomena
The term “psi” was adopted by parapsychologists in the mid-20th century as a deliberately neutral label, declining to specify mechanism or even to assert that the phenomena indexed are the same kind of thing. The original Hyman-Honorton (1986) Joint Communique adopted this convention explicitly: “psi” denoted an unexplained communications anomaly, not a positive mechanistic claim.
Under the umbrella label, however, parapsychology has studied at least four distinguishable phenomena:
- Telepathy: information transfer between two living minds (sender → receiver). Ganzfeld is the primary paradigm.
- Clairvoyance: information acquisition about a distant or hidden physical state, without an active sender. Remote-viewing is the primary paradigm.
- Precognition: information acquisition about future events. Forced-choice precognition is the primary paradigm.
- Psychokinesis (PK): mind influencing matter, typically random-event-generator output. RNG and dice-influence are the primary paradigms.
These have different empirical profiles, different effect-size estimates, different artifact-explanation candidates, and may or may not share underlying mechanisms (if they index a real phenomenon at all). Cross-paradigm meta-analyses — the kind the replication hub features in Section 9 — average across these distinctions. The combined-evidence claim “meta-analytic syntheses across paradigms show small positive effects” averages across phenomena that may be ontologically distinct.
Two implications of the scope ambiguity:
First, evidential weight should be paradigm-specific, not psi-general. The case for forced-choice precognition (large corpus, sensory-leakage-immune by design, fail-safe N very large) is structurally different from the case for ganzfeld (smaller corpus, sensory-leakage concerns require control, mixed independent-replication record). A reader evaluating the evidence should evaluate paradigm-by-paradigm rather than via the umbrella term.
Second, the term “psi” may be doing rhetorical work the evidence doesn’t support. If telepathy, clairvoyance, precognition, and PK are distinct phenomena, calling them all “psi” implies a unification of phenomenon-existence claims that the data may not justify. A reader skeptical of telepathy may be told “but precognition also has positive results, and they all reflect psi” — a move that requires unification not in evidence. The honest editorial frame is to allow each paradigm’s evidence to stand or fall on its own.
This page, and the replication hub it spokes from, use “psi” because that is the term the field uses. But the umbrella conceals scope distinctions that matter for fair evaluation.
6. Implications for replication interpretation
The five measurement-and-methodology points above converge on a single editorial implication:
“Failed to replicate” is harder to interpret than the standard skeptical reading implies, even before any statistical-argument considerations from the hub apply. Specifically:
- If the measurement method discards information that richer methods would capture, the effect-size estimate that “failed to replicate at the expected magnitude” may be a methodologically-floor-estimate of a phenomenon whose true magnitude is unknown.
- If the paradigm whose result “failed to replicate” is sensory-leakage-immune (forced-choice precognition), the leakage-based dismissal does not apply; alternative artifact-explanations must be specified.
- If the replication is conducted under different experimenter-elicitation conditions than the original, the failed result tests a different empirical question than the original demonstrated.
- If the meta-analysis evaluating the “failure” used different inclusion criteria than the meta-analysis evaluating the original, the divergence may reflect inclusion-criteria differences rather than phenomenon-existence differences.
- If the original and replication are studying potentially-distinct phenomena under a shared umbrella label, the meta-level conclusion “psi failed to replicate” averages across paradigm-specific evidential profiles in a way that may obscure real distinctions.
None of these observations proves that any specific psi effect is real. Each constrains the inferential space within which skeptical dismissals are valid. The integrated argument across hub and spokes is that the standard skeptical move — “the effect failed to replicate, therefore it is not real” — requires methodological assumptions that are often unstated and sometimes demonstrably wrong for specific paradigms.
7. What this page establishes and what it does not
This page establishes:
- The dominant scoring methods in psi research (binary direct-hit rankings) discard qualitatively-rich information. The published effect-size estimates are floor estimates of what richer methods might detect; richer methods are an active research frontier.
- Forced-choice precognition is sensory-leakage-immune by design (target does not exist at response time). The sensory-leakage artifact-explanation, often invoked against ganzfeld results, structurally cannot apply to this paradigm.
- Experimenter effects are documented (Wiseman & Schlitz 1997, 2006; recent synthesis Kruth 2022, Graff 2023, Drucker 2023) and are interpretable in both proponent and skeptic directions. Both interpretations require taking experimenter-as-variable seriously, which the standard “any-lab-can-replicate” framing does not.
- Meta-analytic inclusion-criteria choices can produce diametrically opposed conclusions on overlapping data (Milton & Wiseman 1999 vs Storm, Tressoldi & Di Risio 2010). Meta-analytic conclusions are inclusion-criteria-dependent in ways often not transparent to readers.
- “Psi” is an umbrella label covering at least four distinguishable phenomena (telepathy, clairvoyance, precognition, PK), each with its own empirical profile. Cross-paradigm meta-analyses may average across distinctions that matter.
This page does NOT establish:
- That any psi phenomenon exists. The measurement-and-methodology arguments are about how to interpret existing evidence fairly, not about whether the evidence is conclusive.
- That richer scoring methods, when applied retroactively to existing datasets, will produce larger effect estimates. This is an open empirical question and an active research direction.
- That the experimenter effect is necessarily skill-mediated rather than expectancy-mediated. The two interpretations have substantial published defenders.
- Which meta-analytic inclusion criteria (narrower à la Milton-Wiseman or broader à la Storm et al) better reflect “the truth about the ganzfeld database.” That question requires per-included-study scrutiny that this page does not perform.
- That the four phenomena listed under the psi umbrella are ontologically distinct. They may be distinct, may share a deeper structure, or may all be artifacts of distinct kinds — the question is open.
Together with the statistical argument on the hub and the behavioral-science argument in the companion spoke, this page contributes one part of a three-part case: the standard skeptical dismissals of psi research require methodological assumptions that are not always defensible, and a fair evaluation requires acknowledging the measurement, paradigm-specific, meta-analytic-methodology, and scope considerations that this spoke documents.
Sources
- Utts, J. (1995). An assessment of the evidence for psychic functioning. Report prepared for the American Institutes for Research review of the U.S. government’s remote-viewing program; later published as Utts, J. (1996), Journal of Scientific Exploration, 10(1), 3–30. ↩︎
- Honorton, C., & Ferrari, D. C. (1989). “Future telling”: A meta-analysis of forced-choice precognition experiments, 1935–1987. Journal of Parapsychology, 53, 281–308. ↩︎
- Wiseman, R., & Schlitz, M. (1997). Experimenter effects and the remote detection of staring. Journal of Parapsychology, 61, 197–207. ↩︎
- Schlitz, M., Wiseman, R., Watt, C., & Radin, D. (2006). Of two minds: Sceptic-proponent collaboration within parapsychology. British Journal of Psychology, 97(3), 313–322. ↩︎
- Kruth, J. G. (2022). Editorial: Synthesizing thoughts on experimenter psi. Journal of Parapsychology, 86(2), 161–164. ↩︎
- Graff, D. E. (2023). Letters to the editor: Experimenter psi considerations. Journal of Parapsychology, 87(1). ↩︎
- Drucker, D. (2023). Letters to the editor: Expanding the experimenter role. Journal of Parapsychology, 87(1). ↩︎
- Braud, W. (1975). Psi-conducive states. Journal of Communication, 25(1), 142–152. https://doi.org/10.1111/j.1460-2466.1975.tb00563.x ↩︎
- Milton, J., & Wiseman, R. (1999). Does psi exist? Lack of replication of an anomalous process of information transfer. Psychological Bulletin, 125(4), 387–391. https://doi.org/10.1037/0033-2909.125.4.387 ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992–2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136(4), 471–485. https://doi.org/10.1037/a0019457 ↩︎
- Hyman, R. (2010). Meta-analysis that conceals more than it reveals: Comment on Storm et al. (2010). Psychological Bulletin, 136(4), 486–490. https://doi.org/10.1037/a0019676 ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). A meta-analysis with nothing to hide: Reply to Hyman (2010). Psychological Bulletin, 136(4), 491–494. https://doi.org/10.1037/a0019840 ↩︎
Further reading. Landmark work in this literature not cited inline above:
- Hyman, R., & Honorton, C. (1986). Joint communiqué: The psi ganzfeld controversy. Journal of Parapsychology, 50, 351–364.