Replication in Psi Research
Why “failed to replicate” doesn’t mean what you think. A widespread misconception about replication — held even by trained psychologists — drives much of the public conversation about which scientific findings are “real.” The misconception is that repeating an experiment should produce the same significant result every time. Statistical power theory shows this is mathematically wrong. At the modest power levels typical of behavioral-science studies, a real effect is statistically expected to produce only partial replication — often 2-of-3 or 1-of-3 successful results. The skeptical argument that “the effect failed to replicate three times in a row, therefore it isn’t real” commits a Type II reasoning error.
Deeper dives — Replication in Psi Research:
At 50% power — which is roughly what well-conducted social-psychology and parapsychology studies actually achieve — there is a 50/50 chance that fewer than 2 of 3 replication attempts will reach significance, EVEN IF THE EFFECT IS REAL.
1. The replication misconception — and the evidence that scientists themselves hold it
If you ask a non-specialist what it means for a scientific finding to “replicate,” the typical answer is: another team repeats the same experiment and gets the same significant result. Behind this answer is a stronger implicit claim — that a finding which doesn’t replicate, in this sense, isn’t real.
The first piece of evidence that something is wrong with this picture is that even professional psychologists overestimate replication probabilities by nearly 2-to-1. In a study published in 1982, Amos Tversky and Daniel Kahneman[2] — the founders of modern behavioral-decision research — surveyed 84 members of the American Psychological Association and the Mathematical Psychology Group with this question:
Suppose you have run an experiment on 20 subjects, and have obtained a significant result which confirms your theory (z = 2.23, p < .05, two-tailed). You now have cause to run an additional group of 10 subjects. What do you think the probability is that the results will be significant, by a one-tailed test, separately for this group?
The median answer was 0.85. Most respondents believed there was an 85% chance of replication. Only 9 of the 84 respondents gave an answer in the range 0.40–0.60, which is where the actual answer lies. The mathematically correct answer, assuming the first result is close to the true population effect, is approximately 0.47 — slightly less than even odds. (This 0.47 is an optimistic ceiling: in practice, an initial significant result is upward-biased relative to the true population effect — the “winner’s curse” — so the actual replication probability would be somewhat lower. The respondents’ miscalibration vs. the 0.47 ceiling is in fact even larger when winner’s curse is properly accounted for.)
If APA-member psychologists overestimate replication probability by nearly 2-to-1, the entire public conversation about which findings “failed to replicate” is operating on a flawed shared assumption — including assumptions that drive popular-reference articles, science-journalism articles, and skeptical commentary on contested findings.
The same misconception extends to how scientists evaluate “failed” replications. In a companion question, Tversky and Kahneman asked respondents what t-value in a second study they would consider a failure to replicate an original result with t = 2.46 (significant). Most said t = 1.70 (nonsignificant). But if you simply combine the two studies’ data, the pooled result is t = 2.94, p = .003 — substantially more significant than the original. As Utts (1988) summarized this paradox:
The new study DECREASES faith in the original result if viewed separately but INCREASES it when combined with the original data.
Two studies, identical data, opposite conclusions — depending on whether you “vote-count” the individual results or pool the underlying observations.
2. Two kinds of statistical error
To understand why this happens, consider what a statistical test actually does. When a scientist performs a hypothesis test (the typical “p < .05” calculation), two opposite errors are possible:
| Error type | Technical meaning | Plain English |
|---|---|---|
| Type I (false positive) | The test reports an effect that isn’t really there in the population. | “The data looks interesting, but it was actually just random fluctuation.” |
| Type II (false negative) | The test fails to detect an effect that is really present in the population. | “There IS something real going on — but the experiment wasn’t sensitive enough to catch it.” |
The convention “p < 0.05” exists specifically to control Type I error. Researchers set their significance threshold low enough that random fluctuation will produce a “significant” result less than 5% of the time when no real effect exists. What this convention does NOT do is address Type II error.
A study that doesn’t reach p < 0.05 might be:
- Detecting that no effect exists, OR
- Failing to detect an effect that exists — because the experiment was too small, too noisy, or otherwise underpowered to detect it
Both interpretations are statistically compatible with the same null result. A “nonsignificant” study tells you nothing about which of these two scenarios is true. Most science writing — and most skeptical commentary on contested findings — focuses almost entirely on Type I error. Popular-reference articles, popular-science explanations, and adversarial commentary treat “failed to replicate” as if it were unambiguous evidence the original was a false positive. The Type II possibility is systematically ignored.
The argument that skeptical commentary in parapsychology systematically ignores Type II error was made directly by Berger (1989)[1] in his critical examination of Susan Blackmore’s psi-experiment database in the Journal of the American Society for Psychical Research: “It appears that it is Blackmore’s argument that flaws can plausibly only lead to false positives (Type I errors). It is beyond the scope of this paper to elaborate, but there are a number of design flaws that can lead to false negatives (Type II errors). These include, but are not limited to, inadequate sample size (low statistical power), weak or inappropriate statistical tests, sampling from inappropriate populations, experimenter expectancy effects, demand characteristics, and the faulty operationalization of dependent measures” (Berger 1989, p. 137). Berger also identified the asymmetric-standards pattern explicitly: skeptical methodological objections are routinely invoked to dismiss positive results but withdrawn when the same studies produce nulls (p. 135–136). The present page extends Berger’s 1989 argument to the modern statistical literature.
3. What statistical power means
Statistical power is the probability that an experiment will detect a real effect, if one exists, at the chosen significance level. Power is determined by three things:
- The size of the underlying effect — bigger real effects are easier to detect.
- The size of the sample — larger samples produce more precise estimates and detect smaller effects.
- The significance threshold — a stricter threshold (e.g., p < .01 vs p < .05) reduces both false-positive and detection rates.
An experiment with 80% power has an 80% chance of producing a “significant” result if the underlying effect is real. An experiment with 20% power has only a 20% chance — even if the effect is exactly as large as the researchers expect. Power is, in this sense, the conditional probability of getting the result you’re trying to demonstrate, given that the underlying claim is true.
Most behavioral-science studies — including most psi studies, but also most studies in social psychology, education research, and clinical psychology — operate at much lower power than researchers typically assume. Power calculations are routinely skipped or done casually, with researchers implicitly assuming power will be adequate.
Jacob Cohen, author of the 1988 textbook Statistical Power Analysis for the Behavioral Sciences[4] (the canonical reference on this subject), made this point bluntly in 1990:
Despite widespread misconceptions to the contrary, the rejection of a given null hypothesis gives us no basis for estimating the probability that a replication of the research will again result in rejecting that null hypothesis.
Robert Rosenthal — the pioneer of meta-analysis in psychology, commissioned by the U.S. National Academy of Sciences to evaluate parapsychology in 1988 — said the same thing in the same flagship journal:
Given the levels of statistical power at which we normally operate, we have no right to expect the proportion of significant results that we typically do expect, even if in nature there is a very real and very important effect.
These quotes are from American Psychologist — the flagship journal of the American Psychological Association. This is not a parapsychology-specific argument. It is a mainstream-statistics consensus that the conventional way of reading replication results — vote-counting significant studies — is methodologically unsound for low-power literatures.
4. The binomial reality of replication — worked numbers
Concretely: if a real effect produces power = p in each replication attempt, then the probability of N successful replications out of K attempts follows a binomial distribution.
| Power per study | P(3 of 3 significant) | P(2 of 3 or better) | P(exactly 2 of 3) | P(1 of 3 or fewer) |
|---|---|---|---|---|
| 0.80 (high — well-funded mainstream) | 51.2% | 89.6% | 38.4% | 10.4% |
| 0.50 (typical contested-effect study) | 12.5% | 50.0% | 37.5% | 50.0% |
| 0.30 (modest power) | 2.7% | 21.6% | 18.9% | 78.4% |
| 0.18 (typical psi-ganzfeld at N=28) | 0.6% | 8.6% | 8.0% | 91.4% |
Read this table carefully. At 50% power — which is roughly what well-conducted social-psychology and parapsychology studies actually achieve — there is a 50/50 chance that fewer than 2 of 3 replication attempts will reach significance, EVEN IF THE EFFECT IS REAL. Three-of-three significant replications at this power level is a 12.5% outcome — improbable rather than expected.
At the 0.18 power that Utts (1986) calculated[8] for typical 28-trial ganzfeld studies assuming a real 33% hit rate, three-of-three significant replications happens only 0.6% of the time. It would be statistically astonishing for an underpowered series to replicate at p < .05 three times in a row, even if psi were real.
Utts (1988, p. 312) calculated concrete numbers for typical psi ganzfeld studies. If the true hit rate is 33% against a 25% chance baseline:
- A study with 26 trials is expected to reach “significance” only about one-fifth of the time
- A study with 100 trials is expected to reach significance only about half the time
Her conclusion: “It is no wonder that there are so many ‘unsuccessful’ attempts at replication in psi.” The unsuccessful attempts are statistically expected — they are evidence about sample-size choices, not about the absence of an underlying effect.
5. A worked example: the ganzfeld database
The ganzfeld experiments — designed to test whether subjects can identify a randomly selected target image being viewed by a distant sender, under conditions of sensory isolation — provide the cleanest demonstration of why the misconception matters. The ganzfeld database is also one of the most thoroughly analyzed in all of parapsychology, with both proponent and critic publishing meta-analyses of the same studies.
In 1985, parapsychologist Charles Honorton and skeptic psychologist Ray Hyman each analyzed 28 ganzfeld experiments and reached opposite conclusions. Hyman pointed to the count: of the 24 studies that reported “direct hits,” 13 of 24 (54%) were individually nonsignificant at the conventional p < 0.05 threshold (one-tailed). On a vote-count basis, this looks like a failed-replication story.
But Utts (1991, p. 370) computed what happens when those 13 “failures” are pooled — combining the underlying trial-level data rather than the individual study verdicts: 106 direct hits out of 367 trials. z = 1.66, p = 0.0485. Pooled together, the supposed-failures themselves reach statistical significance. The individual studies were not failures of replication — they were each too small to detect the effect alone. Vote-counting throws away the information that pooling preserves.
This is the structural reason vote-counting fails: when individual studies have modest power, their evidential weight only emerges through proper combination. Hedges & Olkin (1985)[9] — the foundational reference on meta-analytic methods — showed mathematically that vote-counting becomes increasingly likely to make the wrong decision as the number of studies grows. The vote-count method discards information at exactly the rate that meta-analysis aggregates it.
Robert Rosenthal — commissioned by the National Academy of Sciences to write a background paper for their 1988 parapsychology report — computed a meta-analytic synthesis of the 28 ganzfeld direct-hit studies:
One of his metrics requires definition: in meta-analytic terms, the file-drawer problem refers to the systematic bias that arises when null or negative studies remain unpublished — sitting in researchers’ file drawers — while only positive results enter the literature. The corresponding fail-safe N estimates how many such hidden null studies would be required to overturn an observed meta-analytic effect.
- Combined z = 6.60 (p = 3.37 × 10⁻¹¹)
- 23 of the 28 studies had positive effect sizes (median Cohen’s h = 0.32)
- Mean direct-hit rate ≈ 0.38 against a 0.25 chance baseline (95% CI 0.30–0.46)
- File-drawer fail-safe N = 423 unreported null studies required to negate the effect
Rosenthal’s conclusion was that the ganzfeld effect was real (Rosenthal 1986)[10]. Ironically, his commissioned analysis was not cited in the NAS report’s eventual conclusion, which relied instead on Hyman’s vote-count critique:
The discussion of the ganzfeld work in the National Academy Report focused on Hyman’s 1985 analysis, but never mentioned the work it had commissioned Rosenthal to perform, which contradicted the final conclusion in the report.
This is a documented case of the asymmetric epistemic standard at work: an independent statistician was commissioned, performed the analysis, returned a conclusion that contradicted the report’s framing, and was simply not cited in the final document. Utts (1991, p. 372) documents the specific case: an independent meta-analytic synthesis commissioned by the NAS was simply not cited in the report’s final conclusion, which relied instead on a vote-count critique. This illustrates how asymmetric-standards problems can operate at the institutional review level. Generalizing from this one documented omission to a “systematic preference” or “known bias” would require broader evidence; the page treats the NAS case as an illustrative datapoint, not as proof of a fieldwide pattern.
6. The “three-of-three” skeptical demand — and why it fails
The clearest articulation of the underlying skeptical argument comes from British psychologist C.E.M. Hansel, in his book[11] ESP and Parapsychology: A Critical Re-evaluation (1980):
If a result is significant at the .01 level and this result is not due to chance but to information reaching the subject, it may be expected that by making two further sets of trials the antichance odds of one hundred to one will be increased to around a million to one, thus enabling the effects of ESP — or whatever is responsible for the original result — to manifest itself to such an extent that there will be little doubt that the result is not due to chance.
Hansel was demanding three consecutive experiments at p ≤ .01 as the standard for accepting an effect as real. This is, in essence, the same argument that appears in popular reference works and contemporary skeptical commentary on contested findings: if the effect were real, it would replicate.
Utts’s response, in the same flagship-statistics paper:
This argument implies that if a particular experiment produces a statistically significant result, but subsequent replications fail to attain significance, then the original result was probably due to chance, or at least remains unconvincing. The problem with this line of reasoning is that there is no consideration given to sample size or power. Only an experiment with extremely high power should be expected to be “successful” three times in succession.
To put numbers on this: for Hansel’s demand to be statistically reasonable, each replication study would need power of at least 0.80 — meaning each study would need to be large enough that, given the true effect size, a 95%-or-greater probability of reaching the one-tailed .05 result is realistic. For ganzfeld studies at the time Hansel was writing, this would have required approximately 345 trials per study (Utts 1991, p. 374).. The actual mean ganzfeld study size in the database Hansel critiqued was approximately 28 trials. Hansel demanded a degree of replication that, by his own logic, would only be statistically achievable in experiments roughly 12 times larger than the ones being criticized. The critique’s implicit null — that real effects should replicate at p < .05 three times in a row — does not match the statistical reality of the studies it was applied to.
7. What successful replication actually looks like: the autoganzfeld
After the 1985 Hyman-Honorton debate, the two researchers held lunch together at the 1986 meeting of the Parapsychological Association and co-authored a Joint Communique outlining stricter methodological standards for future ganzfeld experiments. The Communique included this passage — co-signed by the leading skeptic and the leading proponent:
We agree that there is an overall significant effect in this data base that cannot reasonably be explained by selective reporting or multiple analysis. We continue to differ over the degree to which the effect constitutes evidence for psi, but we agree that the final verdict awaits the outcome of future experiments conducted by a broader range of investigators and according to more stringent standards.
Two further passages from the same Communique are essential context. First: Hyman and Honorton noted that “the old database departs from ideal standards in ways that could have spuriously produced the obtained results,” and that the relationship between flaws and outcomes “supports no firm conclusion at the present time.” Second: the Communique adopted the convention of using “psi” as a neutral label for an unexplained communications anomaly — not as an affirmation of any particular paranormal mechanism. The Communique is properly read as a methodological road-map agreed to between the leading skeptic and the leading proponent, with the substantive verdict explicitly deferred to future replication work, not as a settled finding.
The methodological standards specified in the Communique anticipated, by approximately three decades, the post-2010 “replication crisis” reforms in mainstream psychology: pre-registration of analysis plans, automated procedural execution, audited randomization, mandatory reporting of all conducted trials regardless of outcome, and explicit constraints on multiple testing.
The autoganzfeld experiments that followed — designed to satisfy the methodological standards specified in the Joint Communique — produced the following results across 11 experimental series, all conducted within Honorton’s Psychophysical Research Laboratories (PRL), by 8 experimenters working under one principal investigator:
| Outcome | Result |
|---|---|
| Total trials | 355 |
| Direct hits | 122 (34.4%) |
| Chance hit rate | 25.0% |
| Statistical significance | p = 0.00005 |
| Effect size (Cohen’s h) | 0.20 |
| 95% CI on hit rate | 0.30 – 0.39 |
| Series with positive effect size | 10 of 11 |
| Experimenters with positive effect size | 8 of 8 |
| File drawer (audited period) | No unreported trials during the reviewed PRL period; trial counts planned in advance for most series (two pilot series excepted) |
The autoganzfeld results — produced under methodological standards developed jointly with the leading critic — were statistically significant at p = 0.00005, with positive effect sizes in 10 of 11 series and 8 of 8 experimenters within Honorton’s PRL laboratory. Trial counts were planned in advance for most but not all series (two pilot series were unfinished); during the reviewed PRL period, the Parapsychological Association’s full-reporting policy was honored. Hyman’s subsequent assessment was that the autoganzfeld met “most, but not all” of the stringent standards specified in the Joint Communique, and he flagged two specific predicted effects — static-target performance and sender-friend performance — that failed despite adequate power. The 8/8 experimenter-positive and 10/11 series-positive figures should be read as a within-laboratory consistency result, not as eight independent replications from independent investigators. The effect size — Cohen’s h = 0.20 — is approximately three times that of the famous 1987 aspirin/heart-attack trial (h ≈ 0.068; Utts 1991, p. 374), a finding so consequential that the Steering Committee of the Physicians’ Health Study terminated the trial early on ethical grounds (1988)[13] — withholding aspirin from the placebo arm was no longer defensible once the mortality benefit was clear. The Type II implication is direct: a “small” standardized effect was considered overwhelming evidence when the outcome mattered. The autoganzfeld effect at three times that magnitude cannot be dismissed on standardized-size grounds alone — what determines real-world significance is what the outcome is, not Cohen’s benchmark categories. The limitations of effect-size-only comparisons are nonetheless discussed in Section 11.
The autoganzfeld result was followed by an independent meta-analysis from Milton & Wiseman (1999)[14] covering 30 ganzfeld studies conducted between 1987 and 1997 across seven independent laboratories under the same Joint Communique standards. Their analysis returned Stouffer z = 0.70, p = 0.24, mean effect size d = 0.013 — a clear failure to confirm the autoganzfeld effect under broader independent replication. Storm, Tressoldi & Di Risio (2010)[15] meta-analyzed a broader free-response set covering 29 ganzfeld studies from 1997 to 2008 under revised inclusion criteria, returning Stouffer z = 5.48, p = 2.13 × 10⁻⁸, mean d = 0.142 — a directly opposite conclusion. The same statistical method (Stouffer combined-z) applied to different study-inclusion sets thus produces both p = 0.24 and p < 10⁻⁷ on largely overlapping data; the divergence reflects criteria choices, not different statistics. The disagreement over ganzfeld replicability since 1999 is not resolved; the page’s editorial commitment to “name nulls explicitly” requires Milton-Wiseman to be reported at full strength alongside the Storm rebuttal.
8. The mainstream-science context
The pattern documented above — under-powered individual studies producing apparent replication failures of real effects — is not unique to parapsychology. The 2010s mainstream “replication crisis” in psychology hit exactly the same wall.
The Open Science Collaboration (2015)[16], which attempted to replicate 100 high-profile psychology findings, found that replication effect sizes averaged approximately half the magnitude of the original studies, with only 36% of replications reaching statistical significance (and 39% subjectively rated by the replication teams as having successfully replicated the original). Both findings are load-bearing. The effect-size shrinkage is the substantive result — replicated effects exist but are systematically smaller than originally reported, consistent with winner’s curse plus genuine effect heterogeneity. Magnitude shrinkage at the ~50% scale documented by OSC is, in fact, a stronger indicator of original-study inflation than of replication-only underpowering: power alone changes the rate of significance but not the magnitude of point estimates in replications. The Bem (2011) case, where preregistered replications at 99.92% power for the original effect size returned near-null results, is the cleanest single-paradigm demonstration of this pattern; it cannot be rescued by a power-based interpretation alone. The 39% significance rate alone would be statistically consistent with the original studies’ mean power (approximately 0.50) — but the OSC’s principal finding is the magnitude shrinkage, not the significance count. The same misinterpretation Tversky-Kahneman documented among individual psychologists in 1982 was repeated at the discipline level in 2015: media coverage focused on the significance vote-count and read it as “psychology is broken,” when the underlying data showed that effects exist but are smaller and noisier than originally reported.
The methodological reforms that followed the OSC 2015 report — pre-registration, registered reports, increased sample sizes, full reporting — have a documented precedent, in substantially similar form, in the 1986 Hyman-Honorton Joint Communique. The Communique specified pre-registration, full reporting, automated procedural execution, and explicit constraints on multiple testing. Some parapsychology subfields adopted elements of these standards — most prominently the post-Communique ganzfeld work that produced the autoganzfeld. Many other parapsychology research programs did not. The fieldwide claim “parapsychology has been operating by post-2010 mainstream-reform standards” would be too strong; the narrower documented claim — that one specific reform agenda had a parapsychology precedent — is what the historical record supports.
The same pattern appears in medical research. Gardner & Altman (1986)[18], writing in the British Medical Journal, argued that vote-counting and the use of p-values “to define two alternative outcomes — significant and not significant — is not helpful and encourages lazy thinking.” They advocated confidence intervals and effect-size reporting — which is precisely what meta-analytic synthesis in parapsychology has been doing since the late 1980s.
9. The pattern across paradigms
The ganzfeld experiments are a single paradigm. The methodological argument above applies to any contested behavioral-science finding. Where multiple paradigms have been independently meta-analyzed in parapsychology, the meta-analytic pattern is similar across paradigms: small positive pooled effects across thousands of studies, multiple investigators, and multiple decades. The post-meta-analysis direct-replication evidence is more mixed (see the direct-replication-nulls paragraph below the table), and the strength of the meta-analytic pattern’s evidential implications is contested.
| Paradigm | Studies | Subjects/Trials | Combined evidence | Quality-vs-outcome |
|---|---|---|---|---|
| Forced-choice precognition (Honorton & Ferrari 1989)[19] |
309 | 62 senior authors, >50,000 subjects, ~2M trials | z = 11.41, p = 6.3 × 10⁻²⁵ | r = 0.081 (positive — better quality, slightly larger effect) |
| Random number generator (micro-PK) (Radin & Nelson 1989)[20] |
832 | ~600,000 trials experimental + 235 control studies | z = 4.1 (experimental vs control) | slope = +2.5 × 10⁻⁵ (SE 3.2 × 10⁻⁵; essentially zero) |
| Dice-influence (Radin & Ferrari 1991)[21] |
148 | ~2.5M dice tosses | z = 18.2 (exp) vs 0.18 (control) | Fail-safe N = 17,974; quality-effect slope: negative (p = .004) — better-rated studies produced smaller effects (Radin & Ferrari 1991) |
| Ganzfeld autoganzfeld (Honorton et al. 1990)[22] |
11 series | 355 trials, 241 participants | p = 0.00005, Cohen’s h = 0.20 | 10/11 series, 8/8 experimenters positive |
Winner’s-curse / Type M inflation on positive estimates. The same low-power conditions that make individual nulls statistically uninformative also inflate the effect sizes of significant findings via selection (Gelman & Carlin 2014 on Type M / “magnitude” error). The page’s symmetric principle requires this be acknowledged on the positive side: if power = 0.18 to 0.50 is the typical regime, the meta-analytic effect sizes are themselves built from individual significant findings whose magnitudes are upward-biased. Decline-effect patterns observed across the ganzfeld trajectory (pre-autoganzfeld h ≈ 0.32 → autoganzfeld h = 0.20 → Milton-Wiseman d = 0.013) and the Bem case (preregistered failure at 99.92% power) are statistically consistent with this Type M inflation, although they are also consistent with a real-effect-with-design-noise interpretation; current data does not distinguish these.
Direct-replication nulls in these paradigms. The meta-analytic syntheses above are not the only relevant evidence. Two prominent direct-replication nulls deserve explicit mention: (1) the Milton & Wiseman (1999)[14] 30-study, seven-laboratory ganzfeld meta-analysis (Stouffer z = 0.70, p = 0.24, mean d = 0.013), failing to confirm the autoganzfeld effect under broader independent replication; (2) Ritchie, Wiseman & French (2012)[23], who conducted three preregistered direct replications of Bem’s (2011)[17] forced-choice precognition study with combined statistical power of 99.92% for Bem’s original effect size, returning combined p = 0.83 — a clear failure to replicate. Broader-inclusion ganzfeld meta-analyses (e.g., Storm, Tressoldi & Di Risio 2010[15]) returned smaller positive pooled effects; the Bem replication disagreement remains active. The convergent meta-analytic pattern across paradigms is real; the post-1999 ganzfeld and post-Bem replication evidence is genuinely mixed.
Key features of this pattern:
- Effects are small but consistent. No single study is overwhelming. The pattern emerges from synthesis across many studies.
- Quality of studies does not consistently correlate with reduced effect size, but the picture is mixed. The Radin & Nelson (1989) RNG meta-analysis found a near-zero within-corpus quality-effect-size correlation (slope = −2.5 × 10⁻⁵), inconsistent with the simple prediction that worse-rated studies produce larger effects. However, Radin & Ferrari (1991), in their own dice-influence meta-analysis, reported a significant negative quality-effect-size relationship (weighted regression p = .004) and characterized the pre-1975 dice database as suspect with respect to selective reporting. DMILS analyses have similarly reported negative quality-effect relationships in some surveys. A near-zero within-corpus quality slope rules out one specific form of the flaw hypothesis — the prediction that worse-rated studies produce larger effects — but the quality metrics typically used (e.g., Radin & Nelson’s 16-criterion checklist) assess design-transparency features: randomization documentation, reporting completeness, procedural clarity. They do NOT assess for the systematic methodological pathways Hyman has most consistently flagged — experimenter-expectancy effects, within-session optional stopping, target-pool non-uniformity, selective channel reporting, and analytic flexibility. A near-zero slope on metrics that don’t capture the relevant failure modes does not rule out those failure modes; it rules out flaws as captured by the metric. Modern selection-model meta-analytic tools (p-curve, three-parameter selection models, PET-PEESE) provide stronger tests of flaw-driven inflation and have not been broadly applied to parapsychology meta-analyses, which constitutes a known methods-modernization gap.
- File-drawer effects are implausible at the magnitudes required, under standard fail-safe-N assumptions. The Radin & Ferrari (1991) dice-influence meta-analysis would require 17,974 unreported null studies to negate the observed effect under Rosenthal’s fail-safe formula — roughly 121 unreported studies per retrieved study. Fail-safe N is, however, a blunt and assumption-heavy diagnostic: it assumes unpublished nulls are drawn from the same distribution as published results. Radin & Ferrari themselves flagged the pre-1975 dice database as suspect with respect to selective reporting. If positive bias operates at the design level (researcher degrees of freedom, optional stopping) rather than via withheld null studies, fail-safe N underestimates the bias needed. The 17,974 figure is informative, but is not a substitute for modern publication-bias modeling.
- Effect sizes are comparable in standardized magnitude to landmark medical findings. The forced-choice precognition effect size is comparable, on Cohen’s standardized scale, to the 1987 aspirin/heart-attack trial that was terminated early because its result was considered so dramatic. Standardized-magnitude comparability does not imply evidentiary parity — the aspirin finding had a confirmed biological mechanism, multiple independent RCTs, and high prior plausibility, none of which apply to psi (see the effect-size-ratio caveat box in Section 11). Practical significance is a separate question from effect-size magnitude.
- Multi-author reproduction is well documented in some paradigms but contested in others. The forced-choice precognition meta-analysis (Honorton & Ferrari 1989) includes 62 senior authors at multiple institutions — strong author-diversity evidence for that paradigm. In the ganzfeld paradigm, by contrast, the post-1986 broader-laboratory replication record is mixed (Milton-Wiseman 1999 null; Storm-Tressoldi-Di Risio 2010 smaller positive). The pattern as a whole is real across paradigm meta-analyses; per-paradigm independent-replication strength varies.
Engagement with the major critic objections
The cross-paradigm pattern above is the strongest statistical case the published parapsychology literature supports. Several named methodologists have advanced substantive objections to it. Each deserves engagement on the merits.
Ray Hyman. The 11 autoganzfeld series are not 11 independent investigations — they are one laboratory’s work under one director. Hyman’s specific objection is that the systematic methodological flaws that concern him operate at the design level (procedural drift, optional stopping, shared materials, unblinded feedback) and are invisible to within-corpus quality scoring. The response on this page: Hyman is correct that within-corpus quality-effect nulls cannot address shared design-level bias. What they CAN address is the simpler hypothesis that worse-rated studies produce larger effects — which is the prediction the within-corpus slope tests. The shared-bias hypothesis requires different evidence: adversarial replication by skeptic-led laboratories, which the post-1986 ganzfeld history partially provides (Milton-Wiseman 1999).
Richard Wiseman & Julie Milton. Post-autoganzfeld broader replication failed (1999, z = 0.70, p = 0.24, d = 0.013). The response on this page: Milton-Wiseman is now explicitly engaged in Sections 7 and 9 above. The post-1999 ganzfeld literature has not converged on a single answer; broader-inclusion meta-analyses returned smaller positive pooled effects, and narrower-inclusion confirmation has not been achieved at autoganzfeld magnitudes. This is unresolved evidence, and the page reports it as such.
Eric-Jan Wagenmakers. Frequentist p-values overstate evidence in low-prior-probability domains. Wagenmakers et al. (2011)[24] showed that small p-values can correspond to weak Bayes factors when the alternative hypothesis is diffuse. The response on this page: the meta-analytic combined-z figures (e.g., z = 11.41 for forced-choice precognition) are statistically informative but should not be read as Bayes factors. Confirmatory designs with pre-specified analyses and proper Bayesian treatment of priors are the appropriate next-generation evidential standard for any contested claim. Parapsychology has begun adopting this with preregistration; full Bayesian re-analyses of existing databases remain incomplete.
Dorothy Bishop. Stricter p-thresholds without preregistration do not by themselves establish methodological rigor — a motivated investigator can run trials until p < .01. The response on this page: Bishop is correct on the principle. The historical-threshold documentation in Section 10 below is offered as factual correction to a specific public misconception (“parapsychology operates by looser standards than mainstream behavioral science”), not as evidence of discipline-wide rigor. Preregistered, audited execution is the modern criterion; the 1986 Joint Communique standards anticipated this, and the autoganzfeld is its earliest implementation in the field.
Stuart Ritchie. Psi follows the same shrinkage pattern as the rest of low-power behavioral science — initial high effect sizes, modest replications, near-null in preregistered direct replications. Ritchie, Wiseman & French (2012)[23] preregistered three close replications of Bem (2011)[17] with combined power 99.92% for the original effect size; they returned combined p = 0.83. The response on this page: this is reported in Section 9 above. The shrinkage pattern in the ganzfeld (autoganzfeld h = 0.20 followed by Milton-Wiseman d = 0.013) and the Bem case (preregistered failure) are real and inconsistent with a strong-effect interpretation of the underlying claim. They are also consistent with a small-effect-with-methodological-noise interpretation; the available data does not distinguish these.
The cross-paradigm pattern described in this section is therefore best read as: published meta-analytic models report small positive deviations across several paradigms, with genuinely mixed independent-replication evidence; whether these deviations represent psi, residual artifact, selective reporting, analytic flexibility, or some combination remains unresolved. The methodological case against treating “failed replications” as automatic disproof is independent of this resolution and remains sound regardless of how the underlying question resolves.
10. Parapsychology demanded stricter thresholds than mainstream psychology
One source of confusion in the public conversation is the assumption that parapsychology operates by looser methodological standards than mainstream behavioral science. On significance thresholds specifically, the historical record runs the other direction.
The conventional p < 0.05 threshold has casual origins. It traces to a passage by R.A. Fisher[25] — the founder of modern statistical inference — in a 1926 paper on agricultural field experiments:
It is convenient to draw the line at about the level at which we can say: “Either there is something in the treatment, or a coincidence has occurred such as does not occur more than once in twenty trials.” … If one in twenty does not seem high enough odds, we may, if we prefer it, draw the line at one in fifty (the 2 per cent point), or one in a hundred (the 1 per cent point). Personally, the writer prefers to set a low standard of significance at the 5 per cent point, and ignore entirely all results which fail to reach that level.
This is the entire methodological foundation of the .05 convention. Fisher offered it as a personal preference, in a paper on field-trial planning, not as a discipline-wide standard. Mainstream psychology adopted it by convention after 1940.
Parapsychology, however, used stricter thresholds for most of its history. The documentary evidence comes from the Journal of Parapsychology‘s own glossary of statistical terms:
| Year | Source | Threshold required |
|---|---|---|
| 1917 | Coover (Stanford), Experiments in Psychical Research | [26]p ≤ 2.21 × 10⁻⁵ |
| 1940 | Rhine et al., Extra-Sensory Perception After Sixty Years | [27]p ≤ .0062 (z ≥ 2.5) |
| 1949 | Journal of Parapsychology glossary | p ≤ .01 |
| 1957 | Rhine & Pratt, Parapsychology: Frontier Science of the Mind | p ≤ .01 |
| 1968–1986 | Journal of Parapsychology glossary | p ≤ .02 (with p ≤ .05 explicitly labeled only “strongly suggestive”) |
For most of its history, parapsychology used between 2.5x and 8x stricter significance thresholds than mainstream behavioral science. The methodological reforms in the 1986 Hyman-Honorton Joint Communique — pre-registration, full reporting, automated procedural execution — anticipated the post-2010 mainstream-psychology reform agenda by approximately three decades. Stricter significance thresholds and procedural-reform priority do not by themselves establish discipline-wide methodological superiority. Significance-threshold control (alpha) reduces Type I error rate but does not address the modern methodological agenda — preregistration, blinding, independent randomization audit, full pre-specification of analysis pipelines, public data archiving, independent verification of raw data, and modern publication-bias modeling. A motivated investigator using a strict alpha can still produce inflated effects via researcher degrees of freedom or selective reporting. The historical-threshold evidence contradicts only the specific public assumption that parapsychology operated under looser significance thresholds than mainstream behavioral science; on that narrower point, the documented record is the opposite.
11. Effect size and practical significance
A common follow-up objection to the meta-analytic evidence presented above is: even if the effect is real, the effect sizes are very small. So what does it matter? This is a fair question. Effect sizes in parapsychology — particularly in forced-choice and RNG paradigms — are indeed small. Cohen’s h of 0.20 (autoganzfeld) is small by Cohen’s own benchmarks. RNG effects are smaller still — around 10⁻⁴ to 10⁻³.
Three responses:
First, small effects are statistically informative when they replicate across thousands of trials. A single small study with effect size 0.0002 has essentially no individual evidential weight. But the same effect size sustained across 2 million trials (the Honorton-Ferrari 1989 forced-choice precognition database) produces combined z = 11.41 — a result that exceeds, by orders of magnitude, the statistical evidence underlying many accepted scientific findings.
Second, parapsychology effect sizes are comparable in standardized magnitude to landmark medical findings. The 1987 Steering Committee of the Physicians’ Health Study terminated their aspirin-vs-heart-attack trial early — at chi-square = 25.01, p < 0.00001 — because the effect was considered too dramatic to ethically withhold from the placebo group. The Cohen’s h for that aspirin effect was approximately 0.068 (Utts 1991, p. 374). The autoganzfeld effect size is 0.20 — approximately three times larger on the standardized h scale.
Third, practical significance is a separate question from existence. Whether the effect is “useful” or “important” is a different inquiry than whether it exists. ESP-Nexus’s position is that the existence question deserves an honest answer based on the evidentiary record, and that practical implications can be evaluated separately once the existence question is settled to the standards of behavioral-science evidence.
12. What replication actually demonstrates
The argument of this page can be summarized in seven steps:
- Replication does not mean identical significant results every time. This is a widespread misconception held even by trained psychologists. Tversky & Kahneman (1982) demonstrated this empirically.
- “Failed replications” can mean either no underlying effect OR an underpowered study. A single null result by itself does not distinguish between these. Power analysis is required. Berger’s (1989)[1] review of the Blackmore psi-experiment database documented this pattern directly: Table 2 of that paper lists per-experiment Cohen’s d values, with multiple positive effect sizes — including d = +1.119 in the first Tarot experiment (Berger 1989, p. 132) — that nonetheless individually missed conventional significance because of underpowered N. Berger called for a serious meta-analysis with weighted flaw-correction; pending that, neither a positive nor a negative conclusion about psi could be drawn from those studies (Berger 1989, p. 141).
- For real effects at modest power, partial replication is the statistically expected outcome. At 50% power, getting 2-of-3 successful replications is the median expected result. Three-of-three is improbable.
- Vote-counting individual significant/nonsignificant studies is methodologically inadequate. Hedges & Olkin (1985) showed mathematically that vote-counting becomes increasingly likely to make the wrong decision as the number of studies grows. Meta-analytic synthesis — combining trial-level data — is the appropriate evidentiary frame for low-power literatures.
- The mainstream skeptical demand for “three significant replications in a row” is statistically incoherent unless paired with extremely high per-study power. Hansel (1980) is the canonical example of this argument; Utts (1991) refuted it on the same page where it appears.
- Meta-analytic syntheses of parapsychology research show small but nonzero pooled effects across multiple paradigms. Forced-choice precognition, ganzfeld, RNG/micro-PK, and dice-influence all show convergent positive pooled effects of small magnitude. Quality-effect-size correlations are mixed across paradigms (near-zero in RNG meta-analyses, negative in dice-influence and DMILS). File-drawer estimates are large but not dispositive against design-level positive bias. Independent direct-replication evidence is mixed: Milton-Wiseman (1999) failed to confirm the ganzfeld effect under broader replication, and Ritchie-Wiseman-French (2012) failed to replicate Bem (2011) at preregistered 99.92% power, while broader-inclusion meta-analyses returned smaller positive pooled effects. The cross-paradigm pattern is real; the strength of its evidential implications is contested.
- Some parapsychology venues adopted stricter significance thresholds and procedural-reform priorities decades before mainstream psychology did so. Documentary record: significance thresholds 2.5x to 8x stricter than the .05 convention from 1917 through 1986, full reporting via PA policy from the 1980s, and the 1986 Joint Communique standards (pre-registration, automated execution, full reporting) anticipated the post-2010 mainstream-psychology reform agenda. This is a narrower factual claim than discipline-wide methodological superiority, but it inverts the public assumption that parapsychology has operated under looser standards.
None of this proves that psi phenomena exist. What it demonstrates is that the standard skeptical argument — “the effect failed to replicate, therefore it isn’t real” — is methodologically unsound when applied to studies of the type and size that constitute most of the parapsychology literature. The argument fails on its own statistical premises, independent of any conclusion about whether psi is real.
Honest evaluation of any contested behavioral-science finding requires:
- Meta-analytic synthesis across the available studies, not vote-counts
- Effect-size reporting and confidence intervals, not p-value verdicts
- Power analysis context for individual null results, not bare significance counts
- Attention to file-drawer estimates and quality-vs-outcome correlations
- Joint consideration of all claim-levels in any given study (a single study often addresses multiple distinct hypotheses)
ESP-Nexus applies these standards uniformly across the research it summarizes. The methodological argument on this page is the canonical reference for that editorial discipline — it is cross-linked from every scientist hub and phenomenology spoke where mixed-result literatures are presented.
13. Beyond the statistical argument
The argument above is statistical and methodological. It is the load-bearing case against vote-counting replication and against Hansel-style three-of-three demands. But the statistical argument is one of two argument families needed to evaluate psi research fairly. Psi is a behavioral phenomenon — the science is the behavioral kind — and behavioral science has principles about trait distribution, measurement design, subject selection, experimenter effects, and epistemic posture that the statistical argument by itself does not engage. The deep-dive spokes below cover what the statistical argument leaves out. Both families are needed. If psi is real, dismissing it via either family-incomplete framing is the Type-II failure mode this page exists to document.
Behavioral Science and Psi Replication
The Leonardo and shoe-size analogies for trait-distributed phenomena; the Bem & Honorton (1994) four-predictor framework (extraversion, creative-arts background, prior psi experience, belief); Schmeidler-Lawrence sheep-goat effect; Tellegen Absorption and hypnotic-susceptibility correlates; state-dependence of psi task performance (Cardeña 2018 in American Psychologist). The argument: small mean effect sizes in unselected-sample meta-analyses are floor estimates of a trait-distributed phenomenon, not phenomenon-absence proofs.
Measurement and Methodology
Qualitative data discarded by binary scoring; the sensory-leakage paradox in forced-choice precognition (zero leakage opportunity by design, yet effect sizes comparable to ganzfeld); experimenter effects (the Wiseman-Schlitz 1997/1999/2006 collaboration sequence, with the 2006 2×2 greeter/sender cross-over yielding a failed replication); definitional scope ambiguity across telepathy / clairvoyance / precognition / PK.
Epistemic Frame (forthcoming spoke)
The zetetic vs pseudo-skeptical distinction (Truzzi 1987, 1998) and what it means for evaluating contested evidence; comparative rigor with mainstream subtle-effects research; adversarial collaboration as the appropriate epistemic frame for genuinely contested findings.
Related topics on ESP-Nexus
Sources
- Berger, R. E. (1989). Discussion: A critical examination of the Blackmore psi experiments. Journal of the American Society for Psychical Research, 83 (April), 123–144. (The canonical published source establishing the Type-II / asymmetric-standards argument in parapsychology.) R001 [Berger 1989] ↩︎
- Tversky, A., & Kahneman, D. (1982). Belief in the law of small numbers. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment Under Uncertainty: Heuristics and Biases. Cambridge University Press. R002 [Tversky & Kahneman 1982] ↩︎
- Utts, J. (1988). Successful replication versus statistical significance. Journal of Parapsychology, 52, 305–319. https://www.ics.uci.edu/~jutts/Utts-Successful-Replication.pdf R003 [Utts 1988] ↩︎
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates. R004 [Cohen 1988] ↩︎
- Cohen, J. (1990). Things I have learned (so far). American Psychologist, 45(12), 1304–1312. https://psycnet.apa.org/doi/10.1037/0003-066X.45.12.1304 R005 [Cohen 1990] ↩︎
- Utts, J. (1991). Replication and meta-analysis in parapsychology. Statistical Science, 6(4), 363–403. https://projecteuclid.org/journals/statistical-science/volume-6/issue-4/Replication-and-Meta-Analysis-in-Parapsychology/10.1214/ss/1177011577.full R006 [Utts 1991] ↩︎
- Rosenthal, R. (1990). How are we doing in soft psychology? American Psychologist, 45, 775–777. https://psycnet.apa.org/doi/10.1037/0003-066X.45.6.775 R007 [Rosenthal 1990] ↩︎
- Utts, J. (1986). The ganzfeld debate: A statistician’s perspective. Journal of Parapsychology, 50, 393–402. R008 [Utts 1986] ↩︎
- Hedges, L. V., & Olkin, I. (1985). Statistical Methods for Meta-Analysis. Academic Press. R009 [Hedges & Olkin 1985] ↩︎
- Rosenthal, R. (1986). Meta-analytic procedures and the nature of replication: The ganzfeld debate. Journal of Parapsychology, 50, 315–336. R010 [Rosenthal 1986] ↩︎
- Hansel, C. E. M. (1980). ESP and Parapsychology: A Critical Re-evaluation. Prometheus Books. R011 [Hansel 1980] ↩︎
- Hyman, R., & Honorton, C. (1986). Joint communiqué: The psi ganzfeld controversy. Journal of Parapsychology, 50, 351–364. R012 [Hyman & Honorton 1986] ↩︎
- Steering Committee of the Physicians’ Health Study Research Group. (1988). Preliminary report: findings from the aspirin component of the ongoing Physicians’ Health Study. New England Journal of Medicine, 318(4), 262–264. https://doi.org/10.1056/NEJM198801283180431 R013 [Physicians’ Health Study 1988] ↩︎
- Milton, J., & Wiseman, R. (1999). Does psi exist? Lack of replication of an anomalous process of information transfer. Psychological Bulletin, 125(4), 387–391. https://doi.org/10.1037/0033-2909.125.4.387 R014 [Milton & Wiseman 1999] ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992–2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136(4), 471–485. https://doi.org/10.1037/a0019457 R015 [Storm et al. 2010] ↩︎
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://www.science.org/doi/10.1126/science.aac4716 R016 [Open Science Collaboration 2015] ↩︎
- Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524 R017 [Bem 2011] ↩︎
- Gardner, M. J., & Altman, D. G. (1986). Confidence intervals rather than p-values: estimation rather than hypothesis testing. British Medical Journal, 292, 746–750. https://www.bmj.com/content/292/6522/746 R018 [Gardner & Altman 1986] ↩︎
- Honorton, C., & Ferrari, D. C. (1989). “Future telling”: A meta-analysis of forced-choice precognition experiments, 1935–1987. Journal of Parapsychology, 53, 281–308. R019 [Honorton & Ferrari 1989] ↩︎
- Radin, D. I., & Nelson, R. D. (1989). Evidence for consciousness-related anomalies in random physical systems. Foundations of Physics, 19, 1499–1514. https://doi.org/10.1007/BF00732509 R020 [Radin & Nelson 1989] ↩︎
- Radin, D. I., & Ferrari, D. C. (1991). Effects of consciousness on the fall of dice: A meta-analysis. Journal of Scientific Exploration, 5, 61–83. https://www.scientificexploration.org/docs/5/jse_5_1_radin.pdf R021 [Radin & Ferrari 1991] ↩︎
- Honorton, C., Berger, R. E., Varvoglis, M. P., Quant, M., Derr, P., Schechter, E. I., & Ferrari, D. C. (1990). Psi communication in the ganzfeld: Experiments with an automated testing system and a comparison with a meta-analysis of earlier studies. Journal of Parapsychology, 54, 99–139. R022 [Honorton et al. 1990] ↩︎
- Ritchie, S. J., Wiseman, R., & French, C. C. (2012). Failing the future: Three unsuccessful attempts to replicate Bem’s “retroactive facilitation of recall” effect. PLOS ONE, 7(3), e33423. https://doi.org/10.1371/journal.pone.0033423 R023 [Ritchie et al. 2012] ↩︎
- Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2011). Why psychologists must change the way they analyze their data: The case of psi. Journal of Personality and Social Psychology, 100(3), 426–432. https://doi.org/10.1037/a0022790 R024 [Wagenmakers et al. 2011] ↩︎
- Fisher, R. A. (1926). The arrangement of field experiments. Journal of the Ministry of Agriculture of Great Britain, 33, 503–513. R025 [Fisher 1926] ↩︎
- Coover, J. E. (1917). Experiments in Psychical Research at Leland Stanford Junior University. Stanford University Press. https://archive.org/details/experimentsinpsy01coov R026 [Coover 1917] ↩︎
- Rhine, J. B., Pratt, J. G., Stuart, C. E., Smith, B. M., & Greenwood, J. A. (1940). Extra-Sensory Perception After Sixty Years. Bruce Humphries. R027 [Rhine et al. 1940] ↩︎
Further reading. Landmark work in this literature not cited inline above:
- Utts, J. (1996). An assessment of the evidence for psychic functioning. Journal of Scientific Exploration, 10(1), 3–30. PDF
Deeper dives — Replication in Psi Research: