Tressoldi & Storm (2024)
Stage 2 Registered Report: Anomalous perception in a Ganzfeld condition – A meta-analysis of more than 40 years’ investigation
Tressoldi, P. E., & Storm, L. (2024). Stage 2 Registered Report: Anomalous perception in a Ganzfeld condition – A meta-analysis of more than 40 years’ investigation [version 4; peer review: 2 approved, 1 not approved]. F1000Research, 10, 234. https://doi.org/10.12688/f1000research.51746.4
AI Assessment
A preregistered (Stage-2 Registered Report) meta-analysis of the entire ganzfeld literature from 1974 to 2020: 78 studies, 113 effect sizes, 46 principal investigators. It reports a small but statistically robust overall effect (Hedges’ g of about .08), matched almost exactly by frequentist and Bayesian random-effects models, that survived four publication-bias tests and showed no decline across four decades. Its real strength is procedural: the analysis plan was locked and peer-approved before the data were pooled, the full database and code are open, and null studies are included. Its honest limits are that individual-study effects are tiny and mostly non-significant (median power .088), the effect is dominated by telepathy-style Type-3 designs, and the meta-analysis inherits whatever methodological quality the primary studies had. This audit describes what the paper reports and how it was conducted; it takes no position on whether anomalous perception is real.
Provenance
DOI. 10.12688/f1000research.51746.4 · F1000Research 2024, 10:234 (version 4, published 24 May 2024; first published 24 March 2021). Open access under CC-BY 4.0.
Study type. A Stage-2 Registered Report meta-analysis. All confirmatory analyses were approved at Stage 1 (Tressoldi & Storm, 2021, F1000Research 9:826) before the pooled results were computed; analyses added afterward are reported separately as exploratory.1
Authors. Patrizio E. Tressoldi (Studium Patavinum, University of Padua, Italy) and Lance Storm (School of Psychology, University of Adelaide, Australia). No competing interests and no grant funding were declared.
Scope. Studies of anomalous perception in a ganzfeld condition published between January 1974 and December 2020 inclusive, peer-reviewed and non-peer-reviewed (proceedings), dissertations excluded.
Data availability. The complete database (GZMADatabase1974_2020), power calculations, reference list, R syntax, and forest / cumulative / sequential-Bayes-factor figures are openly deposited on Figshare (doi.org/10.6084/m9.figshare.12674618.v13) under CC-BY 4.0.
Source basis. Every figure below is taken verbatim from the article’s own reported statistics (Abstract, Tables 1–4, and Results text). One transparency caveat, noted by the article itself: the per-study effect-size table is not printed in the article body; it lives in the Figshare underlying data, and the forest / cumulative / sequential-Bayes plots (Figures S1–S3) are supplementary. This audit therefore verifies the article’s summary statistics, not each of the 113 individual effect sizes.
What the paper reports
The ganzfeld procedure places a percipient in a homogeneous sensory field (halved ping-pong balls or diffusing glasses under red light over the eyes, white or pink noise through headphones) and asks them to describe mental imagery that might correspond to a randomly selected target hidden among three or four decoys.1 The paper distinguishes three designs: Type 1 (precognition, target chosen after judging), Type 2 (clairvoyance, target chosen before the ganzfeld phase), and Type 3 (telepathy, target chosen beforehand and viewed by a distant sender). The authors set out to pool every available ganzfeld study from 1974 to 2020 with more advanced statistics than earlier reviews, and to test whether participant type and task type moderate the effect.
The overall picture emerging from this meta-analysis is that there is sufficient evidence to claim that it is possible to observe a non conventional (anomalous) perception in a Ganzfeld environment. The available evidence does not seem to be contaminated by publication bias or questionable research practices.
How it was run
- Reporting standards. The study follows the APA Meta-Analysis Reporting Standard and PRISMA-P; the confirmatory analysis plan was approved at Stage 1 of the Registered Report.
- Retrieval. Studies were drawn from prior ganzfeld meta-analyses plus a Google Scholar, PubMed, and Scopus search for “ganzfeld” in the title or abstract, 1974 to 2020.
- Inclusion criteria. Human participants only; more than two participants per study; target selection randomized by RNG or random-number table and not manipulable by experimenter or participant; enough reported data (trials and hits) to compute a hit rate and effect size; peer-reviewed and proceedings studies both eligible; dissertations excluded.
- Coding. One author coded authors, year, trials, hits, choices per trial, task type, participant type (selected vs. non-selected), and peer-review level; the second author independently checked the studies and discrepancies were resolved against the original papers.
- Effect size. The binomial Z score divided by the square root of the number of trials, then transformed to Hedges’ g to correct small-sample overestimation.
- Pooling. Both a frequentist random-effects model (REML with the Knapp–Hartung adjustment, metafor package) and a Bayesian random-effects model (MetaBMA; a positive-constrained normal prior, mean 0.1, SD 0.03) were run to test robustness, plus robust median and mode estimators and an influence-function outlier check.
- Bias and trend. Publication bias was assessed with four methods (3PSM, p-uniform*, the Mathur & VanderWeele sensitivity analysis, and RoBMA); decline was tested with a cumulative meta-analysis and a meta-regression using year of publication as a covariate.
- Power. Overall power was estimated with the metameta package, and the trials needed for 80% power were computed with G*Power.
Results, as reported
| Metric | Result |
|---|---|
| Database | 78 studies, 113 effect sizes, 46 principal investigators (1974–2020) |
| Overall effect size (frequentist) | Hedges’ g = .074 (95% CI .03–.12), p = .0009 |
| Overall effect size (Bayesian) | g = .084 (95% CrI .05–.12), Bayes factor 89.5 |
| Hit rate above chance | 6.8% (95% CI 4.7–8.9) |
| Heterogeneity | τ² = .03; I² = 63.8 (medium-large) |
| Decline test | meta-regression slope .0012 (95% CI −.002–.005, p = .53); cumulative estimate stable since ~1997, no decline |
| Publication bias (4 tests) | passed all four (p-uniform* .12, 3PSM .15, RoBMA .074); to explain the effect away, significant results would need to be at least 4-fold more likely to be published |
| Moderator: participant type | selected g = .13 (.06–.20) vs non-selected .04 (−.01–.09): almost three-fold |
| Moderator: task type | Type 3 (telepathy) .08, Type 2 (clairvoyance) .04, Type 1 (precognition) .12 (only 5 studies, treat with caution) |
| Moderator: peer-review level | level 1 (proceedings) .073 vs level 2 (journals) .076: no difference |
| Median statistical power | .088; only 30 of 113 (22.5%) individual results were statistically significant |
| Peer-review status (version 4) | 2 reviewers approved, 1 not approved |
Values are reproduced from the article’s Abstract, Results text, and Tables 1–4. The frequentist and Bayesian estimates agree closely and both reject the null with high probability; the effect is small in absolute terms (a hit-rate elevation of under 7 percentage points).
Eleven-dimension audit
Pre-registration
This is the study’s strongest dimension. It is a Stage-2 Registered Report: the analysis plan, inclusion criteria, moderators, and statistical models were peer-reviewed and approved at Stage 1 (Tressoldi & Storm, 2021) before the pooled results were known, and every analysis added afterward is explicitly labelled exploratory (the sequential Bayes-factor trend and the selected-plus-Type-3 combination). That locks the confirmatory/exploratory boundary in a way ordinary meta-analyses cannot.
Randomization
Randomization is enforced at the meta level through an inclusion criterion: a study was eligible only if its target was selected by a true or pseudo-random RNG or random-number table and the procedure could not be manipulated by experimenter or participant. The meta-analysis cannot re-verify the randomization of each primary study beyond what those studies reported, but the criterion excludes manually selected targets by design.
Sensory leakage
Leakage control is likewise handled at the inclusion level and by the paradigm itself: the sender and percipient are isolated, and the research assistant who interacts with the participant is required to remain blind to the target identity until the rating task is complete. As with any meta-analysis, the integrity of that control in each of the 78 studies rests on those studies’ own reporting rather than on independent re-inspection here.
Blinding
The judging process is blind (an independent judge, or a percipient blind to which item is the target among the decoys). The meta-analysis inherits the blinding quality of its constituent studies; it does not grade each study’s blinding, and one reviewer specifically objected that lumping fully peer-reviewed and proceedings studies together may mask quality differences (see the adversarial record).
Optional stopping
Optional stopping is a property of the primary studies, not the meta-analysis. The authors address the broader questionable-research-practices concern directly: they cite a simulation by Bierman, Spottiswoode, and Bijl (2016) on 78 ganzfeld studies showing that such practices could inflate the effect but would not reduce it to zero,5 and they note that specialist journals in this field publish non-significant results, which limits the file-drawer problem.
Outcome measure
Pre-stated and standard for this literature: the binomial Z score over the square root of the number of trials, transformed to Hedges’ g. The same measure was used in the authors’ earlier meta-analyses, making the estimate comparable across reviews. Table 1’s descriptive hit-rate mean (.068) is flagged by the authors as purely descriptive because not all studies used a four-alternative free-choice design.
Effect size
The headline result is a small effect: g = .074 (frequentist) and .084 (Bayesian), corresponding to a hit rate 6.8% above chance. Removing the two influential outliers barely changes it (.078). The absolute size is modest and the authors do not overstate it; the claim rests on the effect’s statistical robustness and consistency, not its magnitude.
Multiple comparisons
Three pre-planned moderators were tested (participant type, task type, peer-review level), which is a contained set. The authors flag the low-powered Type-1 moderator (only 5 studies) as needing caution, and they clearly separate the exploratory analyses (sequential Bayes factor, the selected-plus-Type-3 combination) from the confirmatory ones, limiting the risk of moderator fishing.
Internal replication
The cumulative meta-analysis is the internal-consistency check: the pooled estimate stabilized around the evidence available by roughly 1997 and has stayed stable for more than 20 years, across 46 different principal investigators. That the estimate does not swing as studies accumulate is a stronger internal signal than any single pooled number.
External replication
The result sits within a long line of ganzfeld meta-analyses and broadly agrees with most of them: Honorton’s 1985 review (38% hits vs 25% expected), Bem & Honorton’s 1994 autoganzfeld analysis (32.2%),2 and Storm and colleagues’ 2010 and 2020 analyses (g around .13–.14).3 The notable dissent is Milton & Wiseman’s 1999 near-zero estimate (.013), which a later exact binomial reanalysis of the same trial counts turned into a significant 27% hit rate. This meta-analysis’s smaller estimate (.08) than the two most recent Storm analyses (~.13) is itself worth noting.
Transparency
Strong on openness: the full database, power file, reference list, and R syntax are on Figshare under CC-BY, the Registered Report format makes the plan auditable, and peer review is open and signed (one reviewer, Pavo Orepic, records a “not approved” verdict in the published record). The transparency caveats are that the per-study effect-size table is not in the article body, and the article carries a small internal inconsistency: the Discussion refers to “113 studies” where the Results specify 78 studies yielding 113 effect sizes, and its Abstract calls the evidence “not contaminated by publication bias” while the Discussion hedges more carefully that the analysis is “not immune to publication bias.”
The adversarial record
- Open peer review. As a Registered Report the reviews are published. Version 4 stands at two approvals (Michel-Ange Amorim; Dean Radin) and one “not approved” (Pavo Orepic, University of Geneva). This is an unusually transparent record of live disagreement about the paper.
- The dissenting reviewer. Orepic argued that the title and framing are misleading because the pooled database is dominated by Type-3 telepathy designs (81 of 113 effect sizes, 72%) rather than the ganzfeld phenomenon in general; that including studies that are not fully peer-reviewed weakens the “peer-review level” moderator; and that a Test for Excess Success was not run. The authors did not revise the framing to his satisfaction.
- Reviewer corrections that were accepted. The approving reviewer Amorim caught real errors across versions: spelling errors in the task-type labels (“Precogniyion,” “Clayrvoyance”), a publication-bias sentence that had misstated the sensitivity ratio (corrected to “at least 4-fold”), and a query about the power calculation. The article states that 245 trials are needed for 80% power at the observed effect size; the reviewer independently obtained a similar figure (about 312) under slightly different G*Power settings, a reminder that the exact number depends on the test specification.
- Reading against the study. A committed skeptic can grant the procedural rigor and still argue that a g of .08 with median power .088 means most constituent studies are individually inconclusive, that heterogeneity is substantial (I² = 63.8), and that a meta-analysis cannot repair whatever leakage or randomization weaknesses the primary studies carried. The authors’ own answer is the convergence of frequentist and Bayesian models, the four passed bias tests, and four decades of stability without decline.
Contested record
Sources
- Tressoldi, P. E., & Storm, L. (2024). Stage 2 Registered Report: Anomalous perception in a Ganzfeld condition – A meta-analysis of more than 40 years’ investigation [version 4]. F1000Research, 10, 234. https://doi.org/10.12688/f1000research.51746.4 R001 [Tressoldi & Storm 2024] ↩︎
- Bem, D. J., & Honorton, C. (1994). Does psi exist? Replicable evidence for an anomalous process of information transfer. Psychological Bulletin, 115(1), 4–18. https://doi.org/10.1037/0033-2909.115.1.4 R002 [Bem & Honorton 1994] ↩︎
- Storm, L., & Tressoldi, P. (2020). Meta-analysis of free-response studies 2009–2018: Assessing the noise-reduction model ten years on. Journal of the Society for Psychical Research, 84(4), 193–219. R003 [Storm & Tressoldi 2020] ↩︎
- Hyman, R. (1985). The ganzfeld psi experiment: A critical appraisal. Journal of Parapsychology, 49(1), 3–49. R004 [Hyman 1985] ↩︎
- Bierman, D. J., Spottiswoode, J. P., & Bijl, A. (2016). Testing for questionable research practices in a meta-analysis: An example from experimental parapsychology. PLOS ONE, 11(5), e0153049. https://doi.org/10.1371/journal.pone.0153049 R005 [Bierman et al. 2016] ↩︎
- Cardeña, E. (2018). The experimental evidence for parapsychological phenomena: A review. American Psychologist, 73(5), 663–677. https://doi.org/10.1037/amp0000236 R006 [Cardeña 2018] ↩︎