Meta-Analysis in Parapsychology
Meta-analysis combines effect-size estimates across multiple independent studies to produce a single pooled estimate and examine heterogeneity. Parapsychology has used meta-analytic techniques since the mid-1980s to summarize its most-studied paradigms, ganzfeld ESP, forced-choice precognition, micro-PK RNG experiments, and DMILS, and to test whether apparent effects survive under corrections for publication bias, quality weighting, and pre-registration.
Overview
Meta-analysis is a statistical method for pooling results across many independent studies. Rather than judging one experiment at a time, it combines effect-size estimates, standardized measures of how far a result departs from chance, into a single pooled value, and then tests whether the studies agree among themselves (a property called heterogeneity). In parapsychology the method has been used since the mid-1980s to summarize the field’s most-studied designs: ganzfeld telepathy, forced-choice extrasensory perception (ESP), micro-PK experiments with random number generators (RNGs), and distant mental interaction with living systems (DMILS).
A key point is that meta-analysis is a tool, not a verdict. Positive and negative meta-analyses of the very same paradigm coexist in the literature, and the differences usually trace to which studies were included, how study quality was weighted, and what assumptions the analysts brought. The pooled effects in psi research are also typically small. A small effect is not, by itself, evidence that nothing is there, but it does demand careful handling of publication bias and heterogeneity before any claim is made.
Protocol
A meta-analysis in this field generally follows a fixed sequence. Analysts first define a clear search and inclusion rule, for example, all ganzfeld telepathy trials meeting a stated standard between two dates. Each study is then converted to a common effect-size scale so that experiments of different sizes can be compared. Studies are weighted, often by sample size or by a quality rating, and a pooled estimate is produced together with a test of whether the spread of results is wider than chance would predict.
Recent work has refined this template by separating procedural moderators from the headline estimate. One meta-analysis of telepathy ganzfeld studies, for instance, modeled how features of the sender–receiver setup changed success rates, finding that letting senders hear the receiver’s running commentary raised hit rates while a post-session review period lowered them.[1] Such moderator analysis is now a routine part of the protocol, because a single pooled number can hide systematic differences between procedures.
History
The modern history begins with a 1985 ganzfeld meta-analysis that brought formal effect-size pooling into the field’s mainstream debate. It was answered the same year by a detailed critical appraisal that questioned the database’s quality and the handling of multiple analyses, an exchange that set the template for nearly every psi meta-analysis since: a quantitative summary followed by a methodological rebuttal. The following year the two analysts published a joint communiqué stating that “there is an overall significant effect in this data base that cannot reasonably be explained by selective reporting or multiple analysis,” while continuing to differ over the degree to which the effect constitutes evidence for psi and agreeing that the final verdict awaited experiments conducted by a broader range of investigators according to more stringent standards.[19] In 1991 Jessica Utts reviewed the field’s meta-analyses for a statistics audience in Statistical Science, arguing that most nonstatisticians do not appreciate the connection between statistical power and successful replication of experimental effects.[21]
The most cited episode came in 1994, when a pooled “autoganzfeld” database was published in a mainstream psychology journal and argued for a replicable information-transfer anomaly.[2] Five years later a second meta-analysis of thirty newer ganzfeld studies from seven laboratories reported no overall effect, concluding that the earlier replication had not held up.[3] These two papers, published in the same venue, became the standard illustration that inclusion criteria and study selection can drive opposite conclusions from the same paradigm. Parallel summaries appeared for other designs, including an early pooling of RNG experiments[4] and a meta-analysis of dice-throwing studies spanning half a century.[5]
The exchange that followed the 1999 null became the method’s most instructive case. Storm and Ertel combined the older and newer ganzfeld databases into a unified set of 79 studies with a mean effect size of 0.138 (Stouffer z = 5.66, p = 7.78 × 10−9) and calculated that 857 unpublished nonsignificant studies would be needed to bring that result down to chance.[22] Milton and Wiseman replied that problems in the early ganzfeld studies make it difficult to draw strong conclusions from meta-analyses that include them, and noted that a single strongly positive study published after their cutoff, by Dalton in 1997, was by itself enough to pull the otherwise null database of post-1986 studies into overall statistical significance.[23] Bem, Palmer, and Broughton then had three independent raters score all forty post-communiqué replications for adherence to the standard ganzfeld protocol: the forty studies combined yielded a 30.1% hit rate (ES = .051, Stouffer Z = 2.59, p = .0048, one-tailed), the standard replications a 31.2% hit rate (ES = .096, Z = 3.49), and the non-standard replications 24.0% (ES = −.10, nonsignificant), with effect size significantly correlated with the degree of adherence to the standard protocol.[24] The three papers describe one body of experiments; the conclusions diverge on which studies are pooled and how a replication is defined.
Key findings
Across the free-response designs, later pooled summaries continued to report effects modestly above chance, with ganzfeld procedures producing the strongest results and no sign of a decline over roughly four decades.[6] A combined free-response meta-analysis reached similar conclusions, and a review tying several ganzfeld meta-analyses together reported a statistically significant overall ESP effect across more than a hundred publications.[7]
The forced-choice design tells a quieter story. A meta-analysis of 141 forced-choice ESP studies from 1987 to 2022 found a reliably positive but very weak effect, near the edge of statistical noise yet consistent across experimenters and target types.[8] A broader classical-and-Bayesian review argued that the cumulative evidence from more than two hundred non-local perception studies meets standards comparable to those used in clinical medicine.[9]
A meta-analysis of 141 forced-choice ESP studies from 1987 to 2022 found a reliably positive but very weak effect, consistent across experimenters and target types.[8]
For mind–matter interaction, an early RNG pooling reported a small but persistent influence on random output across many investigators,[4] and later analysts maintained that publication bias and quality concerns do not plausibly erase the cumulative result.[10] Related meta-analytic work covered the sense of being stared at[11] and the anomalous anticipation of future events.[12]
Landmark meta-analyses and framing agreements in the reference library
The entries below are the landmark meta-analyses of parapsychological research paradigms currently held in full text in the ESP-Nexus reference library, together with the methodological agreements that frame them, listed chronologically. Two founding documents of the 1985 exchange, Honorton’s meta-analysis of the original ganzfeld database and Hyman’s critical appraisal of the same studies, are not yet held; that episode is represented by the joint communiqué the two authors published the following year. Each entry reports the figures the paper itself states, and the entries report different kinds of number, Stouffer z values, standardized effect sizes, correlations, Cohen’s d, and raw hit rates, which cannot be combined into one bottom-line figure. For one of these databases the library’s own coverage has been measured: ESP-Nexus holds 57 of the 78 studies in the ganzfeld database analyzed by Tressoldi and Storm (2024), and pooling the held 57 under the same random-effects specification gives 0.085, against 0.077 computed for the full 78 studies; the paper itself prints .074, and the small gap between 0.077 and .074 is the site’s recomputation under the same specification versus the published rounding. That comparison is the site’s own holdings check, not a published finding; it means any summary built from the library’s ganzfeld holdings runs slightly more positive than the full published database.
Analysis. The entries show the range of outcomes the method produces on the same or overlapping databases. Milton and Wiseman’s thirty-study set returned an effect size of 0.013 where Storm and Ertel’s 79-study recombination returned 0.138,[3][22] and Bem, Palmer, and Broughton attributed the gap to protocol adherence, with standard replications at ES = .096 and non-standard replications at ES = −.10.[24] Across paradigms, the pooled values reported in these papers run from 0.01 for forced-choice ESP[28] to 0.21 for presentiment,[29] and the pre-registered 2024 ganzfeld estimate of .074[31] is smaller than the 0.142 the same research line reported for the 1997–2008 period.[6] On the random number generator literature, Radin and Nelson report unequivocal non-chance experimental effects with no quality–effect size relation,[4] while Bösch, Steinkamp, and Boller report that the same class of effect could in principle be a product of publication bias;[27] both readings rest on meta-analyses of overlapping databases.
Methodological critiques
The central dispute in RNG meta-analysis concerns the relationship between effect size and sample size. A 2006 critical meta-analysis published in a mainstream psychology journal concluded that psychokinesis was “not proven.”[27] A formal reply argued that this verdict rested on an assumption that effect size should be independent of sample size, and that correcting it left the cumulative RNG data supporting a genuine effect.[13] A companion paper by the same authors made the same case in a specialist venue.[10]
A second line of criticism focuses on questionable research practices and the file drawer. Defenders have responded that simulations meant to show such practices can manufacture ganzfeld effects rely on flawed prevalence estimates and unsupported assumptions.[14] Earlier replies to ganzfeld critiques argued that file-drawer concerns and meta-analytic methods had been mishandled by the critics.[15] A separate critique aimed at an umbrella review of anomalous-cognition meta-analyses flagged problems in inclusion criteria, heterogeneity reporting, and effect-size metrics, urging cautious reading of its conclusions.[16]
Some analysts within the field have also warned against over-reading patterns. One argument holds that supposed long-term experimenter effects and chronological declines often fail to survive rigorous meta-analysis and reflect a tendency to err in interpretation.[17] A further consideration is statistical power: replication failures for weak effects may stem from underpowered studies rather than from the phenomenon vanishing.[18]
Current practice
Contemporary psi meta-analysis increasingly mirrors reforms in mainstream psychology: pre-registration of protocols, explicit publication-bias tests, and shared databases. The episode that pushed this hardest followed a 2011 report of anomalous anticipation effects, after which dozens of laboratories across many countries posted replication attempts. A 2016 meta-analysis of ninety such experiments pooled this collected record into a single estimate,[12] and the scale of the replication effort, some thirty-three laboratories in fourteen countries, became a reference point for how the field now organizes large-scale tests.
Current meta-analyses also pay closer attention to moderators rather than headline pooled values, modeling how procedural details shape outcomes.[1] The forced-choice synthesis of 1987–2022 illustrates the modern two-stage approach, combining a review phase with a quantitative pooling phase.[8] The unresolved questions remain the ones meta-analysis was built to address: whether small pooled effects survive strict bias correction, and whether independent laboratories converge on the same estimate. On those points the literature still divides, and the method’s main value is to make the disagreement quantitative rather than rhetorical.
References
- Pooley, A., Murray, A., & Watt, C. (2023). Understanding the Factors at Play in the Sender-Receiver Dynamic During the Telepathy Ganzfeld: A Meta-Analysis. Journal of Anomalous Experience and Cognition, 3(1), 42–77. https://journals.lub.lu.se/jaex/article/view/23878 [PDF] R001 [Pooley 2023] ↩︎
- Bem, D., & Honorton, C. (1994). Does psi exist? Replicable evidence for an anomalous process of information transfer. Psychological Bulletin, 115(1), 4–18. https://doi.org/10.1037/0033-2909.115.1.4 [DOI] R002 [Bem 1994] ↩︎
- Milton, J., & Wiseman, R. (1999). Does psi exist? Lack of replication of an anomalous process of information transfer. Psychological Bulletin, 125(4), 387–391. https://psycnet.apa.org/record/1999-05876-001 [DOI] R003 [Milton 1999] ↩︎
- Radin, D., & Nelson, R. (1989). Evidence for consciousness-related anomalies in random physical systems. Foundations of Physics, 19(12), 1499–1514. https://doi.org/10.1007/bf00732509 [DOI] R004 [Radin 1989] ↩︎
- Radin, D., & Ferrari, D. C. (1991). Effects of consciousness on the fall of dice: A meta-analysis. Journal of Scientific Exploration, 5(1), 61–83. https://web.archive.org/web/20230331215150/https://www.scientificexploration.org/docs/5/jse_05_1_radin.pdf [archive.org] R005 [Radin 1991] ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992–2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136(4), 471–485. https://doi.org/10.1037/a0019457 [DOI] R006 [Storm 2010] ↩︎
- Tressoldi, P., Radin, D., & Storm, L. (2010). Extrasensory Perception and Quantum Models of Cognition. NeuroQuantology, 8(4). https://doi.org/10.14704/nq.2010.8.4.353 [DOI] R007 [Tressoldi 2010] ↩︎
- Storm, L., & Tressoldi, P. (2023). Assessing 36 Years of the Forced Choice Design in Extra Sensory Perception Research: A Meta-Analysis, 1987 to 2022. Journal of Scientific Exploration, 37(3), 517–535. https://journalofscientificexploration.org/index.php/jse/article/view/2967 [PDF] R008 [Storm & Tressoldi 2023] ↩︎
- Tressoldi, P. (2011). Extraordinary claims require extraordinary evidence: the case of non-local perception, a classical and Bayesian review of evidences. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.2030810 [DOI] R009 [Tressoldi 2011] ↩︎
- Radin, D., Nelson, R., Dobyns, Y., & Houtkooper, J. (2006). Assessing the evidence for mind-matter interaction effects. Journal of Scientific Exploration, 20(3), 361–374. https://web.archive.org/web/20240712074323/https://www.scientificexploration.org/docs/20/jse_20_3_radin_1.pdf [archive.org] R010 [Radin 2006 JSE] ↩︎
- Radin, D. (2005). The sense of being stared at: A preliminary meta-analysis. Journal of Consciousness Studies, 12(6), 95–100. [citation pending verification] R011 [Radin 2005] ↩︎
- Bem, D., Tressoldi, P., Rabeyron, T., & Duggan, M. (2016). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, article 1188. https://doi.org/10.12688/f1000research.7177.2 [PDF] R012 [Bem 2016] ↩︎
- Radin, D., Nelson, R., Dobyns, Y., & Houtkooper, J. (2006). Reexamining psychokinesis: Comment on Bösch, Steinkamp, and Boller (2006). Psychological Bulletin, 132(4), 529–532. https://doi.org/10.1037/0033-2909.132.4.529 [DOI] R013 [Radin 2006 PB] ↩︎
- Storm, L. (2026). Questions about Questionable Research Practices. Journal of Anomalous Experience and Cognition, 6(1), 92–102. https://doi.org/10.31156/jaex.27502 [PDF] R014 [Storm 2026] ↩︎
- Radin, D. (2007). Finding Or Imagining Flawed Research? The Humanistic Psychologist, 35(3), 297–299. https://doi.org/10.1080/08873260701578384 [DOI] R015 [Radin 2007] ↩︎
- Schmidt, S. (2021). Open Peer Comment to “Anomalous Cognition: An Umbrella Review of the Meta-Analytic Evidence”. Journal of Anomalous Experience and Cognition, 1(1-2), 73–75. https://journals.lub.lu.se/jaex/article/view/23439 [PDF] R016 [Schmidt 2021] ↩︎
- Storm, L. (2023). The Dark Spirit of the Trickster Archetype in Parapsychology. Journal of Scientific Exploration, 37(4), 665–682. https://journalofscientificexploration.org/index.php/jse/article/view/2715 [PDF] R017 [Storm 2023] ↩︎
- Tressoldi, P. (2012). Replication unreliability in psychology: elusive phenomena or “elusive” statistical power? Frontiers in Psychology, 3. https://doi.org/10.3389/fpsyg.2012.00218 [PDF] R018 [Tressoldi 2012] ↩︎
- Hyman, R., & Honorton, C. (1986). A joint communiqué: The psi ganzfeld controversy. Journal of Parapsychology, 50(4), 351–364. https://doi.org/10.30891/jopar.2018S.01.09 [DOI] R019 [Hyman 1986] ↩︎
- Honorton, C., & Ferrari, D. C. (1989). “Future telling”: A meta-analysis of forced-choice precognition experiments, 1935–1987. Journal of Parapsychology, 53. [citation pending verification] R020 [Honorton 1989] ↩︎
- Utts, J. (1991). Replication and meta-analysis in parapsychology. Statistical Science, 6(4), 363–403. https://doi.org/10.1214/ss/1177011577 [DOI] R021 [Utts 1991] ↩︎
- Storm, L., & Ertel, S. (2001). Does psi exist? Comments on Milton and Wiseman’s (1999) meta-analysis of ganzfeld research. Psychological Bulletin, 127(3), 424–433. https://doi.org/10.1037/0033-2909.127.3.424 [DOI] R022 [Storm 2001] ↩︎
- Milton, J., & Wiseman, R. (2001). Does psi exist? Reply to Storm and Ertel (2001). Psychological Bulletin, 127(3), 434–438. https://doi.org/10.1037/0033-2909.127.3.434 [DOI] R023 [Milton 2001] ↩︎
- Bem, D. J., Palmer, J., & Broughton, R. S. (2001). Updating the ganzfeld database: A victim of its own success? Journal of Parapsychology, 65. https://www.infoamerica.org/documentos_pdf/bem06.pdf [PDF] R024 [Bem 2001] ↩︎
- Sherwood, S. J., & Roe, C. A. (2003). A review of dream ESP studies conducted since the Maimonides dream ESP programme. Journal of Consciousness Studies, 10(6–7), 85–109. [citation pending verification] R025 [Sherwood 2003] ↩︎
- Schmidt, S., Schneider, R., Utts, J., & Walach, H. (2004). Distant intentionality and the feeling of being stared at: Two meta-analyses. British Journal of Psychology, 95, 235–247. [citation pending verification] R026 [Schmidt 2004] ↩︎
- Bösch, H., Steinkamp, F., & Boller, E. (2006). Examining psychokinesis: The interaction of human intention with random number generators—A meta-analysis. Psychological Bulletin, 132(4), 497–523. https://doi.org/10.1037/0033-2909.132.4.497 [DOI] R027 [Bösch 2006] ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2012). Meta-analysis of ESP studies, 1987–2010: Assessing the success of the forced-choice design in parapsychology. Journal of Parapsychology, 76, 243–273. [citation pending verification] R028 [Storm 2012] ↩︎
- Mossbridge, J., Tressoldi, P., & Utts, J. (2012). Predictive physiological anticipation preceding seemingly unpredictable stimuli: A meta-analysis. Frontiers in Psychology, 3, article 390. https://doi.org/10.3389/fpsyg.2012.00390 [PDF] R029 [Mossbridge 2012] ↩︎
- Storm, L., & Tressoldi, P. E. (2020). Meta-analysis of free-response studies 2009–2018: Assessing the noise-reduction model ten years on. Journal of the Society for Psychical Research, 84(4), 193–219. https://doi.org/10.31234/osf.io/3d7at [PDF] R030 [Storm 2020] ↩︎
- Tressoldi, P. E., & Storm, L. (2024). Stage 2 Registered Report: Anomalous perception in a Ganzfeld condition—A meta-analysis of more than 40 years investigation. F1000Research, 10, article 234. https://f1000research.com/articles/10-234/pdf [PDF] R031 [Tressoldi 2024] ↩︎