Replication and Meta-Analysis Debates

Whether psi effects replicate, and what counts as a successful replication, has been contested for decades. The debate sits at the intersection of small effect sizes, meta-analytic technique, the broader reproducibility movement in psychology, and the post-2011 turn toward pre-registration.

Key findings

  • Milton and Wiseman (1999) reported a null result in a meta-analysis of 30 post-Honorton ganzfeld studies[1], breaking the apparent meta-analytic trend established in the 1980s[2].
  • Storm, Tressoldi, and Di Risio (2010) reanalyzed free-response studies from 1992-2008 and reported a small but statistically significant positive effect (mean z ≈ 5.48 across 29 ganzfeld studies)[3]; Storm, Tressoldi, and Utts (2013) replied to a Bayesian critique using both frequentist and Bayesian approaches[4].
  • Bem (2011) published nine experiments reporting “retroactive” psi effects in the Journal of Personality and Social Psychology; eight of nine reported statistically significant results[5].
  • Wagenmakers and colleagues (2011) reanalyzed Bem’s data using default Bayes factors and reported substantial evidence for the null[6]; Rouder and Morey (2011) reported a Bayes-factor meta-analysis with similar conclusions[7].
  • Galak, LeBoeuf, Nelson, and Simmons (2012) reported seven failed direct replications of Bem’s retroactive priming paradigm[8].
  • Bem, Tressoldi, Rabeyron, and Duggan (2015) compiled a meta-analysis of 90 replication attempts, reporting an overall positive effect[9].
  • The Open Science Collaboration (2015) Reproducibility Project: Psychology successfully replicated 36 of 100 mainstream social and cognitive psychology studies[10], situating psi replication debates within a broader methodological reassessment.

Overview: replication as a contested concept

Replication in parapsychology has never been a settled concept. Effect sizes in psi experiments are characteristically small, which means that any single replication attempt typically lacks the statistical power to detect the predicted effect even if it is real. Alcock (1987), reviewing the field for Behavioral and Brain Sciences, argued that absence of routine independent replication was a structural problem[11]. Honorton (1985), responding to Hyman’s earlier critiques of ganzfeld research, argued that meta-analytic pooling across many small studies was the appropriate response when individual studies lack power[2].

The decades-long exchange that followed has turned on three intertwined questions: what statistical apparatus correctly summarizes a heterogeneous corpus of small studies, whether meta-analytic positive results survive specification choices about which studies to include, and whether direct replications by independent labs cohere with the meta-analytic record. Each of these questions has been answered differently by different authors. See the sibling spoke Ganzfeld Debate for the parallel exchange specific to ganzfeld telepathy studies, and the methods page Replication for the broader methodological framing.

The Milton-Wiseman 1999 inflection

Milton and Wiseman (1999) published a meta-analysis in Psychological Bulletin covering 30 ganzfeld studies conducted between 1987 and 1997 (i.e., the studies completed after Honorton’s earlier autoganzfeld series)[1]. They reported a mean effect size statistically indistinguishable from chance, and concluded that the ganzfeld effect had not replicated in the post-autoganzfeld corpus.

The paper became the most-cited skeptical reference in the ganzfeld literature and is frequently treated as the methodological inflection point for the debate. Subsequent reanalyses, including Storm, Tressoldi, and Di Risio (2010), proposed that Milton and Wiseman’s pooling decisions, particularly their treatment of trial-count weighting and inclusion of methodologically heterogeneous studies, materially affected the reported effect size[3].

Post-1999 ganzfeld meta-analyses

Storm, Tressoldi, and Di Risio (2010) reanalyzed free-response psi studies published from 1992 through 2008 using updated inclusion criteria and reported a small but statistically significant cumulative effect across 29 ganzfeld studies, with a Stouffer Z ≈ 5.48 and effect sizes in the range typically reported for similarly-sized social-psychology effects[3]. They argued that updating the corpus through 2008 and applying contemporary meta-analytic practice produced a different summary than Milton and Wiseman’s 1997 cutoff.

Storm, Tressoldi, and Utts (2013) returned to the corpus after a Bayesian critique and reported both frequentist and Bayesian analyses, arguing that the choice of prior distribution materially affects the Bayesian summary and that under more diffuse priors the data still favor a non-null effect[4]. The exchange illustrates a recurring pattern: the same data, analyzed with different defensible statistical conventions, produce divergent summary conclusions.

The Bem 2011 episode

Daryl Bem’s 2011 paper “Feeling the Future,” published in the Journal of Personality and Social Psychology, reported nine experiments using standard social-psychology paradigms (priming, recall, habituation) run in time-reversed form[5]. Eight of the nine reported statistically significant results in the predicted direction. Publication in a top mainstream journal, with positive results across a standard battery, made the paper a focal point for methodological discussion in psychology more broadly.

Wagenmakers, Wetzels, Borsboom, and van der Maas (2011) published an accompanying critique using default Bayes factors with the symmetric Cauchy prior recommended by Rouder and colleagues; they reported that the data provided substantial evidence for the null hypothesis under that prior, and argued that the apparent significance of Bem’s results reflected limitations of null-hypothesis significance testing rather than evidence of psi[6]. Rouder and Morey (2011) independently published a Bayes-factor meta-analysis of Bem’s experiments reaching a similar conclusion[7]. Galak, LeBoeuf, Nelson, and Simmons (2012) reported seven direct replication attempts of Bem’s retroactive-priming paradigm and found no evidence of the predicted effect across approximately 3,000 participants[8].

Multi-lab replication of Bem 2011

Bem, Tressoldi, Rabeyron, and Duggan (2015), published in F1000Research, compiled a meta-analysis of 90 replication attempts of the Bem (2011) paradigms drawn from 33 laboratories in 14 countries, including both successful and failed replication attempts identified through systematic search[9]. They reported a small overall positive effect across the pooled corpus.

The Bem et al. 2015 meta-analysis itself drew methodological discussion. The included corpus mixed independent direct replications with conceptual and methodological variants, and the file-drawer correction relied on assumptions about unreported attempts. Galak et al. (2012) provide one external benchmark: seven independent attempts at the retroactive-priming paradigm, the most studied of Bem’s nine experiments, reported null results[8].

The wider reproducibility-crisis context

The Open Science Collaboration (2015) Reproducibility Project: Psychology attempted direct replications of 100 mainstream studies drawn from three high-impact psychology journals, including the journal that published Bem (2011); 36 of 100 replications produced statistically significant results in the original direction, and average effect sizes in replications were roughly half those of original studies[10]. This benchmark situates the Bem 2011 replication exchange within a broader methodological reassessment of social and cognitive psychology, not as an isolated parapsychology dispute.

Bem (2011) is frequently cited as a proximate trigger for the post-2011 methodological tightening in psychology, alongside the Stapel fraud case and the Simmons, Nelson, and Simonsohn “false-positive psychology” paper from the same year. The substantive question of whether psi effects replicate is therefore inseparable from the methodological question of what replication standards apply across psychology as a whole.

The pre-registration era

Pre-registration, in which a study’s hypotheses, sample size, exclusion criteria, and analysis plan are publicly recorded before data collection, became one of the central methodological responses to the post-2011 reassessment. Bem et al. (2015) report pre-registration status as an inclusion variable in their meta-analysis of psi replication attempts[9].

Pre-registered psi studies remain a small subset of the overall corpus. The question of whether pre-registered direct replications by labs without prior commitment to psi research yield positive effect sizes comparable to the Storm et al. (2010) and Bem et al. (2015) meta-analytic summaries is one specific empirical question on which independent published evidence is still accumulating[3][9].

Counterarguments and Debate

Critique 1: Meta-analytic positive results disappear when the corpus is updated to include post-Honorton studies under consistent inclusion criteria.

Skeptic source: Milton and Wiseman (1999), Psychological Bulletin[1].

Rebuttal: Storm, Tressoldi, and Di Risio (2010) reanalyzed the post-1992 corpus through 2008 using updated inclusion criteria and reported a statistically significant cumulative effect across 29 ganzfeld studies (Stouffer Z ≈ 5.48)[3]; Storm, Tressoldi, and Utts (2013) added Bayesian analyses reaching the same directional conclusion under diffuse priors[4].

Analysis. Milton and Wiseman (1999) used a 1987-1997 cutoff; Storm, Tressoldi, and Di Risio (2010) extended through 2008. The two meta-analyses cover overlapping but non-identical corpora and use different inclusion rules for trial-count weighting and methodological heterogeneity. The Storm-Tressoldi-Utts (2013) Bayesian reply directly addresses prior-sensitivity objections by reporting results under multiple prior distributions.

Critique 2: Bem (2011)’s significant results reflect limitations of null-hypothesis significance testing rather than evidence of retroactive psi.

Skeptic source: Wagenmakers, Wetzels, Borsboom, and van der Maas (2011), Journal of Personality and Social Psychology[6]; Rouder and Morey (2011), Psychonomic Bulletin & Review[7].

Rebuttal: Bem, Tressoldi, Rabeyron, and Duggan (2015) compiled a meta-analysis of 90 replication attempts of the Bem (2011) paradigms from 33 laboratories in 14 countries and reported a small overall positive effect[9]. They argue that the multi-lab corpus addresses the original methodological criticisms by enlarging the empirical base beyond any single laboratory.

Analysis. Wagenmakers et al. (2011) used the default Cauchy prior recommended by Rouder and colleagues; Bem and colleagues have argued that this prior is overly conservative for small expected effects. The Bem et al. (2015) multi-lab compilation mixes direct replications with conceptual variants. Galak, LeBoeuf, Nelson, and Simmons (2012) report seven direct replication attempts of the most-studied retroactive-priming paradigm with null results across approximately 3,000 participants[8].

Critique 3: Psi research has produced no protocol that yields a positive result on demand in the hands of any competent experimenter, which is the standard for a replicable phenomenon.

Skeptic source: Alcock (1987), Behavioral and Brain Sciences[11].

Rebuttal: Honorton (1985) argued that meta-analytic pooling across many small studies is the methodologically appropriate response when individual studies characteristically lack statistical power, and reported a cumulative significant effect in the pre-1985 ganzfeld corpus[2]. Subsequent meta-analyses by Storm, Tressoldi, and Di Risio (2010) and Bem et al. (2015) extend the same meta-analytic approach to more recent corpora[3][9].

Analysis. The Open Science Collaboration (2015) reported 36 successful replications out of 100 attempts across mainstream social and cognitive psychology[10]; the on-demand-by-any-experimenter standard described by Alcock (1987) is more stringent than the replication rate documented across mainstream psychology more generally. Whether psi-specific replication rates meet, exceed, or fall short of the mainstream-psychology baseline is one specific empirical question; Bem et al. (2015) report 90 attempts, of which a documented subset are independent direct replications[9].

References
  1. Milton, J., & Wiseman, R. (1999). Does psi exist? Lack of replication of an anomalous process of information transfer. Psychological Bulletin, 125(4), 387-391. https://doi.org/10.1037/0033-2909.125.4.387
  2. Honorton, C. (1985). Meta-analysis of psi ganzfeld research: A response to Hyman. Journal of Parapsychology, 49(1), 51-91.
  3. Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992-2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136(4), 471-485. https://doi.org/10.1037/a0019457
  4. Storm, L., Tressoldi, P. E., & Utts, J. (2013). Testing the Storm et al. (2010) meta-analysis using Bayesian and frequentist approaches: Reply to Rouder et al. (2013). Psychological Bulletin, 139(1), 248-254. https://doi.org/10.1037/a0029506
  5. Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407-425. https://doi.org/10.1037/a0021524
  6. Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2011). Why psychologists must change the way they analyze their data: The case of psi: Comment on Bem (2011). Journal of Personality and Social Psychology, 100(3), 426-432. https://doi.org/10.1037/a0022790
  7. Rouder, J. N., & Morey, R. D. (2011). A Bayes factor meta-analysis of Bem’s ESP claim. Psychonomic Bulletin & Review, 18, 682-689. https://doi.org/10.3758/s13423-011-0088-7
  8. Galak, J., LeBoeuf, R. A., Nelson, L. D., & Simmons, J. P. (2012). Correcting the past: Failures to replicate psi. Journal of Personality and Social Psychology, 103(6), 933-948. https://doi.org/10.1037/a0029709
  9. Bem, D., Tressoldi, P., Rabeyron, T., & Duggan, M. (2015). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, 1188. https://doi.org/10.12688/f1000research.7177.2
  10. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
  11. Alcock, J. E. (1987). Parapsychology: Science of the anomalous or search for the soul? Behavioral and Brain Sciences, 10, 553-643.