Statistical and File-Drawer Critiques

Statistical critiques of parapsychology center on whether reported psi effects survive correction for selective publication, optional stopping, multiple comparisons, and prior-probability adjustment. The exchange between Hyman, Alcock, Wagenmakers, and Rouder on one side and Honorton, Rosenthal, Storm, Tressoldi, and Bem on the other has shaped how meta-analytic claims, Bayes factors, and pre-registration are evaluated in psi research.

Key findings

  • Sterling (1959) first documented selective publication of statistically significant results in psychology; Hyman (1985) applied the framework to the ganzfeld database, while Honorton (1985) calculated a fail-safe N indicating an implausible number of unpublished null studies would be required to nullify the meta-analytic effect.[9][1][2]
  • Alcock (1987) catalogued optional-stopping concerns as a structural feature of pre-1990 parapsychology designs; pre-registration via the Koestler Parapsychology Unit registry[13] and the Bem, Tressoldi, Rabeyron, & Duggan (2015) meta-analysis explicitly addressed this concern.[3][11]
  • Effect sizes in ganzfeld and free-response meta-analyses (Cohen’s d roughly 0.10-0.20) are comparable in magnitude to many effects in mainstream social psychology that have themselves been re-evaluated under replication-crisis scrutiny (Open Science Collaboration, 2015).[5][12]
  • Wagenmakers, Wetzels, Borsboom, & van der Maas (2011) argued that Bem’s (2011) precognition findings would not survive default Bayes factor analysis; Rouder & Morey (2011) computed a Bayes factor meta-analysis of Bem’s nine studies and reported support against the null under their chosen prior, while noting Bayes factors are prior-dependent.[7][6][8]
  • The “extraordinary claims require extraordinary evidence” standard articulated by Hyman and Alcock is contested by parapsychologists who argue the standard should apply symmetrically to skeptical claims of pure chance performance across decades of data.[1][3]
  • Storm, Tressoldi, & Di Risio (2010) reported a free-response meta-analysis covering 1992-2008 with a homogeneous hit-rate effect; Milton & Wiseman (1999) had earlier reported a non-significant ganzfeld meta-analysis over an overlapping but distinct study set.[5][4]

Overview

Statistical and file-drawer critiques of parapsychology are not unique to the field; they form a subset of the broader methodological reform movement that produced the replication crisis in psychology and the rise of pre-registration as standard practice. What distinguishes the psi case is the combination of small but persistent effect sizes, decades of cumulative meta-analyses, and the question of whether the cumulative record reflects a real anomaly, a residual artifact of publication bias and analytic flexibility, or some combination of both.

This page covers the principal statistical-methodology critiques: the file-drawer problem (Sterling 1959; Hyman 1985), optional stopping and multiple comparisons (Alcock 1987), effect-size interpretation, Bayesian reanalysis (Wagenmakers et al. 2011; Rouder & Morey 2011), and the extraordinary-claims standard. Each critique is presented with the response published by parapsychologists in the same literature, and with a neutral synthesis of what the published exchange establishes and what remains open.

The File-Drawer Problem in Psi Research

Sterling (1959) first documented that journals in psychology overwhelmingly published statistically significant results, creating systematic distortion in any literature relying on aggregated published findings.[9] The “file drawer” metaphor — for the hypothetical drawer of unpublished null studies — became central to meta-analytic methodology by the 1980s.

Hyman (1985) applied the file-drawer concern directly to the ganzfeld database in his critical appraisal of Honorton’s meta-analytic claims.[1] Honorton (1985) responded with a fail-safe N calculation using Rosenthal’s framework: the number of unpublished null studies that would need to exist to reduce the observed effect to non-significance. For the early ganzfeld database, Honorton calculated a fail-safe N in the hundreds, arguing that a file drawer of that magnitude was implausible given the small community of psi researchers and the field’s practice of submitting both positive and null results to parapsychology journals.[2]

Rosenthal (1986), commenting on the same exchange, contributed methodological refinements to fail-safe N calculation and noted that the published parapsychology record showed a pattern of submitted null results inconsistent with severe file-drawer suppression at the level skeptics had hypothesized.[10]

Optional Stopping and Multiple Comparisons

Alcock (1987) catalogued optional stopping — terminating data collection when results reach significance — as a structural concern in pre-1990 parapsychology designs that lacked explicit stopping rules.[3] Multiple-comparisons concerns extended to post-hoc subgroup analyses (gender, target type, dynamic vs. static targets in ganzfeld) that, if not pre-specified, inflate the family-wise Type I error rate.

The methodological response in parapsychology took two forms: tighter protocol specification (the autoganzfeld procedure introduced by Honorton in the mid-1980s removed several sources of analytic flexibility), and adoption of pre-registration once the infrastructure became available. The Koestler Parapsychology Unit registry provides public time-stamped protocol records; Bem, Tressoldi, Rabeyron, & Duggan (2015) reported a meta-analysis of 90 precognition experiments that included a subset of pre-registered replications.[11]

Effect Sizes in Context

Meta-analytic effect sizes in psi research are small. Storm, Tressoldi, & Di Risio (2010) reported a free-response meta-analysis covering 1992-2008 with an overall hit rate of 31.5% against a chance baseline of 25%, corresponding to a small Cohen’s d in the range typical of psi meta-analyses.[5] Bem, Tressoldi, Rabeyron, & Duggan (2015) reported a comparable effect size across 90 precognition experiments.[11]

The Open Science Collaboration (2015) reproducibility project reported that mainstream social-psychology effects, when subjected to direct replication, often shrank to effect sizes in the same small range — and in many cases failed to reach significance in the replication studies.[12] This finding reframes the effect-size critique: the question is not whether psi effects are smaller than typical psychology effects (in many cases they are comparable) but whether the cumulative meta-analytic record in psi reflects a real signal that the methodological reforms of the 2010s would preserve.

Milton & Wiseman (1999) reported a non-significant ganzfeld meta-analysis over a study set partially overlapping with later analyses; subsequent re-analyses by parapsychologists disputed the study-inclusion criteria, with the exchange continuing in Psychological Bulletin through follow-up commentaries.[4]

Bayesian Reanalyses and Prior Probabilities

Wagenmakers, Wetzels, Borsboom, & van der Maas (2011) responded to Bem’s (2011) Journal of Personality and Social Psychology precognition paper by arguing that default Bayes factor analysis with a Cauchy prior centered on zero would not support Bem’s findings against the null.[7][6] Their critique extended beyond psi to a broader methodological argument that p-value-based inference systematically overstates evidence for novel effects.

Rouder & Morey (2011) computed a Bayes factor meta-analysis of Bem’s nine experiments using a default JZS prior and reported a Bayes factor of approximately 40 in favor of an effect, while explicitly noting that Bayes factors are prior-dependent and that a skeptic with a stronger prior against psi would draw a different conclusion from the same data.[8] The Rouder & Morey reanalysis is sometimes cited as evidence for Bem’s findings and sometimes cited as illustrating the prior-dependence problem in Bayesian inference — both characterizations appear in the published literature.

The Bayesian-reanalysis exchange illustrated a structural feature of the dispute: when the prior probability assigned to psi is very low, Bayes factor calculations require correspondingly large effect sizes or sample sizes to shift posterior belief, even when the data would convince a frequentist analyst.

The Extraordinary-Claims Standard

Hyman (1985) and Alcock (1987) both invoked the principle that claims contradicting established physics require evidentiary thresholds higher than those applied to ordinary empirical claims in psychology.[1][3] The principle is widely associated with Carl Sagan (“extraordinary claims require extraordinary evidence”) and traces further back to Hume’s argument on miracles.

Parapsychologists have responded that the standard, while reasonable in principle, has been applied asymmetrically: the claim that decades of psi data reflect pure chance also requires evidence, and the burden of demonstrating that a multi-decade meta-analytic record with consistent small effects across independent laboratories arose from artifact alone has not been formally discharged. The dispute over what counts as “extraordinary evidence” — replication count, effect size, mechanism specification, or some combination — is itself a methodological question with multiple published positions.

Pre-Registration as Response

The methodological-reform response to optional-stopping and multiple-comparisons critiques is pre-registration: public time-stamped specification of hypothesis, analysis plan, sample size, and stopping rule before data collection. A study registry associated with the Koestler Parapsychology Unit predates the widespread adoption of pre-registration in mainstream psychology

Bem, Tressoldi, Rabeyron, & Duggan (2015) reported a meta-analysis of 90 precognition experiments that included pre-registered replications as a separately analyzable subset; the authors reported that effect sizes in the pre-registered subset were comparable to the full set.[11] Whether pre-registered replications in psi research will continue to show consistent small effects as the practice becomes more widespread is an empirical question that subsequent published replication attempts will address.

Counterarguments and Debate

Critique 1: The ganzfeld meta-analytic effect is an artifact of selective publication — a sufficiently large file drawer of unpublished null studies would nullify the reported effect.

Skeptic source: Hyman (1985) raised the file-drawer concern as part of his critical appraisal of the early ganzfeld database in Journal of Parapsychology.[1]

Rebuttal: Honorton (1985) calculated Rosenthal’s fail-safe N for the contested ganzfeld database and reported a value in the hundreds of unpublished null studies required to reduce the effect to non-significance, arguing that a file drawer of this magnitude was implausible given the small parapsychology research community and the publication of null results in parapsychology journals.[2]

Analysis. Hyman and Honorton published their critique-response exchange in the same 1985 issue of Journal of Parapsychology. Honorton’s fail-safe N calculation used Rosenthal’s framework; the calculation itself is reproducible from the published study counts. Whether the implied file-drawer size is “implausible” depends on assumptions about submission rates that are not directly measurable. Rosenthal (1986) contributed independent methodological commentary on the exchange.[10]

Critique 2: Bem’s (2011) precognition findings do not survive default Bayes factor analysis and reflect the broader problem of p-value-based inference overstating evidence.

Skeptic source: Wagenmakers, Wetzels, Borsboom, & van der Maas (2011) published a critique in Journal of Personality and Social Psychology arguing that default Bayes factor analysis with a Cauchy prior centered on zero did not support Bem’s findings.[7]

Rebuttal: Rouder & Morey (2011) computed a Bayes factor meta-analysis of Bem’s nine experiments using a JZS prior and reported a Bayes factor of approximately 40 in favor of an effect, while noting prior-dependence as a structural feature of Bayesian inference.[8]

Analysis. Wagenmakers et al. and Rouder & Morey published in the same 2011 timeframe using different default priors and arrived at different Bayes factors over the same data. Rouder & Morey explicitly acknowledged that a skeptic with a stronger prior against psi would draw a different conclusion from the calculation. The exchange illustrates that Bayes factors require explicit prior specification, and that the dispute over Bem (2011) is partly a dispute over prior probabilities rather than over the data themselves.

Critique 3: Optional stopping and post-hoc subgroup analyses in pre-1990 parapsychology designs inflated apparent effect sizes; modern reforms have not been applied uniformly to historical meta-analyses.

Skeptic source: Alcock (1987) catalogued optional-stopping and multiple-comparisons concerns as structural features of pre-1990 parapsychology designs in his target article in Behavioral and Brain Sciences.[3]

Rebuttal: Honorton’s autoganzfeld procedure (mid-1980s) introduced explicit pre-specified stopping rules and removed several sources of analytic flexibility; the Koestler Parapsychology Unit registry provides time-stamped protocol records, and Bem, Tressoldi, Rabeyron, & Duggan (2015) reported a meta-analysis that included pre-registered replications as a separately analyzable subset with effect sizes comparable to the full set.[11]

Analysis. Alcock’s 1987 target article and the subsequent pre-registration response are separated by roughly 25 years. The autoganzfeld procedure (mid-1980s) and the Koestler registry (2012) are documented methodological responses. Whether retrospective application of modern standards to pre-1990 meta-analyses is appropriate is a methodological question discussed in Storm, Tressoldi, & Di Risio (2010) and in Bem et al. (2015), with different authors taking different positions in the published literature.[5][11]

References
  1. Hyman, R. (1985). The ganzfeld psi experiment: A critical appraisal. Journal of Parapsychology, 49(1), 3-49. ↩︎
  2. Honorton, C. (1985). Meta-analysis of psi ganzfeld research: A response to Hyman. Journal of Parapsychology, 49(1), 51-91. ↩︎
  3. Alcock, J. E. (1987). Parapsychology: Science of the anomalous or search for the soul? Behavioral and Brain Sciences, 10, 553-643. ↩︎
  4. Milton, J., & Wiseman, R. (1999). Does psi exist? Lack of replication of an anomalous process of information transfer. Psychological Bulletin, 125(4), 387-391. ↩︎
  5. Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992-2008. Psychological Bulletin, 136(4), 471-485. https://doi.org/10.1037/a0019457 ↩︎
  6. Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407-425. https://doi.org/10.1037/a0021524 ↩︎
  7. Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2011). Why psychologists must change the way they analyze their data: The case of psi. Journal of Personality and Social Psychology, 100(3), 426-432. https://doi.org/10.1037/a0022790 ↩︎
  8. Rouder, J. N., & Morey, R. D. (2011). A Bayes factor meta-analysis of Bem’s ESP claim. Psychonomic Bulletin & Review, 18, 682-689. https://doi.org/10.3758/s13423-011-0088-7 ↩︎
  9. Sterling, T. D. (1959). Publication decisions and their possible effects on inferences drawn from tests of significance — or vice versa. Journal of the American Statistical Association, 54(285), 30-34. https://doi.org/10.2307/2282137 ↩︎
  10. Rosenthal, R. (1986). Meta-analytic procedures and the nature of replication: The ganzfeld debate. Journal of Parapsychology, 50, 315-336. ↩︎
  11. Bem, D., Tressoldi, P., Rabeyron, T., & Duggan, M. (2015). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, 1188. https://doi.org/10.12688/f1000research.7177.2 ↩︎
  12. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. ↩︎
  13. Watt, C., & Kennedy, J. E. (2015). Lessons from the first two years of operating a study registry. Frontiers in Psychology, 6, 173. https://doi.org/10.3389/fpsyg.2015.00173 ↩︎