Daryl Bem experiment, data and criticism

Daryl J. Bem, PhD

Daryl J. Bem, PhD is a social psychologist at Cornell University whose 2011 publication of nine experiments on retroactive influences on cognition and affect appeared in a top-tier mainstream psychology journal and prompted extended methodological debate across experimental psychology [3]. Before summarizing his research record, one scoping note: the share of the published literature on Bem’s precognition work that this library holds has not been measured, so the sections below describe the studies retrieved for this question, not the full field of replication attempts.

Experiments

Bem’s central contribution is the 2011 paper “Feeling the Future,” which reported nine experiments testing whether events occurring after a participant’s response could nonetheless influence that response [3]. The experiments adapted well-established social-psychology paradigms — priming, habituation, recall facilitation — and reversed their temporal order so the causal stimulus followed the measured behavior [3].

Two subsequent lines of work extend that program:

  • The 2016 meta-analysis (Bem, Tressoldi, Rabeyron & Duggan), which aggregated 90 experiments from 33 laboratories across 14 countries testing anomalous anticipation of random future events [2].
  • The 2021 replication study (Schlitz, Bem, Marcusson-Clavertz, Cardeña and a large international group), which ran two replications of a time-reversed priming task and examined the role of experimenter and participant expectancy in reaction times [1].

Bem also authored earlier reviews of psi evidence, including a 1994 Psychological Bulletin piece on replicable evidence for anomalous information transfer, referenced elsewhere on the site.

Methodology

The 2011 design’s defining feature was the temporal reversal: on each trial, the participant’s response was recorded before the purportedly causal stimulus was presented [1]. In the priming variant, the proportion of congruent versus incongruent pairings is held at 0.5 across all trials, so there is no non-psi way for a participant to anticipate the trial type currently on screen [1].

In the 2016 meta-analysis, each of the 90 experiments was categorized as an exact replication of one of Bem’s experiments (31), a modified replication (38), or an independently designed experiment (the remainder), and the authors coded moderators including peer-review status and replication type [2]. The 2021 study by Schlitz and colleagues added a pre-specified interaction test between experimenters and participants — motivated by the long-noted “experimenter orientation” variable — and supplemented its primary analyses with binomial tests against a mean chance expectation of 50% [1].

Data

The figures below are drawn from the retrieved primary sources. Metrics differ across rows (a mean d, a Hedges’ g, a null result), so they are not poolable into a single number; each is reported against its own study.

StudyScopeEffect sizeTest statisticp
Bem (2011) [3]9 experiments, 8 significantmean d = 0.22Stouffer z = 6.662.68 × 10⁻¹¹
Bem et al. (2016), full database [2]90 experiments, 33 labs, 14 countriesHedges’ g = 0.09z = 6.401.2 × 10⁻¹⁰
Bem et al. (2016), Bem’s own studies excluded [1][2]independent replicationsES = 0.06z = 4.161.1 × 10⁻⁵
Bem et al. (2016), p-curve estimate [2]full DB / independent0.20 / 0.24
Bem et al. (2016), 15 precognitive priming experiments [1]priming subsetd = 0.110.003
Schlitz, Bem et al. (2021) [1]2 time-reversed priming replicationsnullnot significant

The 2016 paper reports a Bayes Factor of 5.1 × 10⁹ for the full database and 3.85 × 10³ (≈3,853) with Bem’s own studies removed, both exceeding the criterion of 100 the authors cite for “decisive evidence” [1][2]. Against these positive aggregates, the 2021 Schlitz/Bem replications did not reject the null hypothesis: neither the primary analyses nor the supplementary binomial tests showed above-chance scoring [1]. The direction of the record is therefore mixed — the meta-analytic and p-curve summaries are positive, while the two 2021 priming replications are null.

Skeptical critiques

What critics argue. Wagenmakers, Wetzels, Borsboom, and van der Maas (2011) argued that Bem’s analyses were partly exploratory and that standard frequentist testing overstated the evidence; reanalyzing the nine experiments under Bayesian priors, they found the support far weaker — under a Cauchy prior several experiments yielded Bayes Factors below 1, favoring the null [3]. Schimmack (2012) raised the separate objection that the studies were low-powered, a condition that can inflate the rate of false positives [1]. Francis (2012) argued that the low power of the Bem (2011) studies implied a number of unreported experiments or p-hacking, i.e. publication bias [2].

What the experimental data show. Bem, Utts and Johnson (2011) replied that the choice of prior distribution drives the Bayesian result and that a knowledge-based prior — informed by prior psi and social-psychology effect sizes — yields Bayes Factors supporting H₁ where the Wagenmakers Cauchy prior did not [3]. On the file-drawer objection, the 2016 authors reported a p-curve analysis estimating a true effect of 0.20 (full database) and 0.24 (independent replications), figures they present as evidence against selective reporting, and noted that the critic had acknowledged not reading past page 9 of the article before commenting on it [2]. Independent of Bem’s own data, the 2016 set still returned ES = 0.06, z = 4.16 [1][2].

Analysis. The exchange turns on statistical convention rather than raw counts: Wagenmakers et al. (2011) and Bem/Utts/Johnson (2011) analyzed the same nine experiments and reached opposite conclusions because they specified different prior distributions [3]. On replication, the record within these sources is split by group — the Bem-led 2016 meta-analysis reports a positive aggregate across 90 experiments [2], while the 2021 study co-authored by Schlitz and Bem reports two null time-reversed priming replications [1]. The ESP-Nexus page on Replication and Meta-Analysis Debates traces this “Bem 2011 episode” as a recurring pattern where the same data under different defensible conventions produce divergent summaries.

For the full profile, see Daryl J. Bem, PhD.

References
  1. Schlitz, M. J., Bem, D. J., Marcusson-Clavertz, D., Cardena, E., Lyke, J., Grover, R., Blackmore, S. J., Tressoldi, P. E., Roney-Dougal, S. M., Bierman, D. J., Jolij, J. J., Lobach, E., Hartelius, G., Rabeyron, T., Bengston, W. F., Nelson, S. E., Moddel, G., & Delorme, A. (2021). Two Replication Studies of a Time-Reversed (Psi) Priming Task and the Role of Expectancy in Reaction Times. Journal of Scientific Exploration, 35, pp. 65–90. https://doi.org/10.31275/20211903
  2. Bem, D. J., Tressoldi, P. E., Rabeyron, T., & Duggan, M. W. (2016). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, article 1188. https://doi.org/10.12688/f1000research.7177.2
  3. Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100, 407–425. https://doi.org/10.1037/a0021524
Deeper dives on ESP-Nexus

Ask another question