Muhmenthaler et al. (2022)
The Future Failed: No Evidence for Precognition in a Large Scale Replication Attempt of Bem (2011)
Muhmenthaler, M. C., Dubravac, M., & Meier, B. (2022). The future failed: No evidence for precognition in a large scale replication attempt of Bem (2011). Psychology of Consciousness: Theory, Research, and Practice. Advance online publication. https://doi.org/10.1037/cns0000342
AI Assessment
A high-powered online replication of three of Bem’s (2011) time-reversed “precognition” experiments, run with more than 2,000 participants. All three backward (precognition) tests were null, while the forward control versions reproduced the ordinary, non-time-reversed effects — showing the tasks themselves worked and the study was sensitive enough to detect a real effect had one been present. It is a large, transparent, open-data direct replication; the limits the authors state are that it was not formally preregistered and that a wide battery of exploratory post-hoc analyses (from which a few scattered significant effects emerged) inflates the risk of false positives. This audit describes what the paper reports and how it was run; it takes no position on whether precognition exists.
Provenance
DOI. 10.1037/cns0000342 · Psychology of Consciousness: Theory, Research, and Practice (American Psychological Association) · advance online publication, 29 September 2022.
Study type. A direct replication of three of the nine experiments in Bem (2011) — the two affective-priming experiments (Bem’s Experiments 3 and 4) and a retroactive free-recall experiment (Bem’s Experiment 9) — each run in a backward (“precognition”) version and a forward (control) version.
Authors. Michele C. Muhmenthaler, Mirela Dubravac, and Beat Meier, Institute of Psychology, University of Bern, Switzerland.
Funding. No funding was received; the authors declared no competing interests.
Ethics. Approved by the Ethics Committee of the Faculty of Human Sciences, University of Bern, following the Declaration of Helsinki; participants gave informed consent online.
Data availability. The datasets for all three experiments and the analysis code are openly available (doi.org/10.48620/40) under a Creative Commons Attribution 4.0 licence, de-identified before upload.
Source basis. Every figure below was confirmed against the primary article’s own stored full text (the APA PDF; DOI 10.1037/cns0000342, confirmed on the article’s own title page). One provenance note: the reference database had stored a corrupted short title for this record (“The future failed Bem 2022”); the full title used here is taken verbatim from the article itself.
What the paper reports
Bem’s (2011) “Feeling the Future” reported nine experiments in which well-established psychological effects were time-reversed — the “causal” stimulus was shown after the participant responded — and eight appeared to show above-chance anticipation. This study set out to replicate three of them at large scale. Participants completed two affective-priming tasks and a free-recall task in both the backward (precognition) direction and the ordinary forward direction, the forward version serving as a positive control that the paradigm was working.2 With more than 2,000 online participants the study was powered to detect even small effects. It found none in the backward direction, while the forward controls reproduced the standard effects.
More than 2000 participants participated via the Internet; thus, our study had high statistical power. The results showed no precognition effects at all.
How it was run
- Design. Three experiments, each a within-participants comparison of a backward (precognition) condition against a forward (control) condition. Experiment 1 replicated Bem’s Experiment 3 (affective priming with the prime words “beautiful”/“ugly”); Experiment 2 replicated Bem’s Experiment 4 (fixed positive/negative prime pairs); Experiment 3 replicated Bem’s Experiment 9 (retroactive facilitation of free recall of a 48-word list).
- Setting. Conducted online during the COVID-19 pandemic (2020), with data gathered through an undergraduate course at the University of Bern; participants were invited by email and consented by clicking a link. No experimenter was present, removing in-person experimenter influence.
- Participants. 2,164 participants contributed data across the three experiments (each experiment analysed separately; the per-experiment samples run into the hundreds-to-thousands).
- Outcome measures. For the priming experiments, the reaction-time difference between congruent and incongruent trials; for free recall, the difference in recall between later-practised and control words. The psi prediction was directional, so one-tailed paired t-tests were used.
- Power. The authors sized the study against the smallest effects that would still be plausible and interesting; with n > 2,000 it comfortably exceeded conventional power to detect Bem’s original effect (about d = 0.22) and even substantially smaller ones.
- Openness. Not formally preregistered; datasets and analysis code deposited openly under CC-BY.
Results, as reported
| Metric | Result |
|---|---|
| Exp 1 backward — affective priming (Bem Exp 3) | congruent 1217 ms vs incongruent 1215 ms; t(726) < 1, p = .512, d = 0.001 — no effect |
| Exp 2 backward — fixed-prime priming (Bem Exp 4) | congruent 1170 ms vs incongruent 1168 ms; t(1413) < 1, p = .227, d = 0.034 — no effect |
| Exp 3 backward — retroactive free recall (Bem Exp 9) | practised 11.4 vs control 11.5 words; t(1394) = 1.08, p = .860, d = −0.03 — no effect |
| Exp 1 forward — control (ordinary priming) | congruent 1033 ms vs incongruent 1037 ms; t(704) = 1.84, p = .033, d = 0.07 — ordinary priming present, task valid |
| Total participants | 2,164 across the three experiments |
| Statistical power | >95% for Bem’s original effect (about d = 0.22); adequate for the smallest plausible effect |
| Exploratory post-hoc analyses | a wide battery of variables and questionnaire items was examined; a few reached significance, which the authors flag as needing confirmatory, theory-driven follow-up |
Values are reproduced from the article’s Results section. The three backward (precognition) tests are the confirmatory outcomes; the forward result is a positive control showing the paradigm detects the ordinary, non-time-reversed effect.
Eleven-dimension audit
Pre-registration
The study was not formally preregistered, which the authors state plainly. The confirmatory hypotheses were directional and pre-stated (matching Bem’s original predictions), and the confirmatory tests are clearly separated in the report from the exploratory ones. But without a timestamped public registration, the confirmatory/exploratory boundary and the analysis choices rest on the authors’ own account rather than a locked protocol. This is the study’s main methodological gap, and the authors recommend preregistration for follow-up work.
Randomization
Trial order, prime assignment, and the rehearsed word subset were randomly determined by the software, following the standard priming and free-recall procedures. The task’s critical feature is temporal: in the backward condition the prime or rehearsal set was determined and presented after the participant’s response, so no target existed to be inferred at response time.
Sensory leakage
Classical sensory leakage is excluded by the time-reversed design: the outcome-determining event occurs after the response. As in all Bem-style paradigms, the integrity of that reversal rests on the software’s timing and randomization, delivered here through a controlled online platform with the code shared openly.
Blinding
The study ran online with no experimenter in contact with participants, removing in-person experimenter cueing and expectancy. Participants gave informed consent and were told the study’s general purpose; the report does not describe a specific procedure to keep participants blind to the precognition hypothesis, so participant-level expectancy cannot be fully ruled out, though the forward controls provide a within-study check on task performance.
Optional stopping
The sample size was set by course enrolment rather than by an interim look at the data, so the classic optional-stopping concern (peeking and stopping when results look favourable) does not arise. There was, however, no preregistered stopping rule, so this rests on the design rather than a locked protocol.
Outcome measure
Pre-stated and matched to Bem’s originals: the congruent-minus-incongruent reaction-time difference for the priming experiments, and the practised-minus-control recall difference for free recall, each tested one-tailed in the psi-predicted direction. The measures are well defined and standard.
Effect size
All three backward effects are essentially zero — d = 0.001, 0.034, and −0.03 — far below Bem’s reported effects (about d = 0.22) and below this study’s detection floor. Crucially, the forward control reproduced the ordinary priming effect (d = 0.07, p = .033), demonstrating the paradigm was sensitive, so the backward nulls reflect an absent effect rather than an insensitive task.
Multiple comparisons
The confirmatory tests are few and pre-stated (three backward comparisons plus forward controls), which limits multiplicity there. The real multiple-comparisons exposure is in the large set of exploratory post-hoc analyses across additional variables and questionnaire items; the authors applied Bonferroni correction and, importantly, label the surviving significant effects as exploratory findings that require confirmatory replication rather than as evidence for psi.
Internal replication
Three distinct experiments — two priming paradigms and one free-recall paradigm — all returned null backward effects, and each carried its own forward control. That coherent pattern across paradigms is a stronger internal check than a single null test would be.
External replication
This study is a large, high-powered external replication of Bem (2011), and it did not reproduce the precognition effects. It enters a contested literature: a proponent meta-analysis reports a small positive effect (Bem, Tressoldi, Rabeyron & Duggan, 2015: d = 0.18, Bayes factor > 100),3 while a meta-analysis of retroactive free-recall experiments found an effect not significantly different from zero (Galak et al., 2012: d = .04).4
Transparency
Strong on openness and candour: the datasets and analysis code are openly shared under CC-BY, all three experiments and both directions are fully reported (including the forward controls that could have undercut the story had they failed), and the authors disclose their own limitations directly — the absence of preregistration, the exploratory nature of the post-hoc findings, and sub-conventional power for some secondary correlational analyses. The single missing ingredient, by their own account, is preregistration.
The adversarial record
- Lineage. The target, Bem’s 2011 “Feeling the Future,” is the study most associated with triggering psychology’s reproducibility debate. This replication originated as a University of Bern undergraduate-class project, conducted online during the 2020 pandemic.
- Contested literature. Bem’s claim has been both defended by a proponent meta-analysis (Bem et al., 2015)3 and challenged by failed replications and methodological critiques of researcher degrees of freedom (Wagenmakers et al., 2011; Galak et al., 2012).45 This paper adds a large, clean null to the challenger side.
- Reading against the study. The honest limits are the ones the authors state. Because it was not preregistered, a critic can note that the confirmatory/exploratory line is self-reported; and the exploratory post-hoc battery did yield a few significant results, which the authors themselves caution should not be over-read. A committed proponent may argue that an unsupervised, online, class-run replication is less sensitive than Bem’s original laboratory sessions — to which the paper’s answer is the forward control, which reproduced the ordinary effect and so demonstrates the task was working.
Sources
- Muhmenthaler, M. C., Dubravac, M., & Meier, B. (2022). The future failed: No evidence for precognition in a large scale replication attempt of Bem (2011). Psychology of Consciousness: Theory, Research, and Practice. Advance online publication. https://doi.org/10.1037/cns0000342 R001 [Muhmenthaler et al. 2022] ↩︎
- Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524 R002 [Bem 2011] ↩︎
- Bem, D. J., Tressoldi, P., Rabeyron, T., & Duggan, M. (2015). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, 1188. https://doi.org/10.12688/f1000research.7177.1 R003 [Bem et al. 2015] ↩︎
- Galak, J., LeBoeuf, R. A., Nelson, L. D., & Simmons, J. P. (2012). Correcting the past: Failures to replicate psi. Journal of Personality and Social Psychology, 103(6), 933–948. https://doi.org/10.1037/a0029709 R004 [Galak et al. 2012] ↩︎
- Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2011). Why psychologists must change the way they analyze their data: The case of psi. Journal of Personality and Social Psychology, 100(3), 426–432. https://doi.org/10.1037/a0022790 R005 [Wagenmakers et al. 2011] ↩︎