Daryl J. Bem experiments and data
Daryl J. Bem, PhD is a social psychologist at Cornell University whose entry into parapsychology research — particularly his 2011 paper in a flagship mainstream psychology journal — made his work among the most debated in the field’s recent history. What follows covers his experimental record, methodology, data, and the critical responses his work generated.
Experiments
Bem’s central contribution is a series of nine experiments reported in a single 2011 paper in the Journal of Personality and Social Psychology, titled “Feeling the Future” [3]. Each experiment took a well-validated procedure from mainstream social psychology — priming, habituation, mere exposure, facilitation of recall — and reversed its causal direction, testing whether a stimulus presented after a response could influence that response retroactively. The nine covered distinct psychological phenomena: retroactive facilitation of recall, retroactive priming, retroactive habituation, and precognitive avoidance of negative stimuli, among others.
Following that paper, Bem collaborated with Patrizio E. Tressoldi, Thomas Rabeyron, and Michael W. Duggan on a large-scale meta-analysis published in 2016 in F1000Research, synthesizing 90 experiments from 33 laboratories across 14 countries that had attempted to replicate or extend the anomalous-anticipation paradigm [2].
A later multi-investigator replication study — co-authored with Marilyn J. Schlitz, David Marcusson-Clavertz, Etzel Cardeña, and a large international team — tested a time-reversed priming task specifically, examining whether experimenter orientation (belief vs. skepticism toward psi) modulated results [1].
Methodology
Bem’s core methodological strategy was deliberate: adapt standard, peer-reviewed social-psychology protocols and run them in reverse temporal order, so that the putative cause (a randomly selected stimulus) came after the participant’s response. This grounded his work within mainstream experimental psychology’s validated effect landscape while testing for anomalous retroactive influence.
For example, a standard mere-exposure paradigm shows that repeated presentation of a stimulus increases liking for it; Bem’s retroactive version presented the stimulus after participants rated their liking, testing whether future exposure influenced past ratings. Because the stimuli were randomly assigned after responses were recorded, there is no conventional mechanism by which congruence or incongruence between prime and picture could be anticipated — as the 2021 replication study notes, “there is no (non-psi) way for a participant to anticipate the kind of trial currently on the screen” [1].
The 2021 multi-site replication study introduced an additional variable: experimenter orientation. The design assigned participants to either a “psi-positive” or “psi-skeptical” experimenter condition, testing whether the experimenter’s belief state interacted with participant outcomes. Four hypotheses were pre-specified, including a main psi effect and an experimenter × participant interaction [1].
The 2016 meta-analysis classified replications into three categories — exact replications of one of Bem’s experiments (31 experiments), modified replications (38 experiments), and independently designed experiments assessing anomalous anticipation by alternative means (the remainder) — and applied both frequentist and Bayesian analysis, including p-curve analysis to assess whether the effect distribution was consistent with a genuine effect or with selective reporting [2].
Data
No structured evidence block has been provided for this question, so the figures below come exclusively from the numbered sources, carried in prose with table format where comparisons are useful.
From the original nine experiments [3]:
| Summary statistic | Value |
|---|---|
| Experiments with statistically significant results | 8 of 9 |
| Combined Stouffer z | 6.66 |
| p (two-tailed) | approximately 2.68 × 10⁻¹¹ |
| Mean effect size (d) | 0.22 |
These figures are reported in [3] in Bem’s reply to critics; the retrieved excerpt confirms them directly.
From the 2016 meta-analysis of 90 experiments [2]:
| Metric | Value |
|---|---|
| Laboratories / countries | 33 / 14 |
| Overall z | 6.40 |
| p | 1.2 × 10⁻¹⁰ |
| Effect size (Hedges’ g) | 0.09 |
| Bayes Factor | 5.1 × 10⁹ |
| Effect size excluding Bem’s own studies (ES) | 0.06 |
| z excluding Bem’s studies | 4.16 |
| p excluding Bem’s studies | 1.1 × 10⁻⁵ |
| Bayes Factor excluding Bem’s studies | 3,853 |
| P-curve true effect size estimate (full database) | 0.20 |
| P-curve true effect size estimate (independent replications) | 0.24 |
The meta-analysis notes that these p-curve estimates are “virtually identical” to the mean effect size of Bem’s original experiments (0.22) and to a separate meta-analysis of presentiment experiments (0.21) [2].
For the time-reversed priming paradigm specifically, the 2016 meta-analysis of 15 precognitive priming experiments reported an effect size of d = 0.11, p = 0.003 [1].
From the 2021 multi-site replication [1]:
The two-study replication of the time-reversed priming task returned null results on its primary analyses — across all data transformations and cutoffs examined, the null hypothesis could not be rejected, and binomial tests of whether participants scored positively at above-chance rates were consistent with those primary null findings [1]. The experimenter-orientation hypothesis was also tested; the retrieved excerpt does not carry the final interaction statistics in intact form, so those are not reported here.
Important metric note: the effect sizes above span different metrics (Cohen’s d, Hedges’ g, a raw Rosenthal effect size z/√n) and different paradigms. They are not poolable into a single figure and are presented separately for each study and meta-analysis.
Skeptical critiques
What critics argue.
Wagenmakers, Wetzels, Borsboom, and van der Maas (2011) argued that the analyses in Bem’s nine-experiment paper were partly exploratory rather than confirmatory, and that psychologists should replace standard frequentist analysis with Bayesian methods — under which, they contended, the evidence for psi in Bem’s data was considerably weaker than the frequentist p-values suggested [3]. Applying a Cauchy prior, they reported Bayes Factors ranging from approximately 0.13 to 1.82 across Bem’s nine experiments, most below the threshold for positive evidence [3]. Schimmack (2012), as noted in [1], argued that the original experiments were low-powered, which raises the probability of false positives in the reported results. Francis (2012), cited in [2], argued that the pattern of significant results across Bem’s low-powered studies strongly implies unreported null experiments or p-hacking, with the effect size estimates artificially inflated.
What the experimental data show.
Bem, Utts, and Johnson replied directly to Wagenmakers et al. (2011), arguing that the choice of prior distribution in Bayesian analysis is inherently subjective, and that a knowledge-based prior — informed by prior psi research — yielded Bayes Factors of 1.76 to 5.35 across most experiments, constituting at least some positive evidence [3]. On the p-curve critique, the 2016 meta-analysis reported that p-curve analysis of the 90-experiment database estimated a true effect size of 0.20 (full database) and 0.24 (independent replications only), figures the authors describe as “virtually identical” to the mean effect size from the original nine studies, which they take as evidence against the inflated-by-selection-bias interpretation [2]. On the question of independent replication, the 2016 meta-analysis reported that excluding Bem’s own studies, the remaining independent experiments still yielded z = 4.16 and ES = 0.06 [2]. The 2021 multi-site replication by Schlitz, Bem, and colleagues, however, returned null results on the primary analyses of the time-reversed priming task [1].
Analysis.
The exchange between Bem and Wagenmakers et al. is a documented instance of the broader replication-and-Bayesian-methods debate that Bem’s 2011 paper helped catalyze across experimental psychology. The 2016 meta-analysis directly addresses the prior-specification dispute and the p-hacking concern, but Wagenmakers et al.’s reply — cited in [2] — was prepared, by the authors’ own acknowledgment in [2], without reading the p-curve section of the meta-analysis. The 2021 multi-site replication produced null results, while the 2016 meta-analysis of independent replications did not. These are different paradigms, different investigator teams, and different outcome patterns; the two do not straightforwardly reconcile.
For a broader view of how Bem’s work sits within the ganzfeld literature and the replication debate, the Replication and Meta-Analysis Debates page covers the surrounding context, and Bem’s full profile is at Daryl J. Bem, PhD.
References
- Schlitz, M. J., Bem, D. J., Marcusson-Clavertz, D., Cardena, E., Lyke, J., Grover, R., Blackmore, S. J., Tressoldi, P. E., Roney-Dougal, S. M., Bierman, D. J., Jolij, J. J., Lobach, E., Hartelius, G., Rabeyron, T., Bengston, W. F., Nelson, S. E., Moddel, G., & Delorme, A. (2021). Two Replication Studies of a Time-Reversed (Psi) Priming Task and the Role of Expectancy in Reaction Times. Journal of Scientific Exploration, 35, pp. 65–90. https://doi.org/10.31275/20211903
- Bem, D. J., Tressoldi, P. E., Rabeyron, T., & Duggan, M. W. (2016). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, article 1188. https://doi.org/10.12688/f1000research.7177.2
- Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100, 407–425. https://doi.org/10.1037/a0021524
More questions answered
- Pool the effect sizes across every ganzfeld study you hold and compare your number to the published Storm and Tressoldi meta-analyses—do they agree?
- Has the effect size for PK data increased in the last 50 years?
- Tell me about the clairvoyance work of the last 50 years
- What trends can you see in precognition research?
- What trends can you see in ESP research in the last 50 years?
- What are parapsychology's current arguments to justify that the phenomena are real and should be taken seriously?