Kekecs et al. (2023)

Raising the value of research studies in psychological science by increasing the credibility of research reports: the transparent Psi project

Kekecs, Z., Palfi, B., Szaszi, B., Szecsi, P., Zrubka, M., Kovacs, M., … Aczel, B. (2023). Raising the value of research studies in psychological science by increasing the credibility of research reports: the transparent Psi project. Royal Society Open Science, 10(2), 191375. https://doi.org/10.1098/rsos.191375

AI Assessment

A 10-laboratory, preregistered registered-report replication of Bem’s (2011) precognition experiment, built from the ground up as a credibility showcase — born-open data, tamper-evident software, formal external audit, and a design co-authored by a consensus panel of ESP proponents and opponents. Across 37,836 trials from 2,115 participants it found no above-chance guessing (49.89% versus a 50% chance baseline); all four pre-specified statistical tests supported the no-effect model at the minimum planned sample size. Methodologically it is among the most rigorous studies in this literature; the honest limits it states are that a design powered for a 51% effect cannot exclude a still-smaller one, and that one well-run replication of a single experiment cannot settle the field. This audit describes what the paper reports and how it was run; it takes no position on whether precognition exists.

Provenance

DOI. 10.1098/rsos.191375 · Royal Society Open Science 10(2):191375 (2023) · gold open access (published by the Royal Society under a Creative Commons licence).

Study type. A preregistered Stage-2 Registered Report: a multi-laboratory (10 laboratories, 9 countries) direct replication of Experiment 1 of Bem (2011), paired with a demonstration of a set of research-transparency tools.

Authors. Thirty co-authors led by Zoltan Kekecs, with Balazs Aczel as senior author (Eotvos Lorand University, Budapest, and collaborating sites). The consensus design panel that fixed the protocol comprised 29 experts — 15 who considered ESP plausible and 14 who did not.

Funding. The project was funded by the Bial Foundation (grant no. 122/16); individual co-authors report additional support in the article’s funding statement.

Ethics. Testing was conducted with the approval of the Eotvos Lorand University Faculty of Education and Psychology, and each site’s local requirements.

Data availability. Born-open: raw data were deposited directly and streamed live to a public GitHub repository throughout collection; all data, analysis code, and materials are archived on the Open Science Framework (preregistration osf.io/a6ew3), and the server-side software was version-controlled in a GitLab repository made public after data collection ended.

Source basis. Every figure below was confirmed against the primary article’s own stored full text (the born-digital Royal Society Open Science PDF; DOI DOI-verified against Crossref). The text layer is clean; no extraction defects affect the audited statistics.

What the paper reports

The study asked whether Bem’s (2011) headline precognition finding — that people guess the future random position of an erotic image slightly better than chance — survives when the known sources of bias in psychological research are removed by design. Participants made a two-alternative forced-choice guess (left or right curtain); only after the guess did a random process determine which side actually hid the image.1 Following Bem, the confirmatory hypothesis test was restricted to the erotic trials, and the models were one-sided (M1: success rate above chance; M0: not above chance). Across 2,115 participants and 37,836 erotic trials the observed success rate was 49.89%, against Bem’s originally reported 53.07%, and all four pre-specified tests supported M0 — no evidence of precognition.2

We found 49.89% successful guesses, while Bem reported 53.07% success rate, with the chance level being 50%… All four tests used for the primary hypothesis testing supported M0, thus the stopping rule was triggered with our minimum sample size achieved.

How it was run

Results, as reported

MetricResult
Primary analysis — erotic trials (n = 2,115; 37,836 trials)49.89% successful guesses vs 50% chance; all four pre-specified tests support M0 (no effect); stopping rule triggered at the minimum sample size
Bem (2011) original, for comparison53.07% in 1,560 erotic trials
Mixed-effects logistic regressionestimate 0.4989; 99.75% CI 49.11%–50.67%; upper bound < 51%, so supports M0
Three Bayesian proportion tests (uniform / BUJ / replication priors)BF01 > 25 for all three priors, i.e. ≥25× more likely under M0 than M1
Robustness — frequentist proportion testchance < 51%: p < 0.001 (supports M0); chance > 50%: p = 0.665 (not significant)
Robustness — Bayesian parameter estimationposterior mode 0.5002; 90% HDI 49.57%–50.40%; >95% of the posterior inside the 0–50.6% region of practical equivalence; 0.96% probability the true rate exceeds 50.6%
Exploratory — sheep-goat (subgroup of gifted guessers)odd/even-trial correlation r = 0.026, 95% CI (−0.017, 0.069); no heavy-tailed high-ability subgroup
Exploratory — experimenter belief (ASGS) effect on performanceestimate −0.0003, 95% CI (−0.001, 0.001); no experimenter psi effect detected

Values are reproduced from the article’s Results section and text. A minor internal inconsistency in the paper: Bem’s original erotic-trial count is given as 1,560 (consistent with its own 828 successes + 732 failures) in two places but as “1,650 trials” in the sample-size section; the internally consistent 1,560 is used here. Bem’s rate appears as both 53.1% (rounded) and 53.07%.

Eleven-dimension audit

Pre-registration

Exemplary. This is a Stage-2 Registered Report: the design and analysis plan were preregistered on OSF (osf.io/a6ew3) after in-principle acceptance and before any data collection, and the Stage-1 protocol pre-specified the conclusion text for every possible outcome, including the abstract. Confirmatory analyses were locked in advance and separated explicitly from exploratory ones. This is the strongest form of preregistration available.

Randomization

The target side was randomly determined by the software after each participant’s guess, so at the moment of the guess no target existed to be inferred. There is no participant-allocation randomization to balance (a single within-participant guessing task), and the randomization method was fixed in the preregistered protocol and discussed as an ESP-specific consideration in the supplement.

Sensory leakage

Classical sensory leakage is impossible by design: because the target is generated after the guess, there is no contemporaneous target state that a cue could reveal. The analogous threat in a multi-site online study is manipulation of the software or data, which the tamper-evident architecture (central version-controlled server, born-open data with a GitHub audit trail, independent IT audit) was specifically built to foreclose.

Blinding

No party could know the future-random target, so the outcome-relevant blind is structural. Experimenter and site-leader belief in the paranormal was not merely controlled but measured (ASGS) and tested as a moderator; it showed no effect on participant performance, addressing the “psi experimenter effect” directly. Trial sessions were video-recorded to verify experimenter conduct (the videos were withheld from public release only to protect participant identity).

Optional stopping

A particular strength. The sequential plan had five pre-specified analysis points, and the analysis was deliberately engineered so that stopping could not bias the result: the Bayesian proportion tests pool successes across all trials irrespective of participant, and all completed erotic trials were retained regardless of whether a session finished. Data collection stopped at the first (minimum) analysis point because all four tests already agreed — a stopping rule fixed in advance, not chosen after seeing the data.

Outcome measure

Pre-stated and singular: the success rate on erotic trials only, one-tailed, matching the exact quantity in which Bem reported his original effect. Non-erotic trials were collected to preserve the original protocol but excluded from the confirmatory test by prior rule, removing an obvious researcher degree of freedom.

Effect size

Essentially zero: 49.89% success against a 50% baseline, versus Bem’s 53.07%. The 99.75% confidence interval (49.11%–50.67%) excludes the 51% effect the study was powered to detect. The stated limitation is symmetric and honest: simulations show the design has >95% power to detect a true rate of 51% or higher, but its sensitivity falls away below about 50.7%, and a true rate of 50.2% or lower would still yield strong (false) support for M0 in most experiments — so a very small effect cannot be excluded, only a Bem-sized one.

Multiple comparisons

Handled conservatively. Rather than one test, four had to agree before any conclusion could be drawn, and the sequential confidence intervals were Bonferroni-widened across the five analysis points (99.75% at the first point, widening thereafter). Exploratory analyses were reported separately and explicitly barred from influencing the confirmatory conclusion or abstract.

Internal replication

Built into the design: 10 independent laboratories in 9 countries collected data on the same central software, so the result is a within-study multi-site aggregate rather than a single lab’s finding. Site leaders ranged widely in paranormal belief (ASGS 0–23), reducing the chance that one believer- or skeptic-heavy lab drove the outcome.

External replication

This study is the external replication — a high-powered, adversarially-designed direct replication of Bem (2011) Experiment 1. It did not reproduce the original above-chance effect. The authors are careful that a clean replication of one experiment does not adjudicate the wider precognition literature, which includes a contested meta-analysis reporting small positive effects.4

Transparency

The paper’s central subject and its strongest dimension. Beyond preregistration, it demonstrates a stack of credibility tools rarely seen together in psychology: born-open data streamed live to GitHub during collection, preregistered and public analysis code, cloud-hosted tamper-evident software with a git audit trail, and a formal external audit by independent research and IT auditors who published their own reports. The authors also disclose the tools’ limits candidly — a potential conflict of interest (two of the three auditors work at the University of Padova, where one collaborating laboratory was located, and the IT auditor has co-published with that site’s lead researcher), and the withholding of trial-session videos to protect participant identity.

The adversarial record

Sources
  1. Kekecs, Z., Palfi, B., Szaszi, B., Szecsi, P., Zrubka, M., Kovacs, M., … Aczel, B. (2023). Raising the value of research studies in psychological science by increasing the credibility of research reports: the transparent Psi project. Royal Society Open Science, 10(2), 191375. https://doi.org/10.1098/rsos.191375 R001 [Kekecs et al. 2023] ↩︎
  2. Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524 R002 [Bem 2011] ↩︎
  3. Wagenmakers, E.-J., Wetzels, R., Borsboom, D., & van der Maas, H. L. J. (2011). Why psychologists must change the way they analyze their data: The case of psi. Journal of Personality and Social Psychology, 100(3), 426–432. https://doi.org/10.1037/a0022790 R003 [Wagenmakers et al. 2011] ↩︎
  4. Bem, D. J., Tressoldi, P., Rabeyron, T., & Duggan, M. (2015). Feeling the future: A meta-analysis of 90 experiments on the anomalous anticipation of random future events. F1000Research, 4, 1188. https://doi.org/10.12688/f1000research.7177.1 R004 [Bem et al. 2015] ↩︎