Stedall & Tressoldi (2025)

Who is calling? An Independent Replication of a Telephone Telepathy Test

Stedall, T., & Tressoldi, P. (2025). Who is calling? An Independent Replication of a Telephone Telepathy Test. Journal of Scientific Exploration, 39(2), 203–206. https://doi.org/10.31275/20253543

AI Assessment

A clean, preregistered replication that found no telephone telepathy; its one positive signal is a post-hoc subgroup, not the planned test. Across 604 trials, 117 people tried to guess which of two personally known friends was calling their smartphone, and they were right 48.6% of the time, no better than a coin flip. The preregistered prediction that the hit rate would reach at least 55% failed. A later, unplanned split of the data suggested that people who said they guessed intuitively beat chance while those who reasoned about it fell below chance, but that comparison was exploratory, not the test the study was built to run.

Provenance

DOI. 10.31275/20253543 · Platinum open access, CC BY-NC.

Study type. Preregistered forced-choice telephone telepathy test: the receiver guesses which of two personally known callers is phoning. Classical sender-and-receiver telepathy paradigm, two-alternative choice, 50% chance baseline.

Funding. No funding statement is given in the brief report.

Data availability. Raw data are openly archived on figshare (10.6084/m9.figshare.24574174); the study was preregistered on the Koestler Registry (KPU Registry 1048).

Source basis. Figures confirmed against the primary article (Journal of Scientific Exploration, 39(2), 203–206; local copy).

What the paper reports

Across 604 trials, 117 receivers tried to identify which of two personally known friends was calling their smartphone, a task with a 50% chance baseline. They were correct 294 times, a 48.6% hit rate, marginally below chance. The preregistered prediction, that the hit rate would reach at least 55%, was not met: a one-tailed binomial test gave z = −0.61, p = .76. A post-hoc split by self-reported guessing strategy found that receivers who described an intuitive approach scored 55.7% (220 / 395) while those who described a rational approach scored 35.4% (74 / 209).1

The only above-chance result is the post-hoc intuitive subgroup (55.7%, p = .027, uncorrected for multiple comparisons), not the preregistered test. The confirmatory finding is a clean null: no overall evidence of telephone telepathy.

How it was run

Results, as reported

MetricResult
Overall hit rate48.6% (294/604) vs 50% chance
Preregistered test (hit rate ≥ 55%)z = −0.61, p = .76 (one-tailed)
Intuitive strategy (post-hoc)55.7% (220/395) vs 50% chance, z = 2.21, p = .027
Rational strategy (post-hoc)35.4% (74/209) vs 50% chance, z = −4.28, p = .00003

The two strategy-subgroup tests are post-hoc and exploratory; they are not corrected for multiple comparisons, and the authors present them as hypothesis-generating rather than confirmatory. All of the binomial tests here treat the 604 trials as independent; because each receiver contributed six trials, the trials are clustered within participants, so the tests modestly overstate statistical precision.

Eleven-dimension audit

Pre-registration

Strong. The study was preregistered on the Koestler Registry (KPU Registry 1048) before data collection, with a specific directional prediction (hit rate at least 55%). The confirmatory test is therefore a genuine prediction, not a post-hoc fit.

Randomization

The caller on each trial was selected by the PHP mt_rand function, a Mersenne Twister generator seeded from time, location, and system variables. The authors explicitly note it is not suitable for cryptographic use. Adequate for a two-alternative forced-choice task, though a hardware or vetted random source would be stronger.

Sensory leakage

Well controlled by the architecture. Caller and receiver interact only through the automated phone system; the caller waits on hold and the receiver is connected only after guessing, so the outcome is unknown at the moment of choice. A research assistant supervised the receiver’s room to block ordinary communication. The callers are personal acquaintances, but trial times are system-determined and not disclosed in advance.

Blinding

The receiver is blind to the caller’s identity until after the guess is logged; the system, not a person, selects the caller; alphabetical key ordering removes position cues. Effectively blind by design.

Optional stopping

A target of 600 trials (100 receivers times six) was set in advance, and recruitment was extended as planned to 604 because some receivers contributed fewer than six trials. No sign of data-dependent stopping.

Outcome measure

The primary outcome, overall hit rate against 50% tested against the 55% threshold by a one-tailed binomial, was pre-declared. The strategy comparison is explicitly labeled post-hoc.

Effect size

The confirmatory result is effectively null: 48.6% against 50%, z = −0.61. The exploratory intuitive subgroup reached 55.7% (z = 2.21), a small effect; the rational subgroup fell well below chance (35.4%, z = −4.28).

Multiple comparisons

The preregistered test is single, which is a strength. The post-hoc strategy split runs two further binomial tests without correction; under any multiple-comparison adjustment the intuitive subgroup’s p = .027 weakens. The authors flag the analysis as exploratory. A related caveat is trial dependence: with six trials per receiver, the 604 trials are not fully independent, so the binomial tests, which assume independence, slightly overstate precision.

Internal replication

None within the study, which is a single experiment. It does, however, function as an attempted replication of two earlier positive studies, and it did not reproduce their above-chance overall hit rates.

External replication

This is an independent external replication of the telephone telepathy paradigm, specifically the fourth of Sheldrake’s automated procedures. Its overall null (48.6%) contradicts Sheldrake & Stedall (2024), which reported 57% (p = .01)2, and Wahbeh et al. (2024), which reported 56.3% after removing no-call guesses (p = .02).3

Transparency

High. Platinum open access under CC BY-NC, a public preregistration on the Koestler Registry, raw data openly archived on figshare, and the full automated procedure described in the paper.

The adversarial record

Sources
  1. Stedall, T., & Tressoldi, P. (2025). Who is calling? An Independent Replication of a Telephone Telepathy Test. Journal of Scientific Exploration, 39(2), 203–206. https://doi.org/10.31275/20253543 R001 [Stedall 2025] ↩︎
  2. Sheldrake, R., & Stedall, T. (2024). A comparison of four new automated telephone telepathy tests. Journal of Anomalous Experience and Cognition, 4, 122–141. https://doi.org/10.31156/jaex.25250 R002 [Sheldrake 2024] ↩︎
  3. Wahbeh, H., Cannard, C., Radin, D., & Delorme, A. (2024). Who’s calling? Evaluating the accuracy of guessing who is on the phone. EXPLORE, 20, 239–247. https://doi.org/10.1016/j.explore.2023.08.008 R003 [Wahbeh 2024] ↩︎
  4. Lawrence, T. (1993). Bringing home the sheep: A meta-analysis of sheep/goat experiments. In Proceedings of the 36th Annual Convention of the Parapsychological Association. R004 [Lawrence 1993] ↩︎