Wahbeh et al. (2023)
Who’s calling? Evaluating the accuracy of guessing who is on the phone
Wahbeh, H., Cannard, C., Radin, D., & Delorme, A. (2023). Who’s calling? Evaluating the accuracy of guessing who is on the phone. Explore. https://doi.org/10.1016/j.explore.2023.08.008
AI Assessment
A preregistered, automated test of “telephone telepathy,” whether people can guess who is calling before answering, cleverly split into two trial types the participant could not tell apart. When the caller was chosen before the guess (a telepathy design, with a sender who knows they are calling), accuracy was 50.0% against a 33.3% chance baseline (p < .001). When the caller was chosen after the guess (a precognition design, with no sender yet), accuracy was at chance (31.9%, p = .61). The pre-versus-post contrast is a genuine internal control, and the data are open. But the confidence is tempered by significant caveats the authors disclose: several protocol changes were made mid-study, including fixing a randomization bug, in a study whose validity depends entirely on correct randomization, and the moderator findings (genetic relatedness, communication frequency) are exploratory and mostly null. This audit reports what the study found and how; it takes no position on whether telepathy is real.
Provenance
DOI. 10.1016/j.explore.2023.08.008 · Explore, available online 25 August 2023.
Study type. A preregistered, web-automated experiment in triads, with a pre-selected (telepathy) and a post-selected (precognition) trial type randomly interleaved, plus exploratory regression analyses.
Authors. Helané Wahbeh, Cedric Cannard, Dean Radin, and Arnaud Delorme, of the Institute of Noetic Sciences. (A cataloguing note: the reference database had mis-attributed this record to a single unrelated author, a surname-and-topic collision; the correct authorship is Wahbeh and colleagues, as here.)
Data availability. The data are openly deposited on Figshare (doi.org/10.6084/m9.figshare.21521235).
Source basis. Every figure below is taken from the article’s own Abstract and Results.
What the paper reports
Many people report occasionally knowing who is calling before they answer. This study tested that claim under controlled, automated conditions, using two designs that the participant could not distinguish: in pre-selected (telepathic) trials a web server chose the caller before the callee guessed, so a real sender was directing attention; in post-selected (precognitive) trials the server chose the caller after the guess, so no sender existed at guess time.1 The paper reports above-chance accuracy on the telepathic design but not the precognitive one, and treats the result as favouring the psi hypothesis while acknowledging conventional explanations cannot be fully excluded.
We discuss how these results favor the psi hypothesis, although conventional explanations cannot be completely excluded.
How it was run
- Design. A web-automated system (PHP and Twilio) randomly selected the trial type and the caller for each trial; participants did not know which trial type they were in. Each guess had three options (Person X, Person Y, or No one), giving a 33.3% chance expectation.
- Sample. 177 participants completed at least one trial; 105 “completers” finished all 12 trials.
- Analysis. Exact binomial tests with Agresti-Coull 95% confidence intervals for accuracy, a randomization test (10,000 iterations) for the pre-selected path, and a repeated-measures comparison of pre- versus post-selected accuracy, plus exploratory multilevel logistic regressions relating accuracy to genetic relatedness, emotional closeness, communication frequency, and physical distance.
- Protocol changes. Because of persistent recruitment difficulties, several changes were made during the study: inclusion/exclusion criteria, the number of call attempts, the trial-type randomization schema (including fixing a randomization bug), and participant compensation.
Results, as reported
| Metric | Result |
|---|---|
| Sample | 177 participants completed at least one trial; 105 completed all 12 |
| Telepathic / pre-selected trials | 50.0% accuracy across 210 completer trials vs 33.3% chance, p < .001: above chance |
| Precognitive / post-selected trials | 31.9% across 630 completer trials vs 33.3% chance, p = .61: at chance |
| Genetic relatedness (regression, all trials) | significant overall (Wald chi-square = 53.0, p < .001); 25% relatedness OR = 2.88 (beta = 1.06, z = 2.10, p = .04), but other relatedness levels were not significant |
| Communication frequency | significant positive predictor (beta = 0.006, z = 2.19, p = .03) |
| Physical distance and emotional closeness | not significant (p = .12 and p = .06) |
Values are reproduced from the article’s Abstract and Results. The core finding is the contrast between the two designs: accuracy exceeded chance only when a sender was actually directing attention (the telepathic path), while the precognitive path sat exactly at chance.
Eleven-dimension audit
Pre-registration
The study was preregistered, which is a real credit, but the preregistration was substantially amended during data collection: eligibility, call attempts, the randomization schema, and compensation all changed, and a randomization bug was fixed mid-study. The regression analyses are labelled exploratory. So the confirmatory core (pre- versus post-selected accuracy) is preregistered, but the amendments and the mid-study bug complicate that status.
Randomization
Randomization is the mechanism of the whole design (trial type and caller chosen by the server), and it is also the study’s most sensitive point: the authors disclose that the randomization schema was adjusted and a randomization bug fixed during the study. Because the validity of the chance baseline depends on correct randomization, the earlier, pre-fix trials warrant caution.
Sensory leakage
The automated web system and the participant’s blindness to trial type are designed to prevent cues, and the forced three-choice format limits information leakage. The residual concern is whether anything about a real sender’s behaviour or timing in the pre-selected trials could subtly inform the callee, which the authors implicitly acknowledge in saying conventional explanations cannot be completely excluded.
Blinding
Participants were blind to which trial type they were in, and the system scored responses automatically, so assessor bias is minimal. This blinding to trial type is what makes the pre-versus-post contrast interpretable.
Optional stopping
Not classic optional stopping, but the recruitment-driven protocol changes introduced flexibility that a fully fixed design would not have. The trial counts were not stopped on favourable results, but the sampling frame shifted during the study.
Outcome measure
Accuracy of caller identification against a clean 33.3% chance baseline, analyzed with exact binomial tests and a randomization test. The measure and baseline are well specified and appropriate.
Effect size
The telepathic-path effect is moderate and clear (50.0% vs 33.3%, roughly 17 percentage points above chance), while the precognitive path is a clean null. The moderator effects are small and mostly non-significant.
Multiple comparisons
Beyond the two confirmatory accuracy tests, four exploratory regression predictors were examined, of which only 25% genetic relatedness (p = .04, uncorrected) and communication frequency (p = .03) were significant while the other relatedness levels, distance, and emotional closeness were not. These exploratory results should be read as hypothesis-generating rather than established.
Internal replication
The pre-selected result aligns with previously reported telephone-telepathy studies, which the authors note; the internal pre-versus-post contrast is itself a strong design feature, showing the effect appears only where a sender exists.
External replication
This is a partial replication of an established telephone-telepathy paradigm, and it reproduces the above-chance pre-selected result while adding a matched precognitive null. Independent replication with a fixed, un-amended protocol would strengthen the picture.
Transparency
Strong: preregistered, openly deposited data, a measured conclusion, and, importantly, full disclosure of the mid-study protocol changes and the randomization bug rather than concealment. The transparency is what lets a reader properly discount those changes.
The adversarial record
- The randomization bug. For a study whose entire inference rests on a correct chance baseline, disclosing that a randomization bug existed and was fixed mid-study is the most serious caveat. It is to the authors’ credit that they report it, but it means the pre-fix data cannot be treated with full confidence.
- Amended preregistration. Changing eligibility, call attempts, the randomization schema, and compensation during the study widens the researcher degrees of freedom and weakens the “preregistered” guarantee, even though each change is attributed to recruitment difficulty.
- Exploratory moderators. The genetic-relatedness and communication-frequency findings are exploratory, uncorrected, and inconsistent (only one relatedness level is significant), so they should not be read as establishing that biological or social closeness drives accuracy.
- Author context. The team is from a proponent research institute, and the telephone-telepathy paradigm has a long, contested history. The authors’ own hedge, that conventional explanations cannot be completely excluded, is the appropriate posture.
- What is genuinely strong. The pre-versus-post design is an elegant internal control, the automation and blinding to trial type are real safeguards, the chance baseline is clean, the data are open, and the deviations are disclosed rather than hidden. On its own terms it is an honestly reported above-chance telepathic-path result with clearly flagged limitations.
Sources
- Wahbeh, H., Cannard, C., Radin, D., & Delorme, A. (2023). Who’s calling? Evaluating the accuracy of guessing who is on the phone. Explore. https://doi.org/10.1016/j.explore.2023.08.008 R001 [Wahbeh et al. 2023] ↩︎