Tressoldi et al. (2024)
Stage 2 Registered Report: Testing mediumship accuracy with a triple-blind protocol
Tressoldi, P. E., Liberale, L., & Sinesio, F. (2024). Stage 2 Registered Report: Testing mediumship accuracy with a triple-blind protocol. Registered Report (Stage 1: F1000Research, 10, 351). https://doi.org/10.31234/osf.io/a756f
AI Assessment
A triple-blind Registered Report testing whether self-defined mediums can retrieve accurate information about deceased strangers, using 64 readings scored by the sitters themselves. Of three preregistered hypotheses, only one was confirmed: sitters rated the reading intended for their own deceased slightly higher on a 0-to-6 scale than a control reading (1.48 vs 1.09, p = .01), but at a weak Bayes factor of 2.8. Reading identification did not beat chance (53.1%, p = .08), and the correct-minus-incorrect information difference was non-significant (p = .06). The design is a genuine strength (three levels of blinding that shut down cold reading), but the result is predominantly null, one confirmed hypothesis is only anecdotally supported, the stronger post-hoc finding was not preregistered, and 78% of the readings concerned just 11 deceased people, so the readings are not independent. This audit reports what the paper found and how it was run; it takes no position on survival of consciousness.
Provenance
DOI. 10.31234/osf.io/a756f · Stage-2 Registered Report, version dated 22 February 2024. The Stage-1 protocol was published as Tressoldi, Sinesio, Liberale & Testoni (2021), F1000Research, 10:351.2
Study type. A preregistered triple-blind experiment. To our knowledge, the authors state, this is the first Registered Report on mediumship.
Authors. Patrizio E. Tressoldi (Science of Consciousness Research Group, Studium Patavinum, University of Padua, Italy) with Laura Liberale and Fernando Sinesio (Gruppo di Ricerca Italiano sulla Medianità).
Funding. Funded by the Bial Foundation under contract 343/20. The Bial Foundation is a long-standing funder of parapsychological research; the authors disclose the grant.
Data availability. The reading database (RRDatabase.xlsx) is openly deposited on Figshare (figshare.com/articles/dataset/Mediumship/13311710) under CC-BY 4.0.
Source basis. Every figure below is taken from the article’s own Results section and reported statistics. The article carries a few surface typos (for example “olus” for “plus,” “This stud was funded”) and mixes future-tense protocol language with past-tense results, an artifact of a Stage-1 protocol carried into the Stage-2 report; none of these affect the reported numbers.
What the paper reports
The study asks whether self-defined mediums can retrieve accurate information about a specific deceased person without any conventional route to that information, in particular without cold reading (verbal and non-verbal cues from the sitter). To exclude that route it uses three levels of blinding, and it asks the sitters, who are the only people able to judge accuracy, to evaluate two anonymized readings (one intended for their deceased, one a same-gender control) without knowing which is which.1 Three hypotheses were preregistered: that intended readings would be identified above chance, would earn higher overall scores, and would carry a higher correct-minus-incorrect information balance than controls.
Among the three main hypotheses … Only hypothesis (b) was confirmed, [plus] the comparison of correct information between the intended and control readings added at this stage of the study.
How it was run
- Mediums and sitters. Eight professional self-defined mediums, prescreened for accuracy in a prior study, were planned; because of unavailability, five more mediums and sixteen more sitters were recruited to reach the target of 64 readings. Sitters were people with a serious interest in a reading for a deceased relative or friend.
- Triple-blind procedure. Research assistant B held the deceased’s identifying details. Research assistant A contacted the medium (by Skype or WhatsApp) and gave only the deceased’s first name. The medium answered ten fixed questions (plus a few sitter-specific ones) for each of two paired, same-gender deceased. Research assistant A transcribed the reading, stripping generic statements, and sent it to B. The two anonymized readings were then sent to the sitter to evaluate, blind to which was intended for their deceased.
- Scoring. Each detailed item marked “perfectly correct” or “clearly wrong” scored +1 or −1; each less-detailed item scored +.5 or −.5; “somewhat correct” scored 0.5. Sitters also gave each reading an overall 0-to-6 score (Beischel’s scale).3 An independent judge re-scored all readings for reliability.
- Analysis. Pre-stated and directional (one-sided, alpha .05), run in both frequentist and Bayesian forms in JASP: a binomial test for identification, and one-sided permutation t-tests (5000 bootstrap samples) plus Bayesian dependent-sample t-tests for the score and information comparisons.
- Power. The Stage-1 plan set 64 readings as sufficient (power .80) to detect a 1.5-point overall-score difference, a 15% above-chance identification rate, or a 25% correct-information difference, taken as the smallest effects of interest.
Results, as reported
| Metric | Result |
|---|---|
| H1: reading identification (preregistered) | 34 of 64 correct (53.1%); binomial z = .38, p = .08, Bayes factor H0/H1 = 4.1: not above chance (not confirmed) |
| H2: overall 0-to-6 score (preregistered) | intended 1.48 (SD 1.18) vs control 1.09 (SD 1.2); paired difference 0.391 (95% CI 0.0469, 0.719); permutation t-test p = .01, Bayes factor H1/H0 = 2.8: confirmed |
| H3: correct-minus-incorrect information (preregistered) | intended −31.7% (SD 33.6) vs control −40% (SD 39.7); paired difference 0.082 (95% CI −.028, .17); p = .06, Bayes factor H0/H1 = 1.19: not confirmed |
| Percentage of correct information (post-hoc, added at Stage 2) | intended 24.6% (SD 14.6) vs control 19.3% (SD 15.9); paired difference 5.3% (95% CI 1.2, 9.2); p = .005, Bayes factor H1/H0 = 6.4: supported, but not preregistered |
| Readings and sources | 64 readings from 13 mediums; 50 of 64 (78%) concerned only 11 deceased persons (each consulted 2 to 8 times) |
| Internal comparison | results lower than the authors’ own 2022 study using the identical protocol |
Values are reproduced from the article’s Results section. One of three preregistered hypotheses was confirmed (the overall-score comparison), at a Bayes factor of 2.8, which on conventional scales is only anecdotal evidence; the stronger result (the percentage of correct information, Bayes factor 6.4) was added post-hoc at Stage 2 and is labelled a deviation from the plan.
Eleven-dimension audit
Pre-registration
A genuine Stage-2 Registered Report: the three hypotheses, scoring rubric, sample size, and one-sided analyses were all fixed at Stage 1 (Tressoldi et al., 2021). Two deviations are disclosed: extra mediums and sitters were recruited to reach the planned 64 readings, and a fourth analysis (percentage of correct information) was added at Stage 2. The authors flag the latter as post-hoc, which is the correct and transparent handling, but it means the study’s strongest number is not a confirmatory one.
Randomization
Each trial paired two same-gender deceased, and the sitter evaluated both readings without knowing which was intended for their loved one. The pairing and the blind evaluation remove any systematic cue to the correct reading; the identification test used chance probability of one-half accordingly.
Sensory leakage
This is the study’s central concern and its main strength. Cold reading, the ordinary explanation for apparent mediumship, is blocked because the medium never meets or hears the sitter, the interviewer (research assistant A) knows only a first name, and generic statements are stripped before scoring. The design is built specifically to close the sitter-to-medium information channel.
Blinding
Triple-blind, as named: the medium is blind to the deceased’s identity (first name only), research assistant A is blind to identifying details, and the sitter is blind to which of the two readings is the intended one. This is the strongest form of blinding used in laboratory mediumship research.
Optional stopping
The sample size was set in advance by the Stage-1 power analysis (64 readings). The additional mediums and sitters were recruited to reach that fixed target, not in response to interim results, so optional stopping is not a concern here.
Outcome measure
Three pre-stated measures with a defined scoring rubric (detailed items scored plus or minus 1, less-detailed plus or minus 0.5, “somewhat correct” 0.5), plus an overall 0-to-6 score on Beischel’s established scale, and an independent judge re-scored all readings for reliability. The measures are well specified and standard for this literature.
Effect size
The effects are small and mostly null. The one confirmed hypothesis is a 0.39-point difference on a 0-to-6 scale (Bayes factor 2.8, anecdotal). Identification (53.1%) sits essentially at the 50% chance line, with the Bayesian test actually favouring the null (H0/H1 = 4.1). The correct-minus-incorrect balance was non-significant. The post-hoc correct-information measure is the only moderately strong signal (Bayes factor 6.4), and it was not preregistered.
Multiple comparisons
Three preregistered directional tests plus one post-hoc test. That is a small, contained set, but with only one preregistered confirmation at an anecdotal Bayes factor, and the strongest result coming from the added analysis, the overall evidential picture is modest rather than decisive.
Internal replication
The study explicitly under-performed the authors’ own 2022 study, which used an identical protocol and reported stronger accuracy.5 The authors attribute the drop to a structural feature of this sample: 78% of the readings concerned only 11 deceased persons, each consulted multiple times, so the readings are not independent, which both violates an assumption of the analysis and may have depressed accuracy.
External replication
The result is weaker than, but broadly in the family of, the wider blinded-mediumship literature: a meta-analysis by Sarraf, Woodley of Menie, and Tressoldi (2020) reported accuracy 6 to 14% above chance, with better performance from prescreened mediums.4 This study adds a rigorously blinded, mostly-null data point to that record.
Transparency
Strong: the Registered Report format, the open Figshare dataset, the disclosed Bial Foundation funding, the independent re-scoring, and the explicit reporting of deviations and of the non-independence problem all sit on the credit side. The debits are cosmetic (surface typos, mixed protocol tense) and one substantive limitation that the authors themselves surface: the 11-deceased clustering that undermines the independence of the readings.
The adversarial record
- Mostly null. The honest headline is that one of three preregistered hypotheses was confirmed, and that one only at an anecdotal Bayes factor (2.8), while the identification test slightly favoured the null. A skeptic can fairly read this as a predominantly negative result presented with a positive gloss.
- Preregistered vs post-hoc. The strongest number in the paper, the percentage-of-correct-information difference (Bayes factor 6.4, p = .005), was added at Stage 2 and is not part of the confirmatory plan. Weighting it as if it were confirmatory would over-read the evidence; the authors correctly label it a deviation.
- Non-independent readings. With 50 of 64 readings concerning just 11 deceased people, the effective sample is far smaller than 64 and the observations are correlated. The authors raise this themselves and note they cannot say whether it caused the lower accuracy, which is the appropriately cautious position.
- Funding. The work is Bial Foundation funded, a body that supports parapsychological research; this is disclosed in the paper and does not by itself bear on the internal validity of a triple-blind protocol, but it is part of the full context.
- What survives. What the design does establish, regardless of the effect, is that the result is not an artifact of cold reading or sitter cueing, because those channels were closed. The contribution is methodological as much as empirical: a demonstration of how to test a contested claim with preregistered, triple-blind rigor.
Sources
- Tressoldi, P. E., Liberale, L., & Sinesio, F. (2024). Stage 2 Registered Report: Testing mediumship accuracy with a triple-blind protocol. https://doi.org/10.31234/osf.io/a756f R001 [Tressoldi et al. 2024] ↩︎
- Tressoldi, P. E., Sinesio, F., Liberale, L., & Testoni, I. (2021). Stage 1 Registered Report: Testing mediumship accuracy with a triple-blind protocol. F1000Research, 10, 351. https://doi.org/10.12688/f1000research.52208.2 R002 [Tressoldi et al. 2021] ↩︎
- Beischel, J. (2007). Contemporary methods used in laboratory-based mediumship research. Journal of Parapsychology, 71, 37–68. https://windbridge.org/papers/BeischelJP71Methods.pdf R003 [Beischel 2007] ↩︎
- Sarraf, M., Woodley of Menie, M. A., & Tressoldi, P. (2020). Anomalous information reception by mediums: A meta-analysis of the scientific evidence. Explore. https://doi.org/10.1016/j.explore.2020.04.002 R004 [Sarraf et al. 2020] ↩︎
- Tressoldi, P., Liberale, L., & Sinesio, F. (2022). Is there someone in the hereafter? Mediumship accuracy of 100 readings obtained with a triple level of blinding protocol. OMEGA – Journal of Death and Dying. https://doi.org/10.1177/00302228221146376 R005 [Tressoldi et al. 2022] ↩︎
- Roe, C. A., & Roxburgh, E. (2013). An overview of cold reading strategies. In C. Moreman (Ed.), The Spiritualist Movement (Vol. 2, pp. 177–203). Praeger. R006 [Roe & Roxburgh 2013] ↩︎