Watt et al. (2020)
Testing Precognition and Alterations of Consciousness with Selected Participants in the Ganzfeld
Watt, C., Dawson, E., Tullo, A., Pooley, A., & Rice, H. (2020). Testing precognition and alterations of consciousness with selected participants in the ganzfeld. Journal of Parapsychology, 84(1), 21–37. http://doi.org/10.30891/jopar.2020.01.05
AI Assessment
A pre-registered precognitive ganzfeld study reporting above-chance hitting on its planned primary test, with null results on both secondary altered-state hypotheses. The paper reports 22 direct hits in 60 trials (37% hit-rate against a 25% chance baseline, exact binomial p = .03, 1-tailed), while finding no relation between altered-state measures and psi task performance. Every figure on this page was verified against the open-access primary article; the authors disclose their limitations, including sub-optimal power and missing ratings on four trials, candidly within the paper itself.
Provenance
DOI. http://doi.org/10.30891/jopar.2020.01.05. Open access, distributed under the terms of the Creative Commons Attribution License.
Study type. Pre-registered experimental study: an automated precognitive ganzfeld ESP experiment with selected participants, described by the authors as the first study to contribute to a registration-based prospective meta-analysis of ganzfeld ESP studies.
Funding. No funding statement appears in the article. The acknowledgments thank Laurène Vuillaume and Paul Nothdurft for assistance in creating the target pool.
Data availability. Session data were time-stamped and uploaded in duplicate, locally and to a remote University of Edinburgh server, with the programmer holding sole server access; the article contains no public data-availability statement.
Source basis. All figures on this page were confirmed against the primary article as published in the Journal of Parapsychology, 2020, Vol. 84, No. 1, pp. 21–37.
What the paper reports
Watt, Dawson, Tullo, Pooley, and Rice report a 60-trial precognitive ganzfeld experiment with participants selected for self-reported creativity, prior psi experience or belief, or practice of a mental discipline.1 On the planned test of the primary precognition hypothesis (H1), participants obtained 22 direct hits out of 60 trials, a 37% hit-rate where mean chance expectation is 25%, statistically significant on the pre-specified exact binomial test (Z = 1.94, p = 0.03, 1-tailed, ES (Z/√N) = 0.25). The count of 22 hits was independently verified by the programmer against duplicate session data held on a remote server.
The secondary altered-state hypotheses were null. For H2a, Spearman correlations between session z-scores and the 12 major Phenomenology of Consciousness Inventory (PCI) dimensions were “either rather small or near-zero” and none was significant (df = 54; no correction for multiple analyses was applied). For H2b, the predicted relation between time under-estimation and psi performance was not found (rho = 0.075, df = 54, p = 0.582). The authors note that both hypotheses were framed as exploratory:
Both hypotheses were exploratory for three reasons: 1. It is the first study run using a new automated ganzfeld testing program at the KPU; 2. There are few previous ganzfeld studies using a precognition design. 3. Limited resources mean that the study has a sub-optimal power for a confirmatory study.
How it was run
- Design. Automated precognition ganzfeld: the target clip was randomly selected by the program only after the participant’s judging had been completed, recorded, and uploaded, and the target identity was uploaded to the remote server before being revealed to participant or experimenter.
- Procedure. An audio recording delivered approximately 9 minutes of a progressive relaxation exercise, approximately one minute of guidance on reporting mentation, then 25 minutes of white noise during which the participant reported thoughts, feelings, and imagery; mentation was audio recorded and noted by the experimenter.
- Targets. 50 target pools, each of four dynamic visual targets (60–90-second color film clips with audio) obtained from the Internet and grouped to be orthogonal to one another, giving a pool of 200 clips; pools and targets were sampled with replacement.
- Randomization apparatus. The TrueRNG3 USB hardware RNG (produced by ubld.it) was used for target pool selection, order of presentation of clips for judging, and target selection; it was tested by Tullo before formal data collection by simulating 1,500,000 trials, with relative frequencies deviating from MCE by no more than 0.03% (pool and clip-of-four selection) and 0.02% (selection among 200 clips).
- Sample. Sixty volunteers (28 male, 32 female, Mage = 34.2, SD = 18.13, Range 18–80 years; 34 students), one trial each, recruited primarily from the KPU volunteer panel and selected for at least one of: practice of a mental discipline, previous psi belief or experience, or creative/artistic ability (scoring at least 3 on a 5-point scale).
- Experimenters. Three female final year undergraduate psychology students each conducted 20 trials between December 2017 and March 2018; Watt designed, pre-registered, and supervised the study and had no contact with participants beyond initial thank-you emails.
- Judging. The participant rated each of four clips from 1 (no correspondence) to 100 (perfect correspondence), with no tied ratings permitted; experimenter and participant were still blind to the target identity at this point.
- Controls. Three data uploads (session start; after ratings were submitted and before target selection; after target selection but before reveal) were server time-stamped with IP recording; all formal trials were reported and none excluded, and the programmer, who alone held the server password, independently verified the hit count with no discrepancy found.
Results, as reported
| Metric | Result |
|---|---|
| Direct hits (H1, planned primary test) | 37% (22/60) vs 25% chance (MCE) |
| Exact binomial test (H1) | Z = 1.94, p = 0.03, 1-tailed |
| Effect size (H1) | ES (Z/√N) = 0.25 |
| Target ranked first (target rankings table, N = 59) | Frequency 22, rate 37% |
| Session z-scores | Range 17.14 to -15.61; M = 1.35, SD = 8.21 |
| H2a: PCI dimensions vs session z-scores (12 Spearman correlations, df = 54) | rho from -.16 to .17; none significant; no support for H2a |
| H2b: time estimation vs session z-scores (Spearman, df = 54) | rho = 0.075, p = 0.582; H2b not supported |
| Time estimates of session duration (objective duration about 35 minutes) | Range 8 to 90 minutes; M = 26.30, SD = 13.49 |
The paper reports no confidence intervals for the hit-rate or for the effect size, and the authors explicitly state that no correction for multiple analyses was applied to the twelve PCI correlations. Due to a software error, the target ranking was unavailable for one trial where a miss occurred, so the rankings table is based on N = 59.
Eleven-dimension audit
Pre-registration
The study was pre-registered on the KPU Study Registry (Watt, 2017b) with a planned number of 60 participants, one trial each, and is the first study to contribute to a prospectively registered meta-analysis of ganzfeld ESP studies (Watt, 2017a). The hypotheses, tests, and alpha levels (exact binomial, p hit = .25, alpha p of .05 or less, one-tailed for H1) were specified in advance, although the authors themselves classify both hypotheses as exploratory.
Randomization
A commercially available hardware RNG (TrueRNG3) handled target pool selection, clip presentation order, and target selection. Pre-study testing simulated 1,500,000 trials, with the relative frequency of each of 4 clips and each of 50 pools deviating from MCE by no more than 0.03%, and each of the 200 clips by no more than 0.02%. Sampling was with replacement, and randomness testing was performed by the programmer, who had no KPU affiliation.
Sensory leakage
The precognition design is itself the principal leakage control: the target did not exist at the time of mentation or judging, having been selected only after ratings were submitted and uploaded. The authors note that, compared to telepathy and clairvoyance designs, precognition protocols minimize possible leakage of target-related information that could artifactually inflate hit rates. Sessions took place in a windowless metal chamber with no adjoining rooms, about 100 meters from the reception room.
Blinding
Experimenter and participant were blind to the target identity throughout judging, since no target had yet been selected. Watt remained unaware of session outcomes while monitoring progress, made the missing-ratings handling decision while still unaware of the outcomes of the affected trials, and the programmer verified the hit count while unaware of the number of hits the experimenters had recorded.
Optional stopping
The planned sample of 60 trials was fixed at registration and completed; all sessions included were formal, no formal sessions were excluded, and all formal trials were reported. When missing ratings were discovered on four trials about three-quarters of the way through, the decision to include the computer’s hit or miss score for those trials in the H1 test (and discard them from the ASC tests) was made before the outcomes of those trials were known, rather than by adding replacement sessions.
Outcome measure
The primary outcome was the pre-specified count of direct hits against a one-in-four baseline, tested by exact binomial probability. For the secondary ASC hypotheses, session z-scores calculated from target and decoy ratings were used as a potentially more sensitive index than hit or miss; the formula and its inputs are given in the paper. The hit-rate result and the z-score-based results are kept distinct, exactly as registered.
Effect size
The paper reports ES (Z/√N) = 0.25 for the primary test, without a confidence interval. This figure sits close to the selected-participant effect size of 0.26 reported in the Storm et al. database that motivated the recruitment strategy, although with N = 60 the study has, by the authors’ own account, sub-optimal power for a confirmatory test.
Multiple comparisons
Twelve Spearman correlations were computed between session z-scores and the major PCI dimensions, and the paper states plainly that no correction for multiple analyses was applied; all were non-significant in any case. The authors go further in the discussion, suggesting that significant PCI correlations in other studies may be spurious Type I errors and recommending correction for multiple analysis in future work.
Internal replication
The design distributed the 60 trials evenly across three experimenters (20 trials each), but the paper does not report hits broken down by experimenter, so no within-study replication comparison is available. The independent verification of the 22-hit count against duplicate server data is a data-integrity check rather than a replication. A software error left target ratings incomplete on four trials and the target ranking unavailable on one, which the paper documents in full.
External replication
The paper situates its 37% hit-rate within a small precognitive ganzfeld literature in which significant positive scoring was found in all but one prior study (its Table 1 lists results from Dunne et al., 1977 through Roe et al., 2020), and within a contested wider database: Milton and Wiseman (1999) found 30 studies at chance (27% hit rate), while Storm et al. (2010) found a significant 32% hit rate in 30 later studies after discarding one outlier, with a 30% rate when all are combined.2 The ASC null results fail to replicate Cardeña and Marcusson-Clavertz (2020) and trend in the reverse direction to Bierman (1988) on time contraction.
Transparency
The article is open access, discloses the registration documents, names who did what (design, sessions, programming, independent analysis checks), tabulates the four trials with missing ratings, reports the experimenters’ own psi-belief ratings (3, 4, 5, and Watt’s 4), and reports both null secondary hypotheses without hedging. No public data-archive link is given, and no confidence intervals accompany the headline statistics, but the reporting of what was done and found is detailed and candid.
The adversarial record
- Precursor work. The study builds on a handful of precognitive ganzfeld experiments (Dunne et al., 1977; Rogo, 1977; Sargent & Harley, 1982; Wezelman et al., 1997; Roe et al., 2020), and on the Storm et al. (2010) and Baptista et al. (2015) recommendations to use selected participants, whose selected-study effect size was 0.26 versus 0.05 for unselected participants.
- Contested literature. The ganzfeld database has been disputed since the competing Honorton (1985) and Hyman (1985) meta-analyses; Milton and Wiseman (1999) concluded that 30 studies from 1987 to February 1997 were at chance (27% hit rate), and Bierman, Spottiswoode, and Bijl (2016) argued from simulations that the database hit rate is probably inflated by questionable research practices, though it remains statistically significant when QRPs are accounted for.
- Reading against the study. The authors themselves label both hypotheses exploratory and the study under-powered for confirmation; the primary p-value (.03, 1-tailed) is modest, no confidence intervals are reported, a software error corrupted ratings on four trials, and the two pre-registered ASC hypotheses, which tested the mechanism the ganzfeld is assumed to engage, both returned null results.
- Mechanism left open. The failure to find any PCI or time-estimation relation to scoring, together with inconsistent ASC findings across Cardeña and Marcusson-Clavertz (2020) and the three Roe et al. (2020) experiments, means the study’s positive primary result does not by itself support the altered-state rationale on which the ganzfeld method rests.
Contested record
Sources
- Watt, C., Dawson, E., Tullo, A., Pooley, A., & Rice, H. (2020). Testing precognition and alterations of consciousness with selected participants in the ganzfeld. Journal of Parapsychology, 84(1), 21–37. http://doi.org/10.30891/jopar.2020.01.05 R001 [Watt 2020] ↩︎
- Storm, L., Tressoldi, P. E., & DiRisio, L. (2010). Meta-analysis of free-response studies, 1992–2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136, 471–485. https://doi.org/10.1037/a0019457 R002 [Storm 2010] ↩︎