Lieb et al. (2024)
VR Video Game-induced Psi Communication With Red and Green Ganzfeld: A Proof-of-Principle Study
Lieb, Y., Schult, B., & Wittmann, M. (2024). VR video game-induced psi communication with red and green ganzfeld: A proof-of-principle study. Journal of Anomalistics, 24, 303–322. https://doi.org/10.23793/zfa.2024.303
AI Assessment
A null confirmatory result, reported transparently. The study’s main confirmatory hypothesis registered 15 hits out of 48 attempts against a chance level of 12, with a binomial probability of p = .199, not significant, and the green-versus-red ganzfeld comparison was likewise non-significant. All figures on this page were verified verbatim against the open-access primary article; the principal limitations are low statistical power, acknowledged by the authors themselves, and strongly unequal appeal of the four target games.
Provenance
DOI. https://doi.org/10.23793/zfa.2024.303. The article is open access.
Study type. Proof-of-principle experimental parapsychology study using a sender-receiver ganzfeld paradigm with two novel features: interactive VR video games as the sender’s target material, and randomly selected red or green ganzfeld light for the receiver.
Funding. The study was funded by a grant from the Society for Psychical Research to the three authors with € 4.054, disclosed in the Acknowledgments. Ethics approval was granted by the local Ethics Committee of the Institute for Frontier Areas of Psychology and Mental Health (IGPP, Freiburg, Germany; IGPP_2021_07).
Data availability. No data or code availability statement appears in the article.
Source basis. Every figure on this page was confirmed against the primary article as published in the Journal of Anomalistics, Volume 24 (2024), pp. 303–322.
What the paper reports
Lieb, Schult, and Wittmann recruited N = 48 young couples in a romantic relationship as sender-receiver pairs, yielding 48 trials.1 In each trial the sender played one of four interactive VR video games, randomly selected, while the receiver lay in a shielded cabin wearing goggles producing a randomly selected red or green ganzfeld. Regarding the main confirmatory hypothesis (H1a), the experiment registered 15 hits out of 48 attempts (31.25%), where the chance level lies at 12 (25%); according to a binomial test the probability of exactly, or more than, 15 hits (K) out of 48 trials (n) is p = .199 (z = .83), corresponding to an effect size of .12. The confirmatory result was null.
The second confirmatory test (H1b), in which two independent raters matched the receivers’ written ganzfeld descriptions to the games, was also null: both raters assigned the correct target 10 times out of 48 (p = .795). Among the exploratory hypotheses, receivers’ hit rates in the green as compared to the red ganzfeld were not significantly different (χ² = .814; p = .367), and the assessed experiential state variables for the video game and ganzfeld sessions, as well as the measured trait variable absorption, did not affect the hit rate. A post-hoc analysis found that, independent of the hit rate, the four games were identified as targets a strongly unequal number of times.
Since this is the very first study conducting such a comparison, the achieved null effect of color difference has to be treated with caution, especially since there was no significant overall effect of target hits.
How it was run
- Design. Sender-receiver ganzfeld paradigm with a four-choice judging task (25% chance of a hit). Only one trial was run per couple (n = 48 couples, n = 48 trials); the authors deliberately did not switch sender and receiver roles, in order not to influence the receiver’s impressions with knowledge about the four video games.
- Procedure. Both partners first completed the Tellegen Absorption Scale and a 10-minute progressive muscle relaxation, then were separated. The VR game and the ganzfeld session started at the same time, 10 minutes after separation, and lasted exactly 25 minutes; the two experimenters synchronized their watches and followed a planned, precise schedule.
- Targets. Four commercially available VR video games (Beat Saber, Superhot, Eagle Flight, and Explore Fushimi Inari), selected to contrast on flow potential, arousal level, agency and control, color theme, mode of action control, and type of challenge; the selection was supervised by video game scholar Federico Alvarez Igarzábal.
- Apparatus. Sender: Oculus Rift S head mounted display (1280×720 pixels) on a Windows 10 PC (Intel i7-9700K; 16 GB RAM; GeForce RTX 2080Ti). Receiver: Kasina DeepVision ganzfeld goggles, RGB values red (255-0-0) and green (0-255-0), luminance standardized to ca. 105 cd/m² to 115 cd/m², with brown noise over Sennheiser HD 201 headphones.
- Randomization. Game selection, ganzfeld color, and partner-to-experimenter assignment were all randomized through www.random.org; the order of the four 2-minute judging clips was predetermined via a Latin square design.
- Sample and recruitment. Couples in a romantic relationship, ages 18 to 39, recruited by advertisement on a local university website and by word of mouth, screened by telephone for exclusion criteria (any psychiatric or neurological disorder), and compensated €30 for a session lasting approximately 90 minutes.
- Controls. Sender and receiver were in separate, non-adjacent rooms on the same floor; the receiver lay in an electromagnetically shielded EEG cabin; each participant-experimenter pair was separated from the other until the end of the session so that no unwanted information transfer could happen; the played game was revealed only after the receiver’s ranking and questionnaires.
- Judging. The receiver ranked the four game clips from 1 to 4 (rank 1 used for confirmatory analysis); separately, two external raters who had played all four games independently read each of the 48 experience descriptions and ranked the clips, unaware of the receivers’ choices.
Results, as reported
| Metric | Result |
|---|---|
| Confirmatory hit rate (H1a) | 15/48 (31.25%) vs 12 expected by chance (25%) |
| Binomial test, H1a | p = .199 (z = .83) |
| Effect size, H1a | ES = .12 (z divided through the square root of n = 48) |
| Confirmatory rater matching (H1b) | 10/48 correct by each of two raters vs 12 expected by chance |
| Binomial test, H1b | p = .795 |
| Inter-rater agreement (Spearman) | rs = .326, p = .024 |
| Green ganzfeld hits (separate binomial) | K = 7 hits in 27 trials, p = .529 |
| Red ganzfeld hits (separate binomial) | K = 8 hits in 21 trials, p = .130 |
| Green vs red hit rate (exploratory H2) | χ² = .814 (df = 1), p = .367 |
| Absorption regression, sender (exploratory H3) | β = −.005, [−.002; .013], p = .161 |
| Absorption regression, receiver (exploratory H3) | β = −.002, [−.008; .005], p = .506 |
| Distribution of games played (post-hoc one-way test) | χ² = 1.83, df = 3, p = .608 |
| Hits vs misses across the four games (post-hoc; test assumptions violated per the authors) | χ² = 14.7, df = 3, p = .002 |
| Eagle Flight signal detection (post-hoc) | hit rate .8, false alarm rate .632, d′ = .504 |
95% confidence intervals are reported only for the bootstrapped regression coefficients (H3); no confidence intervals are given for the hit rates, the binomial tests, or the correlations. The Spearman correlations between the hit rate and the state variables of senders and receivers (exploratory H4 and H5: SAM valence, SAM arousal, passage of time, engagement) are all reported as non-significant, with p values ranging from .517 to .913.
Eleven-dimension audit
Pre-registration
The article does not report a pre-registration. The paper itself distinguishes confirmatory from exploratory tests: H1a and H1b are labelled the confirmatory psi hypotheses, while H2 through H5 are explicitly introduced as “several exploratory alternative hypotheses.” The authors themselves, citing prior failed replication attempts, call for “more strictly controlled, preregistered multi-lab studies,” which indicates they did not present this proof-of-principle study as meeting that standard.
Randomization
Randomization was external and documented: the ganzfeld color was assigned before the session by a random number generator, the game was selected by the experimenter in the sender’s room via smartphone access to the same service, and partner-to-experimenter assignment was randomized before arrival, with all randomization procedures done through www.random.org. The order of the four judging clips was predetermined via a Latin square design. The realized distributions (27 green vs 21 red; unequal game counts) reflect simple random assignment rather than balancing.
Sensory leakage
The protocol addressed leakage in several ways: sender and receiver were placed in separate, non-adjacent rooms on the same floor; the receiver was inside an electromagnetically shielded EEG cabin wearing ganzfeld goggles and brown-noise headphones; each participant-experimenter pair was separated from the other until the end of the session; and the sender’s lab door was marked to prevent interruptions while the receiver’s experimenter occupied the anteroom. The played game was disclosed only after the receiver had ranked the clips and completed the questionnaires.
Blinding
The game was generated in the sender’s room after the pairs had been separated, so the experimenter and participant on the receiver side had no stated channel to the target identity before judging. The two external raters for H1b read the 48 descriptions independently of one another and were unaware of the receivers’ choices. The experimenters were not blind to condition within their own room, which is inherent to the design; the paper does not report any additional masking beyond the separation procedure.
Optional stopping
The sample was a fixed n = 48 couples with exactly one trial per couple, a design decision the authors justify on methodological grounds (avoiding contamination of receiver impressions in a second session). No interim analyses are reported. The authors transparently state that with 48 trials and a 25% chance rate they would have had to achieve at least n = 18 hits (37.5%) for significance in the binomial test.
Outcome measure
The confirmatory outcome was clearly defined in advance of analysis: the receiver’s rank-1 choice among four video clips, scored as a hit or miss against a 25% chance baseline, with K > 12 hits out of 48 required to accept H1. The parallel rater-based outcome (H1b) used the same probabilities. Ranks 2 through 4 were also collected, and the post-hoc ranking analysis (Table 3 of the paper) drew on them, but the primary endpoint itself is simple and unambiguous.
Effect size
The paper reports an effect size of .12 for the confirmatory test, computed as the z-score (.83) divided through the square root of n = 48 trials. The authors candidly frame the power problem: their 31.25% hit rate “conforms to” the overall hit rates of approximately 30% (Storm et al., 2010) and 27% (Storm & Tressoldi, 2020) reported in meta-analyses of four-choice designs, rates that “become only significant in meta-analyses, given the large number of studies included.”
Multiple comparisons
Beyond the two confirmatory tests, the paper reports exploratory tests H2 through H5 (color comparison, absorption regression, and eight state-variable correlations) plus post-hoc analyses of game distribution, hits versus misses per game, and signal detection for Eagle Flight, with no correction for multiplicity mentioned. The authors themselves flag the most striking post-hoc value (χ² = 14.7, p = .002 for hits versus misses across games) as unreliable because low cell counts “violate the assumption of the χ² test,” and they interpret it as a stimulus-appeal artifact rather than evidence of psi.
Internal replication
There is no internal replication: each couple contributed a single trial, and the design deliberately excluded a second session with switched roles. The closest internal cross-check is the independent rater analysis (H1b), which used the same 48 descriptions and was also null (10 of 48 correct, p = .795), with significant inter-rater agreement (rs = .326, p = .024) indicating the raters at least converged on similar, if incorrect, matchings.
External replication
The authors state that to the best of their knowledge this is the first study in which the sender of a ganzfeld psi experiment played VR video games, and the first to compare red and green ganzfeld in this context; no external replication of this specific design exists. The broader four-choice ganzfeld literature against which they situate the result, including the meta-analysis of free-response studies 1992–2008 by Storm, Tressoldi, and Di Risio (2010),2 reports aggregate hit rates near this study’s 31.25%, but the present design awaits its own replication, which the authors explicitly invite.
Transparency
Transparency is generally good: the article is open access, funding (Society for Psychical Research, € 4.054) and ethics approval (IGPP_2021_07) are disclosed, null results are reported without spin, and the authors volunteer the power calculation and the stimulus-appeal problem. Two weaknesses: there is no data availability statement, and the article contains an internal inconsistency in the game counts, with the text reporting the random game distribution as Beat Saber n = 17, Superhot n = 11, Eagle Flight n = 9, Fushimi Inari n = 11, while Table 2 and the post-hoc analyses give 16, 11, 10, and 11 respectively.
The adversarial record
- Precursor work. The design builds directly on the group’s own prior studies: Kübel, Fiedler, and Wittmann (2021) on differential experience of red versus green ganzfeld light, Rutrecht et al. (2021) on flow states in VR video gaming, and Müller and Wittmann (2017), an earlier proof-of-principle remote viewing study in the same journal. The standard sender-receiver ganzfeld paradigm itself descends from the literature the paper cites, including Bem and Honorton (1994).
- Contested-literature context. The paper situates itself within a long-disputed evidence base, citing the Hyman and Honorton (1986) joint communiqué on the psi ganzfeld controversy, and itself acknowledges that “attempts to replicate specific findings have often failed,” citing Kekecs et al. (2023) and Maier et al. (2020), and pointing to Walleczek and von Stillfried (2019) on the need for advanced meta-experimental control. The meta-analytic baselines it invokes (Storm et al., 2010; Storm & Tressoldi, 2020) are themselves products of this contested literature.
- Reading against the study. The confirmatory result was null, and by the authors’ own account the study was underpowered for the effect sizes the meta-analyses suggest, so the color, state, and trait null findings are uninterpretable as evidence about moderators. More structurally, the post-hoc analysis shows the four-choice assumption was compromised: Eagle Flight was chosen as the target 32 times out of 48 trials despite being played only 10 times, yielding a false alarm rate of .632 and a d′ of .504, which the authors attribute to unequal stimulus appeal rather than psi. Future versions of the design would need empirically appeal-matched targets, as the authors concede.
- Source-fidelity note. The article’s text and Table 2 disagree on how often each game was played (17, 11, 9, 11 versus 16, 11, 10, 11); the discrepancy does not touch the confirmatory statistic (15 hits in 48 trials) but should be resolved in any replication or reanalysis.
Contested record
Sources
- Lieb, Y., Schult, B., & Wittmann, M. (2024). VR video game-induced psi communication with red and green ganzfeld: A proof-of-principle study. Journal of Anomalistics, 24, 303–322. https://doi.org/10.23793/zfa.2024.303 R001 [Lieb 2024] ↩︎
- Storm, L., Tressoldi, P. E., & Di Risio, L. (2010). Meta-analysis of free-response studies, 1992–2008: Assessing the noise reduction model in parapsychology. Psychological Bulletin, 136, 471–485. https://doi.org/10.1037/a0019457 R002 [Storm 2010] ↩︎