Radin et al. (2021)
Psychophysical interactions with a double-slit interference pattern: Exploratory evidence of a causal influence
Radin, D., Wahbeh, H., Michel, L., & Delorme, A. (2021). Psychophysical interactions with a double-slit interference pattern: Exploratory evidence of a causal influence. Physics Essays, 34(1), 79–88. https://doi.org/10.4006/0836-1398-34.1.79
AI Assessment
The long-delayed full report of a 2012 to 2013 double-slit experiment (25 people, 250 sessions, no real-time feedback), notable because its pre-planned analysis found nothing. This is the unusually candid one: the paper’s title says “exploratory,” and its abstract leads with the fact that the planned directional analysis showed no psychophysical effect. Every significant result reported here comes from two metrics the authors developed after the planned analysis came up null (a simplified spectral measure and a fringe-visibility measure), which then produced effects while the matched sham sessions stayed null. The authors are explicit that these are post-hoc and require independent replication, and the study is also the one a funder-affiliated group publicly called a false positive. The transparency is a real strength; the post-hoc status of the positive results is the real limitation. This audit reports what the paper planned, ran, and found; it takes no position on whether consciousness influences quantum systems.
Provenance
DOI. 10.4006/0836-1398-34.1.79 · Physics Essays 2021, 34(1), 79–88. Received 21 November 2020, accepted 8 February 2021. Physics Essays is a small, specialist venue rather than a mainstream physics journal.
Study type. A controlled laboratory experiment with a pre-planned (encrypted) primary analysis that was null, plus three exploratory analyses developed afterward that carry the positive findings.
Authors. Dean Radin, Helané Wahbeh, and Leena Michel (Institute of Noetic Sciences) and Arnaud Delorme (University of California, San Diego).
Notable history. The experiment was funded by J. Walleczek, who held the decryption key; the attention assignments were encrypted so data could not be analyzed until all sessions were complete (decrypted March 2013 in the funder’s presence). Public reporting was withheld eight years at the funder’s request, and a funder-affiliated group (Walleczek and von Stillfried, 2019) published a critique calling a control condition a false positive, to which this paper responds.
Funding. The Federico and Elvia Faggin Foundation, the Bial Foundation, Richard and Connie Adams, and an anonymous institute (stated in the Acknowledgments).
Source basis. Every figure below is taken from the article’s own Abstract, Methods, Results, and Discussion.
What the paper reports
Participants directed attention toward (X) or away (O) from a sealed double-slit system, with no real-time feedback, to test whether focused attention shifts the interference pattern.1 The design used four epoch-pair conditions (OO, XX, OX, XO) borrowed from biological causal-inference designs, with sham (no-observer) sessions run immediately before and after each experimental session as matched controls.
Based on the planned analysis, no evidence for a psychophysical effect was found… Future studies using the same protocols and analytical methods will be required to determine if these exploratory results are idiosyncratic or reflect a genuine psychophysical influence.
How it was run
- Apparatus. A 5 mW HeNe laser through a 10 µm double slit separated by 200 µm, recorded by a 3000-pixel CCD line camera 14.0 cm from the slits, in a matte-black sealed housing inside a double steel-walled electromagnetically shielded chamber.
- Participants. 25 people selected for prior performance or an active attention discipline, each contributing ten sessions (two per appointment) plus matched sham sessions, seated about 2 m from the apparatus and asked not to approach it.
- Design. Four epoch-pair types (OO, XX, OX, XO), 30 s epochs, with no real-time feedback (removed because earlier participants found the audio tone distracting). Attention assignments were encrypted in real time so the data could not be analyzed until the experiment ended.
- Controls. Sham sessions used the same attention sequence with no observer present, with a 60 W bulb placed in the participant’s chair to mimic body heat, and four thermocouples logged temperature throughout.
- Analyses. The planned analysis compared the four epoch-pair types with a double-detrended spectral metric (Stouffer Z across sessions, FDR correction). Three exploratory analyses were added later: a simplified spectral metric, a fringe-visibility metric, and a talent-persistence test.
Results, as reported
| Metric | Result |
|---|---|
| Sample | 25 participants; 250 experimental (249 plus one redo) and 250 sham sessions, February 2012 to February 2013 |
| Planned analysis (double-detrended spectral, 8 comparisons) | NULL after FDR correction; Table I experimental Stouffer Z: OO = 0.75 (neg), XX = 0.72, OX = 1.28 (neg), XO = 0.66 (neg) |
| Exploratory spectral, H2 (all X vs O) | experimental z = 2.829 (negative), p = 0.002; sham z = 0.580 (negative), p = 0.281 |
| Exploratory spectral, H3 (12-element causal model) | experimental Spearman rho = 0.873 (p = 0.0001 in text, p = 0.0003 in Fig. 4); sham rho = 0.106, p = 0.37 |
| Exploratory spectral, H4 (talent) | top-5 last-5 sessions z = 3.22, p = 0.0006; worst-5 z = 0.35 (non-significant) |
| Exploratory fringe visibility, H3 | 20 of 22 fringes in predicted direction; 7 survive FDR at p = 0.01; best fringe about 4.3 sigma, p = 0.00001 |
| Fringe-visibility talent (H4) | not confirmed (top-5 and worst-5 both non-significant) |
| Artifact checks | temperature X vs O not different between experimental and sham (z = 0.229); variance experimental > sham (z = 4.058, p = 0.00002) but X vs O within not different (z = 0.424) |
Values are reproduced from the article’s Abstract, Results, and Table I. The structure is unusually clear: the pre-committed analysis is null, and every positive number comes from a metric introduced afterward, with the matched sham sessions null throughout.
Eleven-dimension audit
Pre-registration
Stronger than most in this line, then weaker. The primary hypothesis and metric were fixed in advance and the attention assignments were encrypted so the data literally could not be analyzed early, which is a real safeguard. But that pre-committed analysis returned nothing, and the reported positive effects come from three exploratory analyses developed after the fact. The authors name this as the study’s principal limitation.
Randomization
Attention sequences were pseudorandomly generated per session and the same sequence was reused for the matched sham. Participants were selected for talent rather than randomly sampled. The statistical nulls were built by nonparametric randomized circular-shift resampling.
Sensory leakage
Well controlled physically: sealed housing, shielded chamber, participant 2 m away and instructed not to approach, headphones and steel walls preventing the assistant from overhearing conditions, and gross movement visible as data artifacts (none observed). The honest weakness the authors themselves flag is that the sham environment (a 60 W bulb, no person) did not perfectly match the experimental environment.
Blinding
Real-time encryption of the conditions is an effective analytic blind, and the same assistant ran all sessions to keep interactions uniform. The measure is instrumental, and the matched sham provides the no-observer comparison.
Optional stopping
Limited scope for it: the design fixed 25 participants at ten sessions each (250 planned), and encryption prevented interim analysis. One experimental session that failed to record was simply repeated to complete the planned count.
Outcome measure
This is the crux. The planned outcome (a double-detrended spectral metric) was null. Two new outcome measures (a simplified spectral mean over wavenumbers 25 to 50, and fringe visibility across 22 fringes) were then constructed, motivated by a stated concern that the planned metric was not directly tied to the interference pattern, and those measures produced the positive results. Changing the outcome measure after a null primary is the central interpretive caution here.
Effect size
The exploratory effects are moderate (the simplified spectral X-vs-O at about 2.8 sigma; a single best fringe at about 4.3 sigma), and the 12-element model correlation is high (rho = 0.873) but is a within-design pattern-match rather than a simple mean shift. None of these is a pre-committed effect.
Multiple comparisons
FDR correction is applied within the fringe array and to the planned comparisons, which is appropriate. The deeper multiplicity is across analyses: several exploratory metrics and hypotheses were examined, and the reported effects are the subset that reached significance, which FDR within a single analysis does not address.
Internal replication
Mixed within the paper. The two exploratory metrics (spectral and fringe visibility) broadly agree on direction, which is some internal corroboration, but the talent hypothesis (H4) was confirmed on the spectral metric and not on the fringe-visibility metric, so the internal picture is not uniform.
External replication
None here. The authors are explicit that the exploratory results can only be validated by future replications using the same protocol, and they recommend pretesting participants for talent, redesigning the sham to include a present-but-distracted person, adding accelerometers, and using a non-mode-hopping laser.
Transparency
The strongest dimension. The paper leads with its null primary result, labels the positive findings exploratory in the title and throughout, discloses the encryption and eight-year reporting delay, reports the funder’s false-positive critique and its own rebuttal, documents a significant decline in participant mood and attention over sessions, and openly concedes the sham did not perfectly match the experimental condition. Two small internal discrepancies a reader should note: the 12-element spectral correlation p-value is given as 0.0001 in the Results text but 0.0003 in the Figure 4 caption, and the FDR threshold for fringe #2 is stated as p = 0.05 in the text but alpha = 0.01 in the Figure 5 caption.
The adversarial record
- The primary result is null. On the pre-planned, encrypted analysis, this study found no psychophysical effect. That is the result a strict reading records.
- Positive findings are post-hoc. Every significant number comes from metrics built after the planned analysis failed. However well motivated, developing new outcome measures until one is significant is the classic route to false positives, which is precisely why the authors call for replication.
- A funder-affiliated group called a control a false positive. Walleczek and von Stillfried (2019) published that critique; the authors respond that it failed to correct for eight tests (a 33.6% chance of at least one false positive), which is a fair statistical point but also underscores how much this dataset’s interpretation is contested.
- Imperfect sham and a fatiguing protocol. The authors themselves note the sham environment differed from the experimental one and that participants’ mood and attention declined significantly over the ten sessions, both of which complicate the comparison.
- What the authors do right. This is a model of candor: a genuinely pre-committed and encrypted primary analysis, a null result reported up front and in the title, explicit labeling of everything exploratory, direct engagement with the false-positive critique, disclosure of mood decline and sham mismatch, and a concrete replication roadmap. A reader is given everything needed to discount the positive results appropriately.
Sources
- Radin, D., Wahbeh, H., Michel, L., & Delorme, A. (2021). Psychophysical interactions with a double-slit interference pattern: Exploratory evidence of a causal influence. Physics Essays, 34(1), 79–88. https://doi.org/10.4006/0836-1398-34.1.79 R001 [Radin 2021 PE] ↩︎