Walach et al. (2021)

Nailing Jelly: The Replication Problem Seems to Be Unsurmountable – Two Failed Replications of the Matrix Experiment

Walach, H., Kirmse, K. A., Sedlmeier, P., Vogt, H., Hinterberger, T., & von Lucadou, W. (2021). Nailing Jelly: The Replication Problem Seems to Be Unsurmountable – Two Failed Replications of the Matrix Experiment. Journal of Scientific Exploration, 35(4), 788–828. https://doi.org/10.31275/20212031

AI Assessment

A rare and admirable piece of scientific self-correction: the same group that had reported a positive psychokinesis result (the “matrix experiment”) preregistered a strict protocol, ran two independent replications, and reports that both failed. Neither replication produced more significant participant-to-generator correlations than chance at the predefined counting level; one showed a very small non-significant effect, the other nothing at all, and the control analyses were also null. The authors conclude plainly that the matrix experiment is not a replicable paradigm. Methodologically this is close to exemplary, preregistered, adequately powered, with a clean randomization test and control matrices, and publishing the failure of one’s own prior positive finding is exactly the practice critics say the field lacks. The one point a reader should weigh is the interpretive move: the authors read the failure as confirming their theory that psi effects are inherently elusive, a framing that is hard to falsify. This audit reports what the study found and how; it takes no position on whether psychokinesis is real.

Provenance

DOI. 10.31275/20212031 · Journal of Scientific Exploration 2021, 35(4), 788 to 828. Open access under CC BY-NC.

Study type. Two preregistered, comparatively strict (direct) replications of the authors’ own earlier positive study (Walach et al., 2020).

Authors. Harald Walach, Karolina A. Kirmse, Peter Sedlmeier, Hans Vogt, Thilo Hinterberger, and Walter von Lucadou, a group associated with the pragmatic-information (Model of Pragmatic Information) tradition in psi research.

Data availability. A consensus protocol was deposited before commencement on the Open Science Framework (osf.io/cx2tf).

Source basis. Every figure below is taken from the article’s own Abstract and Results.

What the paper reports

The matrix experiment is a psychokinesis setup in which a random event generator (REG) drives a display that participants try to influence; unusually, instead of testing the REG’s deviation from randomness, it tests a large array of 2,025 correlations between the participant’s behaviour and the REG’s output.1 Having found a positive effect previously, the authors preregistered a consensus protocol and ran two independent replications with the same setup. The paper’s finding is that both replications failed, and its conclusion is that the paradigm does not replicate.

Neither of the two experiments was significant … Our conclusion is that the matrix experiment in and of itself is not a replicable paradigm in PSI research.

How it was run

Results, as reported

MetricResult
Overallneither experiment produced more significant correlations than chance at the predefined p ≤ .1 counting level (10,000-iteration randomization test)
Experiment 1 (n = 64, power 0.88)a very small, non-significant effect; significance appeared only at some isolated lower thresholds (p ≤ .02, .0005, .0001) in the full matrix, while the predefined p ≤ .1 level and the time-forward upper triangle were non-significant
Experiment 2 (n = 40, power 0.69)no detectable effect at any level
Control matricesalso non-significant

Values are reproduced from the article’s Abstract and Results. Both preregistered replications failed at the level fixed in advance; the scattered low-threshold “significances” in Experiment 1 are the kind of artifact expected when scanning thousands of correlations, and the authors do not treat them as a positive result.

Eleven-dimension audit

Pre-registration

A clear strength: a consensus protocol, including the predefined p ≤ .1 counting level and the analysis, was deposited on OSF before data collection. That is what allows the two experiments to count as genuine confirmatory replications rather than post-hoc reinterpretations.

Randomization

The task is driven by a random event generator and evaluated with a 10,000-iteration randomization test against control matrices, an appropriate way to establish the chance baseline for a large correlation array.

Sensory leakage

Not applicable in the ganzfeld sense; this is a machine-based psychokinesis task with no sender, and the control matrices provide the relevant check that structure does not arise by chance.

Blinding

The outcome is computed automatically from REG data, so assessor bias is minimal; the two experiments used different experimenters at different sites, which also guards against a single-experimenter artifact.

Optional stopping

Sample sizes were fixed by the preregistered protocol (64 and 40), so optional stopping is not a concern.

Outcome measure

A pre-specified count of how many of 2,025 correlations exceed chance at p ≤ .1, tested by randomization. The measure is unusual but clearly defined in advance, and the multiplicity is handled by the randomization framework rather than ignored.

Effect size

Effectively zero: Experiment 1 a very small non-significant effect, Experiment 2 nothing. Both fall short of the predefined threshold, which is the honest bottom line.

Multiple comparisons

Handled well. The whole design is multiplicity-aware, counting significant correlations against a randomization baseline, and the authors explicitly decline to read the scattered low-threshold hits in Experiment 1 as evidence, which is the correct treatment of a thousands-of-correlations scan.

Internal replication

Two independent replications at two sites, both null, is strong internal consistency for the negative result, more convincing than a single failed attempt would be.

External replication

This is itself a replication effort, and it establishes that the parent positive result does not reproduce under a strict, preregistered protocol. Whether the original was a false positive is the natural inference.

Transparency

Exemplary: preregistered, adequately powered (at least in Experiment 1), control matrices analyzed identically, and a group publishing the failure of its own prior positive finding with a candid conclusion that the paradigm does not replicate. This is the practice the field is often accused of avoiding.

The adversarial record

Sources
  1. Walach, H., Kirmse, K. A., Sedlmeier, P., Vogt, H., Hinterberger, T., & von Lucadou, W. (2021). Nailing Jelly: The Replication Problem Seems to Be Unsurmountable – Two Failed Replications of the Matrix Experiment. Journal of Scientific Exploration, 35(4), 788–828. https://doi.org/10.31275/20212031 R001 [Walach et al. 2021] ↩︎