Psi in a Skeptic’s Lab: A Successful Replication of Ertel’s Ball Selection Test

Article Analysis — AI-characterized, source-grounded

Contents

Provenance

What this document is

This is a research article in which the author, Suitbert Ertel, re-analyzes data collected by two undergraduate students — Johanna Körting and Luke Hagstrom — at the Anomalistic Psychology Research Unit (APRU), Goldsmiths College, University of London, headed by Professor Chris French, whom the paper identifies as editor of The Skeptic. The students had used Ertel’s Ball Selection Test materials for their BSc Final Year Projects (both dated 2002) after Ertel lectured at APRU and left a test kit behind. The paper frames its central question as whether the APRU students replicated results Ertel had obtained at the Georg-Elias-Müller Institut (GEMI), Göttingen University, and whether psi effects would survive a skeptical laboratory setting.

The stakes, as the document frames them, concern the “hotly debated problem” of replication in parapsychology, the possible role of experimenter effects — with skeptical attitudes said to “obliterate” psi manifestations — and the claim that this low-tech, less monotonous procedure is more psi-conducive and more replicable than conventional multiple choice psi tests.

How the study was run and re-analyzed

The test. Fifty ping pong balls in an opaque bag, each marked with an integer 1–5 and red or green dots, yielding 10 equiprobable ball types (10% chance of a “double hit” on both number and color). Participants jumble the balls, guess number and color, draw a ball, record the result, and replace the ball. A run is 60 trials taking 10–15 minutes; APRU participants completed 6 runs (360 trials), usually over 2 or 3 sessions.

Departures from the GEMI protocol. The GEMI standard has two stages — unsupervised at-home testing, then supervised retesting of significant scorers. APRU forwent home tests: all 40 unselected participants (14 male, 26 female, ages 19 to 56, mean 28.6) were tested under supervision, 20 by each student experimenter (Hagstrom, 26; Körting, 23). Körting used the standard version II balls; Hagstrom substituted “air-flow” balls (white vs. orange replacing the dot colors), citing “availability and time constraints”. Hagstrom let participants record their own results while he monitored; Körting filled out the record sheets herself. Questionnaire data (extraversion, paranormal belief, pre/post self-ratings) were collected by the students but were not available and are not part of the re-analysis; Hagstrom supplied only per-participant totals, so trial-by-trial re-analysis was possible only for Körting’s participants (and only 17 of her 20).

The re-analysis. Ertel argues the students’ one-sample t tests were unsuitable and underpowered, illustrating with hypothetical samples how a t test can rank deviations from chance perversely. His preferred analyses are a one-tailed binomial test on summed hits and a summed-Z² (Chi²) procedure attributed to Timm (1983:222), plus a kurtosis indicator of bi-directional (psi-hitting and psi-missing) spread. He also disputes Hagstrom’s Bonferroni correction as inadmissible because the double-hit score is, by instruction, “the ultimate success measure”, and identifies Körting’s two-sided p = .052 (versus Hagstrom’s p = .038 on the same data) as “apparently due to a calculation error”; the correct one-sided t-test p is given as .019.

Reported results

The paper’s Table 2 compares double-hit results (10% expected) across the APRU replication and two GEMI datasets. Key figures as printed:

VariableAPRU (unselected, supervised)GEMI I (unselected, unsupervised)GEMI II (selected, supervised)
Participants40479
Trials per participant360480480
Total trials14,40022,5604,320
Total hits1,5482,620748
Observed hits per run (6.00 expected)6.456.9710.39
Zbin / p2.99 / 0.0028.07 / 10−1416.00 / <10−50
ES10.0250.0540.243
Chi² (df = N) / p76.3 / 0.0003288.0 / <10−50279.4 / <10−50

Additional reported findings: the abstract states the 40 APRU participants “achieved a hit rate of 10.75” (percent sign not printed), significant at p = .002 (binomial) and p = .0003 (summed Z²); the APRU hit rate was significantly lower than GEMI’s (abstract: p = .02; Results section: Chi² = 6.55, df = 1, p = .01) and this difference was predicted. GEMI II’s deviation from the unselected GEMI sample is given as Chi² = 28.4, df = 1, p = 10−7. The top APRU scorer (Table 1) had 67 double hits, Z = 5.36, p = 10−6; Hagstrom notes this participant was his mother. Two GEMI high scorers, expected 48 hits in 480 trials, scored 80 and 89 at home and 230 and 143 under supervision — “(230 is no typo!)”. Correlations between number and color hits are reported as low (APRU: .40, GEMI: .30). Number hits (20% expected) and color hits (50% expected) are broken out in appendix Tables 4 and 5; for APRU neither reached significance by summed hits (Zbin 0.78 and 1.53, n.s.), though summed-Z² values were significant (p = 0.006 and p = 10−5 respectively).

What the document itself concludes

Ertel concludes that “Highly significant hit scores have been obtained” under the supervision of “an eminent British skeptic”, and that with appropriate statistics the students’ own “hypothesis one” found support even though the students — “through inexperience” — had reached negative or ambivalent conclusions (Körting: “The results obtained by Ertel (2002) were not replicated”). He attributes the lower APRU scores to supervision as a psi-detrimental condition and possibly to the skeptical social embedding. He rebuts the objection that unsupervised home scores cannot be trusted, argues that fraud speculation “disregards the reality of trust” in ordinary social interactions, criticizes die-hard skeptics for “denaturalizing the test environment”, and calls for more joint research between parapsychologists and critics, noting French’s cooperation was “a rare exception”. He suggests the test could eventually serve as a recruiting tool for psi-gifted participants, and acknowledges that attempts to find non-psi explanations should continue even though further success “is deemed unlikely”. The two GEMI participants whose scores rose sharply under supervision are described as “a surprising and as-yet-unexplained observation”.

Document-integrity notes

References cited in the document

The article carries a reference list of 21 entries; per site policy for long lists it is not reproduced here. Entries discussed in the analysis above include: Hagstrom, L. (2002) and Körting, J. (2002), the two Goldsmiths BSc theses supervised by Pr. Chris French; Masuhr, B. (2000), a GEMI Diplomarbeit; Ertel, S. (2007), Zeitschrift für Anomalistik, 7(3), 236–269; Timm, U. (1983), Zeitschrift für Parapsychologie und Grenzgebiete der Psychologie, 2, 195–229; and Watt, C. (2006), Journal of Parapsychology, 70(2), 335–356. The list also includes works on experimenter effects (Rhine & Pratt, Honorton et al., Palmer, Schlitz & LaBerge, Watt & Ramakers, Smith, White) cited in the introduction and discussion.