Psi in a Skeptic’s Lab: A Successful Replication of Ertel’s Ball Selection Test
Article Analysis — AI-characterized, source-grounded
Contents
- Provenance
- What this document is
- How the study was run and re-analyzed
- Reported results
- What the document itself concludes
- Document-integrity notes
- References cited in the document
Provenance
- Author as printed: SUITBERT ERTEL, Georg-Elias-Müller-Institut für Psychologie, Waldweg 26, 37073 Göttingen, Germany.
- Bibliographic record: Journal of Scientific Exploration, Vol. 24, No. 4, pp. 581–598, 2010 (the running head and page footer in the text confirm this range).
- The stored copy appears to be OCR text extracted from the typeset journal pages, with page numbers, running heads, tables, two appendices of instructions/record sheets, and two appendix tables included, though with characteristic OCR artifacts (see integrity notes).
- read the complete extracted text
What this document is
This is a research article in which the author, Suitbert Ertel, re-analyzes data collected by two undergraduate students — Johanna Körting and Luke Hagstrom — at the Anomalistic Psychology Research Unit (APRU), Goldsmiths College, University of London, headed by Professor Chris French, whom the paper identifies as editor of The Skeptic. The students had used Ertel’s Ball Selection Test materials for their BSc Final Year Projects (both dated 2002) after Ertel lectured at APRU and left a test kit behind. The paper frames its central question as whether the APRU students replicated results Ertel had obtained at the Georg-Elias-Müller Institut (GEMI), Göttingen University, and whether psi effects would survive a skeptical laboratory setting.
The stakes, as the document frames them, concern the “hotly debated problem” of replication in parapsychology, the possible role of experimenter effects — with skeptical attitudes said to “obliterate” psi manifestations — and the claim that this low-tech, less monotonous procedure is more psi-conducive and more replicable than conventional multiple choice psi tests.
How the study was run and re-analyzed
The test. Fifty ping pong balls in an opaque bag, each marked with an integer 1–5 and red or green dots, yielding 10 equiprobable ball types (10% chance of a “double hit” on both number and color). Participants jumble the balls, guess number and color, draw a ball, record the result, and replace the ball. A run is 60 trials taking 10–15 minutes; APRU participants completed 6 runs (360 trials), usually over 2 or 3 sessions.
Departures from the GEMI protocol. The GEMI standard has two stages — unsupervised at-home testing, then supervised retesting of significant scorers. APRU forwent home tests: all 40 unselected participants (14 male, 26 female, ages 19 to 56, mean 28.6) were tested under supervision, 20 by each student experimenter (Hagstrom, 26; Körting, 23). Körting used the standard version II balls; Hagstrom substituted “air-flow” balls (white vs. orange replacing the dot colors), citing “availability and time constraints”. Hagstrom let participants record their own results while he monitored; Körting filled out the record sheets herself. Questionnaire data (extraversion, paranormal belief, pre/post self-ratings) were collected by the students but were not available and are not part of the re-analysis; Hagstrom supplied only per-participant totals, so trial-by-trial re-analysis was possible only for Körting’s participants (and only 17 of her 20).
The re-analysis. Ertel argues the students’ one-sample t tests were unsuitable and underpowered, illustrating with hypothetical samples how a t test can rank deviations from chance perversely. His preferred analyses are a one-tailed binomial test on summed hits and a summed-Z² (Chi²) procedure attributed to Timm (1983:222), plus a kurtosis indicator of bi-directional (psi-hitting and psi-missing) spread. He also disputes Hagstrom’s Bonferroni correction as inadmissible because the double-hit score is, by instruction, “the ultimate success measure”, and identifies Körting’s two-sided p = .052 (versus Hagstrom’s p = .038 on the same data) as “apparently due to a calculation error”; the correct one-sided t-test p is given as .019.
Reported results
The paper’s Table 2 compares double-hit results (10% expected) across the APRU replication and two GEMI datasets. Key figures as printed:
| Variable | APRU (unselected, supervised) | GEMI I (unselected, unsupervised) | GEMI II (selected, supervised) |
|---|---|---|---|
| Participants | 40 | 47 | 9 |
| Trials per participant | 360 | 480 | 480 |
| Total trials | 14,400 | 22,560 | 4,320 |
| Total hits | 1,548 | 2,620 | 748 |
| Observed hits per run (6.00 expected) | 6.45 | 6.97 | 10.39 |
| Zbin / p | 2.99 / 0.002 | 8.07 / 10−14 | 16.00 / <10−50 |
| ES1 | 0.025 | 0.054 | 0.243 |
| Chi² (df = N) / p | 76.3 / 0.0003 | 288.0 / <10−50 | 279.4 / <10−50 |
Additional reported findings: the abstract states the 40 APRU participants “achieved a hit rate of 10.75” (percent sign not printed), significant at p = .002 (binomial) and p = .0003 (summed Z²); the APRU hit rate was significantly lower than GEMI’s (abstract: p = .02; Results section: Chi² = 6.55, df = 1, p = .01) and this difference was predicted. GEMI II’s deviation from the unselected GEMI sample is given as Chi² = 28.4, df = 1, p = 10−7. The top APRU scorer (Table 1) had 67 double hits, Z = 5.36, p = 10−6; Hagstrom notes this participant was his mother. Two GEMI high scorers, expected 48 hits in 480 trials, scored 80 and 89 at home and 230 and 143 under supervision — “(230 is no typo!)”. Correlations between number and color hits are reported as low (APRU: .40, GEMI: .30). Number hits (20% expected) and color hits (50% expected) are broken out in appendix Tables 4 and 5; for APRU neither reached significance by summed hits (Zbin 0.78 and 1.53, n.s.), though summed-Z² values were significant (p = 0.006 and p = 10−5 respectively).
What the document itself concludes
Ertel concludes that “Highly significant hit scores have been obtained” under the supervision of “an eminent British skeptic”, and that with appropriate statistics the students’ own “hypothesis one” found support even though the students — “through inexperience” — had reached negative or ambivalent conclusions (Körting: “The results obtained by Ertel (2002) were not replicated”). He attributes the lower APRU scores to supervision as a psi-detrimental condition and possibly to the skeptical social embedding. He rebuts the objection that unsupervised home scores cannot be trusted, argues that fraud speculation “disregards the reality of trust” in ordinary social interactions, criticizes die-hard skeptics for “denaturalizing the test environment”, and calls for more joint research between parapsychologists and critics, noting French’s cooperation was “a rare exception”. He suggests the test could eventually serve as a recruiting tool for psi-gifted participants, and acknowledges that attempts to find non-psi explanations should continue even though further success “is deemed unlikely”. The two GEMI participants whose scores rose sharply under supervision are described as “a surprising and as-yet-unexplained observation”.
Document-integrity notes
- OCR damage is pervasive but mostly cosmetic: ligatures are split throughout (“signi fi cant”, “fi rst”, “in fl uence”), and superscripts/subscripts have collapsed into the line (e.g., “10−14”, “Z bin”, “Chi2”).
- The abstract prints the APRU hit rate as “10.75” with no percent sign, unlike the other rates (“11.6%”, “17.3%”).
- Two apparent internal inconsistencies: the Results text gives GEMI II vs. GEMI I averages as “8.51 vs. 6.97 hits per run” while Table 2 lists the GEMI II observed rate as 10.39; and Table 3 gives a GEMI Znc Chi² of 288.6 where Table 2 prints 288.0.
- Citation mismatches: quoted passages cite “Ertel (2000)” and “Ertel (2002)”, neither of which appears in the reference list; the entries for Ertel (2004) and Ertel (2005c) both give Journal of Consciousness Studies, 12, 61–80. A Hagstrom quotation about “participant no. 6” is bracketed by the author as “[see participant rank #1]”.
- Appendix B (the blank record sheet) did not survive extraction as a usable form — it renders as repeated fragmentary rows — and the effect-size formulas in the table legends are garbled by the OCR.
References cited in the document
The article carries a reference list of 21 entries; per site policy for long lists it is not reproduced here. Entries discussed in the analysis above include: Hagstrom, L. (2002) and Körting, J. (2002), the two Goldsmiths BSc theses supervised by Pr. Chris French; Masuhr, B. (2000), a GEMI Diplomarbeit; Ertel, S. (2007), Zeitschrift für Anomalistik, 7(3), 236–269; Timm, U. (1983), Zeitschrift für Parapsychologie und Grenzgebiete der Psychologie, 2, 195–229; and Watt, C. (2006), Journal of Parapsychology, 70(2), 335–356. The list also includes works on experimenter effects (Rhine & Pratt, Honorton et al., Palmer, Schlitz & LaBerge, Watt & Ramakers, Smith, White) cited in the introduction and discussion.