Harris & Rosenthal (2024)
Parapsychology
Harris, M. J., & Rosenthal, R. (2024). Parapsychology. Journal of Anomalous Experience and Cognition, 4(1), 18–33. https://doi.org/10.31156/jaex.26222
AI Assessment
The meta-analysis of ganzfeld ESP research written for the U.S. National Research Council by Monica Harris and Robert Rosenthal (a founder of modern meta-analysis), prepared around 1988 and republished in 2024. Across the 28 direct-hit studies in Honorton’s 1985 database it found a combined Stouffer z of 6.60 (mean Cohen’s h of .28, a 38% hit rate against 25% expected), a file-drawer tolerance of 423 null studies, and, most consequentially, no statistical relationship between Hyman’s six methodological-flaw variables and study outcomes. The authors’ own conservative estimate, after allowing for statistical errors and the file drawer, was an accuracy of about one in three (33%). Its lasting value is the flaw analysis and the episode around it: asked to evaluate psi, the authors reported that ganzfeld research followed the most rigorous protocols of the areas they reviewed. This audit reports what the paper argues and how it was done; it takes no position on whether ESP is real.
Provenance
DOI. 10.31156/jaex.26222 · Journal of Anomalous Experience and Cognition 2024, 4(1), 18–33. Open access under CC-BY.
Study type and vintage. A meta-analysis and methodological review. It was prepared around 1988 as a background paper for the U.S. National Research Council’s Committee on Techniques for the Enhancement of Human Performance (whose report was Druckman & Swets, 1988), and analyzes ganzfeld studies from the early-to-mid 1980s. It carries the title “Parapsychology” because that was its heading as a commissioned section. The authors are deceased; per the journal’s editorial note, they had granted the editor permission to publish this original report almost unchanged, and the editor added an abstract for indexing.
Authors. Monica J. Harris and Robert Rosenthal, then of Harvard University. Rosenthal is a principal developer of meta-analytic methods (the Stouffer combined-z procedure and the fail-safe N used here are his).
Data basis. A secondary analysis: the 28 effect sizes are taken from Honorton’s (1985) published compilation of direct-hit ganzfeld studies, not from an independent retrieval of the primary literature.
Source basis. Every figure below is taken from the article’s own Results, flaw-analysis, and new-evidence sections. This 2024 scan carries visible OCR artifacts (for example rendering “3.37 × 10⁻¹¹” awkwardly and printing some 1s as the letter l); the intended numeric values are reported here.
What the paper reports
The authors set out to evaluate whether the ganzfeld ESP database, at the time the most methodologically promising area of parapsychology, could be explained away by chance or by design flaws. They summarize the 28 direct-hit studies with a combined significance test and an effect-size distribution, then subject the studies to the standard rival hypotheses (sensory leakage, recording error, fraud, file drawer, multiple testing, randomization, statistical error, and non-independence), and finally test directly whether Hyman’s catalogued flaws predict study outcomes.1 A postscript adds ten new flaw-controlled studies from Honorton’s laboratory.
Our analysis of the effects of flaws on study outcome lends no support to the hypothesis that ganzfeld research results are a significant function of the set of flaw variables.
How it was run
- Database. The 28 direct-hit ganzfeld studies compiled by Honorton (1985), conducted by 10 investigators or laboratories.2
- Significance. The Stouffer combined-z method across the 28 studies.
- Effect size. Cohen’s h (the difference between arcsine-transformed obtained and chance hit proportions), chosen over the raw proportion difference because equal h values are equally detectable.
- File drawer. Rosenthal’s fail-safe N: the number of unretrieved null studies (mean z = 0) that would be needed to push the combined p above .05.
- Flaw analysis. Hyman’s (1985) six flaw variables (documentation, feedback, randomization, security, single target, statistical analysis), each coded adequate or not, were related to the outcomes (z and Cohen’s h) by canonical correlation and by separate multiple regressions.
- Sensitivity checks. The authors recomputed the effect after omitting the six studies with agreed statistical errors, and after omitting the nine studies from the investigator (Sargent) whose randomization was later questioned.
- New evidence. Ten later Honorton studies designed to control the earlier flaws were combined with the original set.
Results, as reported
| Metric | Result |
|---|---|
| Database | 28 direct-hit ganzfeld studies (Honorton 1985), from 10 investigators/labs |
| Combined significance | Stouffer z = 6.60, p = 3.37 × 10⁻¹¹ |
| Mean effect size | Cohen’s h = .28 (95% CI .11 to .45), equivalent to a 38% hit rate vs 25% expected |
| Direction | 82% of studies show a positive effect (p = .0004) |
| File drawer (fail-safe N) | 423 null studies would be needed to raise the combined p above .05 |
| Statistical errors | 6 of 28 studies; omitting them lowers mean h from .28 to .26 (accuracy 38% to 37%) |
| Flaws vs outcome (canonical) | adjusted canonical correlation .46, F(12,40) = 0.91: not significant |
| Flaws vs outcome (regression) | flaws vs h F(6,21) = 0.84, p = .56; vs z F(6,21) = 1.65, p = .18; of 36 t-tests, none reached p < .05 |
| Study independence | median h identical (.32) for the 28 studies and the 10 labs; 82% vs 80% positive |
| New evidence (10 Honorton studies) | combined z = 2.791, p = .0026, mean h = .23 |
| Combined (28 + 10) | z = 7.10, mean h = .27 (z = 5.74, h = .25 omitting Sargent’s 9 studies) |
| Authors’ conservative estimate | after shrinkage for errors and file drawer, h about .18, accuracy about one in three (33%) |
Values are reproduced from the article’s Results, Postscript, and flaw-analysis sections. The headline combined significance is extreme, but the authors deliberately shrink the effect-size estimate downward (from a 38% hit rate to about 33%) to account for statistical errors and probable file-drawer studies, and they emphasize that the flaws Hyman catalogued do not statistically predict which studies succeeded.
Eleven-dimension audit
Pre-registration
Not applicable to a late-1980s meta-analysis, which predates preregistration norms. There is no registered protocol; however, the flaw-analysis design was not improvised to favour a conclusion, it followed Hyman’s own 1986 recommendation to examine flaw-outcome relationships multivariately, which is the fair way to run such a test.
Randomization
Randomization enters twice: as one of Hyman’s coded flaw variables (found not to predict outcome) and as a targeted sensitivity check. The authors note that the median significance of the 16 studies using random-number tables or generators (z = .94) was essentially identical to that of all 28, and that omitting the nine Sargent studies whose randomization was questioned lowers the mean h only from .28 to .26. Poor randomization does not appear to be driving the result.
Sensory leakage
Treated as the primary rival hypothesis. The authors review the long history of unintentional cueing and note Honorton’s finding that studies controlling for handling cues yielded at least as many significant effects as those that did not. As a secondary meta-analysis, the paper cannot independently re-verify leakage control in each study, but it engages the concern directly rather than waving it away.
Blinding
Blinding is a property of the primary studies rather than of the meta-analysis; the paper does not code it as a separate variable beyond Hyman’s security flaw. This is the expected limit of a study-level meta-analysis.
Optional stopping
Addressed through the file-drawer machinery: the fail-safe N of 423, the Parapsychological Association’s norm of reporting negative results, and Blackmore’s survey (7 of 19 unretrieved studies, 37%, judged significant, not appreciably lower than the published proportion) together argue that a large hidden pool of null results is unlikely. The multiple-testing concern (using several dependent variables until one is significant) is acknowledged as a real threat that the flaw analysis partly addresses.
Outcome measure
The direct-hit rate, converted to Cohen’s h and combined via Stouffer z. The choice of h over the raw proportion difference is justified on detectability grounds and is standard Rosenthal practice; the measures are pre-stated and well defined.
Effect size
Mean h = .28 (a 38% hit rate), which the authors themselves shrink to about h = .18 (roughly a 33% hit rate) once statistical errors and file-drawer effects are allowed for. This self-imposed conservatism is a mark of careful meta-analysis: the reported effect is smaller than the raw figure, not larger.
Multiple comparisons
The flaw regression is explicit about multiplicity: 36 t-tests (three partialing methods by six predictors by two outcomes) were computed and none reached the .05 level, so the null relationship between flaws and outcomes is not an artifact of a single lucky test. The multiple-dependent-variable problem within primary studies is named as a residual concern.
Internal replication
The ten later Honorton studies, designed specifically to control the flaws Hyman and Honorton had identified, combined to z = 2.791 (mean h = .23), close to the original set and consistent with it. That a flaw-controlled replication reproduces the effect is the paper’s strongest internal check.
External replication
The quantitative results agree closely with both Honorton’s (1985) and Hyman’s (1985) independent analyses of the same database; the disagreement between those parties was over interpretation, not the numbers. Placed against later work, the effect here (a 38% raw rate) is larger than more recent, much larger ganzfeld meta-analyses report, consistent with a modest decline in estimated magnitude as databases grew.
Transparency
Strong for its era: every rival hypothesis is named and engaged, the flaw analysis and its sensitivity checks are reported in full, and the effect is deliberately adjusted downward. The transparency limits are inherent: it is a secondary analysis of Honorton’s compilation (so it inherits his inclusion decisions), the raw study-level data are not deposited (this predates open-data norms), and the 2024 scan carries OCR noise that a reader must see past.
The adversarial record
- The commissioning episode. The journal’s editorial abstract states that the authors were commissioned by the National Research Council, found ganzfeld research the most rigorous of the areas they reviewed, were pressured to withdraw their supportive evaluation, and refused. The published NRC volume (Druckman & Swets, 1988) reached a markedly more skeptical public conclusion about parapsychology, and this contrast is part of the historical record; the “pressured to withdraw” characterization is the republication’s framing.
- Agreement on the numbers. Notably, the paper records that Hyman (the leading skeptic) and Honorton (the leading proponent) agreed closely on the basic quantitative results, both on significance and on effect size. The dispute was about what the numbers meant, not what they were.
- Secondary, and dated. A committed critic can note that this is a re-analysis of Honorton’s own compilation rather than an independent literature search, that some included studies (Sargent’s) later attracted data-handling concerns, and that the whole analysis is now more than three decades old. The authors’ own sensitivity analyses answer the Sargent point (removing his studies barely moves the estimate), but the vintage is real: this is a foundational document, not current evidence.
- What survives. The durable contribution is the flaw analysis. Whatever one concludes about ESP, the paper demonstrates, with an explicit multivariate test, that the specific methodological flaws critics had catalogued did not statistically account for the ganzfeld results, and that a flaw-controlled replication reproduced them.
Contested record
Sources
- Harris, M. J., & Rosenthal, R. (2024). Parapsychology. Journal of Anomalous Experience and Cognition, 4(1), 18–33. https://doi.org/10.31156/jaex.26222 R001 [Harris & Rosenthal 2024] ↩︎
- Honorton, C. (1985). Meta-analysis of psi ganzfeld research: A response to Hyman. Journal of Parapsychology, 49(1), 51–91. R002 [Honorton 1985] ↩︎
- Hyman, R. (1985). The ganzfeld psi experiment: A critical appraisal. Journal of Parapsychology, 49(1), 3–49. R003 [Hyman 1985] ↩︎
- Hyman, R., & Honorton, C. (1986). A joint communiqué: The psi ganzfeld controversy. Journal of Parapsychology, 50, 351–364. R004 [Hyman & Honorton 1986] ↩︎
- Druckman, D., & Swets, J. A. (Eds.). (1988). Enhancing human performance: Issues, theories, and techniques. National Academy Press. R005 [Druckman & Swets 1988] ↩︎