Rick E. Berger, PhD Sources:
Forensic Critique of Psi Datasets
Rick Berger‘s forensic audit work addressed a problem that cuts across the entire replication debate in parapsychology: the quality of the underlying data. Rather than simply running new experiments, he went back to raw source materials, dissertations, unpublished reports, original data logs, and compared them against what had been published, finding systematic discrepancies in datasets cited by both sides of the psi debate.
Deeper dives — Berger:
Key findings
- Berger’s 1989 forensic audit of Susan Blackmore’s ESP corpus found that discrepancies between unpublished dissertation data and published reports were pervasive enough that no conclusions, for or against psi, could be drawn from the database as a whole.1
- Berger documented that Blackmore’s published reports invoked methodological flaws selectively: significant results were dismissed by citing flaws, while non-significant results from equally flawed studies were reported without similar caveats.1
- The Spinelli dissertation audit addressed the opposite problem: data so implausibly strong that methodological problems, rather than genuine psi, were the more parsimonious explanation.2
- The 1998 PRL autoganzfeld target-distribution audit by Bierman, Broughton, and Berger found a peculiar over-representation of certain target numbers that could not be explained by the selection procedure, though this anomaly did not account for the overall hit excess in the autoganzfeld series.3
- Taken together, Berger’s audits established a bidirectional forensic standard: methodological scrutiny must be applied symmetrically to datasets that appear too weak and to datasets that appear too strong.
Overview
A recurring vulnerability in parapsychology’s replication debate is that meta-analyses and narrative reviews depend on the fidelity of the underlying published reports to the original data. When published summaries diverge from raw source materials, whether through selective reporting, post-hoc reframing of flaws, or chronological reordering of studies, the entire evidentiary structure built on those summaries becomes unreliable. Berger’s forensic audit work at the Science Unlimited Research Foundation (SURF) during 1985–1989 addressed this vulnerability directly, by returning to primary source documents and comparing them against what had reached print. The typical effect-size estimates in this literature are small to moderate, meaning that even modest analytic inconsistencies can shift a dataset from significant to non-significant or vice versa, making source-level auditing a high-leverage methodological intervention.
Why Source-Level Auditing Matters for Small-Effect Literatures
In research domains where true effect sizes are small (Cohen’s h or d in the range of 0.1–0.3), the difference between a significant and non-significant result often hinges on analytic decisions that are invisible in a published summary: which trials were included, how flawed sessions were classified, whether chronological ordering of studies was preserved. Berger’s approach, procuring unpublished dissertations and comparing them line-by-line against published counterparts, is a form of data provenance verification that is rarely attempted but is especially consequential in small-effect domains where publication bias and selective reporting have outsized leverage.1
The Blackmore Psi Experiments Audit
Susan Blackmore’s ESP experiments from the late 1970s and early 1980s reported null (non-significant) results across a range of psi paradigms, and her corpus was widely cited by skeptics as evidence against the existence of psi. Berger’s 1989 audit, published in the Journal of the American Society for Psychical Research, began as a meta-analysis of Blackmore’s published reports but pivoted when he obtained her unpublished doctoral dissertation, the original source for the subsequent publications. Comparison of the two revealed systematic discrepancies serious enough to invalidate the published record as a basis for any conclusion.1
Discrepancies Between Dissertation and Published Reports
Berger’s audit of the Blackmore corpus documented several categories of discrepancy between the unpublished dissertation and the published papers: (1) published reports did not veridically reflect the original data; (2) flaws were invoked selectively, significant results were dismissed by citing methodological problems, while non-significant results from studies Blackmore herself acknowledged as flawed in the dissertation were published without equivalent caveats; (3) admittedly flawed experiments were mixed with supposedly unflawed studies in published summaries without segregation, creating a misleading impression of methodological uniformity; (4) the chronological ordering of studies was reordered in at least two instances, potentially affecting how readers interpreted the trajectory of the research program. Berger concluded that no conclusions, for or against psi, could be drawn from the database as a whole, because the vast majority of the studies were, by Blackmore’s own assessment in the dissertation, individually flawed.1
Asymmetric Flaw Invocation as an Analytic Artifact
One of the most methodologically significant findings in the Blackmore audit was the asymmetric application of quality criteria: the same categories of procedural flaw were used to dismiss studies that produced significant results, but were not applied to dismiss studies that produced non-significant results. This asymmetry, what Berger characterized as a double standard in evaluating extraordinary claims, is a specific form of confirmation bias in analytic reporting.1 The artifact it produces is a database that appears to show consistent null results not because the experiments were uniformly well-conducted and uniformly non-significant, but because the reporting framework filtered out significant results via flaw-invocation while passing through non-significant results from equally flawed studies. Under the editorial standard adopted by ESP-Nexus (the C28 evidence-quality doctrine, which treats results from methodologically broken studies as carrying zero weight in either direction), the corpus contributes no decisive evidence in either direction: the Blackmore corpus, as audited, supports neither a pro-psi nor an anti-psi conclusion.1
The Spinelli Dissertation Audit
Where the Blackmore audit addressed a null dataset that skeptics cited as evidence against psi, the Spinelli dissertation audit addressed the opposite problem: a dataset whose reported effects were so implausibly large that they invited suspicion of methodological failure rather than genuine psi. Berger’s 1989 audit of the Spinelli dissertation data, published in the Journal of the Society for Psychical Research, documented procedural and analytic problems consistent with that suspicion.2
The “Too Good to Be True” Problem in Psi Research
In small-effect literatures, implausibly large effect sizes are a diagnostic signal for methodological problems, optional stopping, selective trial inclusion, undisclosed analytic flexibility, or outright data irregularities. The Spinelli dissertation reported ESP effects of a magnitude that fell well outside the range of the broader ganzfeld and forced-choice literature. Berger’s audit documented specific procedural and analytic problems in the dissertation that were consistent with the implausibility of the reported effects.2 Under the same C28 evidence-quality doctrine (methodologically broken studies carry zero evidential weight), the documented flaws mean the Spinelli data yields no conclusion in either direction, neither confirming nor disconfirming psi, because the methodological problems are sufficient to explain the reported results without invoking any anomalous process.
Target-Set Forensics: PRL Autoganzfeld Distributions
A decade after the Blackmore and Spinelli audits, Berger joined Dick J. Bierman and Richard S. Broughton in a forensic re-examination of the PRL autoganzfeld’s own target selection records, applying the same source-level scrutiny to the dataset that Berger himself had helped generate. The 1998 paper in the Journal of Parapsychology distinguished between a randomization procedure that is sound in principle and a resulting target sequence that may contain empirical anomalies, and examined both for the PRL autoganzfeld series.3
Target Frequency Anomalies and Their Implications
Bierman, Broughton, and Berger found that while the PRL autoganzfeld’s hardware RNG-based target selection procedure was sound in principle, using a noise-based generator that excluded sequential dependencies, the resulting target sequence contained a peculiar empirical anomaly: targets numbered 77, 78, 79, and 80 were over-represented, producing a strong over-representation of Set 20 (which contained those four targets). This distributional skew could not be explained by the selection procedure itself. The authors applied a conservative post-hoc correction for response bias and found a slight reduction in overall statistical significance and a weakening of the conclusion that dynamic targets were superior to static targets. Critically, however, the documented target-frequency anomaly could not account for the overall hit excess reported in the PRL autoganzfeld series, meaning the anomaly partially mitigated but did not explain away the positive results.3 The specific artifact addressed was response-bias confounding target-frequency imbalance; the specific control was a post-hoc correction adjusting hit probability for the observed distributional skew; the strength of ruling-out was partial, the correction reduced effect size estimates but left a residual excess of hits unexplained by the artifact.
Distinguishing Procedure Integrity from Sequence Integrity
A key conceptual contribution of the 1998 paper was the formal distinction between two levels of randomization quality: (1) the procedure level, whether the algorithm or hardware generator is theoretically sound and excludes sequential dependencies, and (2) the sequence level, whether the empirically realized sequence of targets, even from a sound procedure, happens to contain distributional anomalies that could interact with subject response biases. A hardware RNG can be procedurally valid while still producing, by chance, a sequence that over-represents certain targets. The paper provided post-hoc correction methods applicable when a closed-deck (balanced) design was not used, and noted that such corrections are unnecessary when the design is balanced.3
The Bidirectional Forensic Standard
The most distinctive feature of Berger’s forensic audit program is its symmetry. Most methodological critiques in parapsychology are unidirectional, skeptics scrutinize positive results for artifacts, while proponents scrutinize null results for suppression. Berger applied the same forensic standard to both tails: the Blackmore audit examined a null corpus that skeptics cited as evidence against psi and found it methodologically insufficient to support that conclusion; the Spinelli audit examined an implausibly positive corpus and found it methodologically insufficient to support a pro-psi conclusion. The PRL autoganzfeld target-distribution audit applied the same scrutiny to Berger’s own lab’s data.3
Implications for File-Drawer and Replication Debates
The file-drawer problem in parapsychology is typically framed as a concern about suppressed null results inflating apparent effect sizes in meta-analyses. Berger’s audits add a second dimension: even the null results that are published may not accurately represent the underlying data if published reports diverge from source materials. The Blackmore audit demonstrated this concretely, a corpus of null results that had been widely cited as a replication failure turned out, on source-level examination, to be methodologically too compromised to serve as evidence in either direction.1 This reframes the replication debate: the question is not only whether positive results replicate, but whether the null results used as the comparison baseline are themselves methodologically sound. A null result from a broken study, under the C28 evidence-quality doctrine, carries zero evidential weight against the hypothesis, just as a positive result from a broken study carries zero weight for it.
Open Questions for Future Forensic Work
Berger’s audits were conducted in an era before data-sharing mandates, preregistration infrastructure, or open-science norms in parapsychology. The specific artifacts he documented, discrepancies between dissertation and published data, asymmetric flaw invocation, target-frequency distributional anomalies, are addressable in principle by modern open-science practices: preregistered analysis plans, public data repositories, and independent data audits prior to publication. Whether the broader ganzfeld and forced-choice literatures contain additional source-level discrepancies of the kind Berger documented in the Blackmore corpus remains an open empirical question that would require systematic dissertation-to-publication comparison studies to answer.1
Modern Context
Berger’s forensic-audit methodology — fine-grained comparison of raw dissertation data against published-paper claims, with systematic accounting of every discrepancy — anticipated by roughly two decades the data-forensics literature that emerged from the replication crisis. Simmons, Nelson, and Simonsohn (2011) showed that undisclosed analytic flexibility produces selective-reporting bias of exactly the kind Berger documented in the Blackmore corpus.4 The Open Science Collaboration (2015) reproducibility project formalized the principle that single-study significant effects systematically over-estimate true effect magnitudes — a concern Berger raised in 1989 by demonstrating that selectively-reported subsets of the Blackmore data shifted the effect-size estimate by an order of magnitude.5 The general principle that bidirectional auditing (applying the same forensic standards to skeptical and supportive datasets) is essential to disciplinary integrity is now mainstream in meta-research; Berger’s symmetric audits of Spinelli (a pro-psi dataset) and Blackmore (a skeptical dataset) were unusual for their era in modeling this norm.
Skeptical Critiques and Discussion
Critique 1: The Blackmore null corpus constitutes genuine replication failure, and auditing it is motivated reasoning by a psi proponent
Skeptic source: Blackmore’s own published accounts across multiple papers from 1980 through 1988 presented her ESP experiments as scrupulously conducted and consistently null, framing the corpus as evidence that a committed researcher found no sign of psi after a decade of work. The implicit critique of Berger’s audit is that it selectively targets a prominent skeptic’s work while the broader null replication literature is left unexamined.1
Response: Berger’s audit was not initiated as a targeted attack on Blackmore’s conclusions but began as a straightforward meta-analysis of her published reports, one that initially suggested possible psi effects in the database. The pivot to source-level auditing occurred only after procurement of the unpublished dissertation revealed that the published reports did not veridically reflect the original data.1 The bidirectional character of Berger’s forensic program, applying identical scrutiny to the implausibly positive Spinelli data2 and to the PRL autoganzfeld’s own target distributions3, is the strongest structural rebuttal to the motivated-reasoning charge: a researcher engaged in motivated reasoning would not audit their own lab’s data for artifacts.
Analysis. The core finding, that published reports diverged from dissertation source materials, is documented in the audit itself and has not been formally rebutted at the data level. The motivated-reasoning charge is a meta-level critique that the bidirectional audit record addresses in part by documenting source-level comparisons in both directions; independent replication of the source-level comparison by investigators outside the audit team has not been published.
Critique 2: The PRL autoganzfeld target-frequency anomaly undermines the positive results of the entire autoganzfeld series
Skeptic source: The finding that targets 77–80 and Set 20 were over-represented in the PRL autoganzfeld sequence, an anomaly that could not be explained by the hardware RNG procedure, raises the concern that subject response biases interacting with this distributional skew could account for some or all of the reported hit excess, even if the randomization procedure was sound in principle.3
Response: Bierman, Broughton, and Berger applied a conservative post-hoc correction for the response-bias artifact introduced by the target-frequency imbalance and found that the correction produced only a slight reduction in overall statistical significance.3 The specific artifact, response-bias confounding from target over-representation, was partially mitigated by the correction, but the residual hit excess remained unexplained by the distributional anomaly alone. The authors explicitly stated that the documented deviations from a balanced target-frequency distribution cannot explain the excess of hits reported for the PRL autoganzfeld study. The anomaly weakened one secondary conclusion (dynamic target superiority) but left the primary positive result standing after correction.
Analysis. The artifact is real and documented; the correction is methodologically appropriate; the residual effect after correction is not explained by the artifact. The finding is self-reported by the same team that ran the original study, which is a strength (no motivated suppression) but also means independent replication of the correction analysis has not been published.
References
- Berger, R. E. (1989). A critical examination of the Blackmore psi experiments. Journal of the American Society for Psychical Research, 83(2), 123–144. R001 [Berger 1989-Blackmore] ↩︎
- Berger, R. E. (1989). A critical examination of the Spinelli dissertation data. Journal of the Society for Psychical Research, 56, 28–34. R002 [Berger 1989-Spinelli] ↩︎
- Bierman, D. J., Broughton, R. S., & Berger, R. E. (1998). Notes on Random Target Selection: The PRL Autoganzfeld Target and Target Set Revisited. Journal of Parapsychology, 62(4), 339–350. R003 [Bierman 1998] ↩︎
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. https://doi.org/10.1177/0956797611417632 R004 [Simmons 2011] ↩︎
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716 R005 [Open Science Collaboration 2015] ↩︎
Deeper dives — Berger: