Roger D. Nelson, PhD Sources:

Meta-Analysis and Statistical Methods for Consciousness Effects

Roger Nelson‘s contributions to parapsychology extend well beyond designing experiments, he has been a central figure in developing and defending the statistical frameworks used to aggregate and evaluate evidence for mind-matter interaction (MMI) effects across decades of random number generator (RNG) research. His meta-analytic work, conducted primarily in collaboration with Dean Radin, established quantitative benchmarks for the field and became the focal point of sustained methodological debate about publication bias, effect-size heterogeneity, and the assumptions underlying RNG-based consciousness research.

Key findings

  • A 1989 meta-analysis of over 800 RNG experiments found non-chance effects in experimental conditions while control conditions conformed to chance expectation, providing early quantitative evidence for consciousness-related anomalies in random physical systems.1
  • An updated meta-analysis covering 1959–2000 extended the earlier findings across a larger corpus of mind-matter interaction experiments, maintaining the pattern of small but statistically significant effects.2
  • Nelson and colleagues argued that the “influence-per-bit” assumption, the premise that MMI effects operate uniformly on each random bit regardless of psychological context, is empirically unwarranted and, when applied, artificially generates the appearance of publication bias.3
  • In a direct response to the Bösch, Steinkamp, and Boller (2006) meta-analysis published in Psychological Bulletin, Nelson and co-authors agreed that the data show statistical evidence for a psychokinetic effect and that study quality is generally high, but contested the interpretation that heterogeneity is explained by selective reporting.3
  • Publication bias and low experimental quality were assessed as implausible primary explanations for the cumulative RNG meta-analytic results, with specific arguments addressing why the file-drawer model does not account for the observed data structure.4

Overview

A persistent challenge in evaluating parapsychological evidence is that individual experiments are typically underpowered to detect small effects, making cumulative synthesis essential. Nelson’s meta-analytic work addressed this directly: by aggregating hundreds of RNG experiments, he and Radin sought to determine whether the pattern of results across the literature was consistent with a genuine small effect or with statistical noise amplified by selective reporting. A key Type-II vulnerability in this literature is that any genuine effect, if small (as the data suggest), would be invisible in most individual studies, meaning dismissal based on single-study failures to replicate would be premature without accounting for the aggregate signal.1

Scope of the RNG Literature Reviewed

The 1989 Radin and Nelson meta-analysis in Foundations of Physics identified more than 800 relevant experiments in the parapsychological literature bearing on consciousness-related anomalies in random physical systems, contrasting sharply with only three studies found in physics journals proper at that time. The review applied meta-analytic techniques to assess both methodological quality and overall effect size across this corpus. Control conditions conformed to chance expectation; experimental conditions showed unequivocal non-chance effects. The authors noted agreement with two earlier independent reviews, strengthening the cumulative inference.1

Foundational Meta-Analyses

The 1989 Radin–Nelson meta-analysis was the first systematic quantitative review of RNG-based consciousness research to appear in a mainstream physics journal, establishing a methodological template for the field. A subsequent update covering experiments from 1959 to 2000 extended the corpus and reaffirmed the pattern of small positive effects across individual intention studies.2 1

1989 Foundations of Physics Meta-Analysis, Methodology and Results

Radin and Nelson’s 1989 review, published in Foundations of Physics, applied meta-analytic techniques to a corpus of over 800 experiments examining whether human intention could influence the outputs of random physical systems. The analysis assessed methodological quality across studies and computed an overall effect size. The key finding was a dissociation between experimental and control conditions: control runs conformed to chance expectation, while experimental runs showed statistically significant non-chance deviations. The authors framed this as quantitative evidence for a consciousness-related anomaly, situating it within ongoing physics discussions about the role of observation in quantum mechanics.1 The competing non-psi explanation most directly addressed was that the effect was an artifact of selective reporting (file-drawer problem); the authors argued that the number of null studies required to nullify the observed effect was implausibly large given the structure of the data.

1959–2000 Update, Extended Corpus

The 2000–2003 update by Radin and Nelson extended the meta-analytic corpus to cover mind-matter interaction experiments from 1959 through 2000, providing a longer time-series view of the literature.2 The update maintained the pattern of small positive effects in individual intention studies. The analysis also addressed the specific non-psi alternative of optional stopping, the possibility that experimenters terminated runs selectively when results were favorable, by examining whether effect sizes were correlated with study characteristics that would be expected under optional stopping but not under a genuine effect model.

Individual Intention vs. Group Effects, Scope Distinction

Nelson and Radin’s meta-analytic work distinguished between individual intention studies (single operators attempting to influence RNG outputs) and group or mass-consciousness studies. The individual intention corpus, reviewed in detail in their 2003 chapter,2 formed the primary evidentiary base for the meta-analyses. The group-consciousness literature, which includes the Global Consciousness Project and FieldREG studies, involves different experimental designs and is treated as a separate research stream. This scope distinction matters for interpreting effect sizes: individual intention studies show small but consistent effects, while group studies involve different statistical models and event-selection criteria.

The Publication Bias Debate

The most sustained methodological challenge to the Radin–Nelson meta-analyses has been the claim that the observed cumulative effects are artifacts of selective reporting, that null results remain unpublished in file drawers, inflating the apparent effect size in the published literature. Nelson and colleagues engaged this critique directly and at length, arguing that the file-drawer model is implausible given the specific structure of the RNG data.4

File-Drawer Analysis, How Many Null Studies Would Be Required?

A standard tool for assessing publication bias in meta-analyses is the fail-safe N calculation: how many unpublished null studies would need to exist to reduce the observed cumulative effect to non-significance? Radin, Nelson, Dobyns, and Houtkooper (2006) argued in their Journal of Scientific Exploration response that the fail-safe N for the RNG corpus is implausibly large, far exceeding what the research infrastructure of the field could plausibly have produced and suppressed.4 They further noted that the specific artifact of publication bias, selective reporting, would be expected to produce a particular pattern of effect-size distribution (funnel-plot asymmetry), and that the observed distribution does not straightforwardly match that pattern. The multiple-comparisons artifact was also addressed: the authors argued that the primary analyses were pre-specified in terms of the outcome measure (RNG deviation from chance), limiting the scope for post-hoc selection of favorable statistics.

Bösch et al. (2006) Meta-Analysis, Points of Agreement and Disagreement

The Bösch, Steinkamp, and Boller (2006) meta-analysis, published in Psychological Bulletin, independently reviewed the RNG psychokinesis literature and reached a nuanced verdict. Radin, Nelson, Dobyns, and Houtkooper’s published comment3 noted three points of agreement with Bösch et al.: (1) the existing data indicate the existence of a statistically significant PK effect; (2) the studies are generally of high methodological quality; and (3) effect sizes are distributed heterogeneously across studies. The disagreement was entirely about the source of that heterogeneity. Bösch et al. attributed heterogeneity to selective reporting and supported this with an ad hoc Monte Carlo simulation. Nelson and colleagues argued that Bösch et al.’s simulation rested on the influence-per-bit assumption (see next section), which they contended is empirically unjustified, and that once this assumption is relaxed, the heterogeneity is expected under a genuine effect model rather than a publication-bias model.

The Influence-Per-Bit Assumption

A technical dispute at the heart of the RNG meta-analysis debate concerns what Nelson and colleagues called the “influence-per-bit” assumption: the premise that if mind-matter interaction exists, it must operate uniformly on each randomly generated bit, independent of the number of bits per sample, the generation rate, or the psychological conditions of the experiment. Nelson and co-authors argued this assumption is not only unwarranted but, when applied, mechanically produces the appearance of publication bias even in data generated by a genuine effect.3 4

Why the Influence-Per-Bit Assumption Generates Spurious Heterogeneity

Radin, Nelson, Dobyns, and Houtkooper (2006) illustrated the problem with the influence-per-bit assumption through a thought experiment contrasting two studies with identical physical setups but radically different psychological contexts: one involving 1,000 experienced meditators with high motivation and pre-training, the other involving a single unmotivated undergraduate with no feedback and no consequences.3 Both studies generate 1,000 random bits, but the psychological conditions differ maximally. If MMI is sensitive to psychological context, as virtually all human performance phenomena are, then the two studies would be expected to produce different effect sizes. The influence-per-bit assumption treats them as equivalent, predicting identical effect sizes. When the observed data show heterogeneity (as they do), the assumption leads to the conclusion that the heterogeneity must be artifactual (i.e., due to selective reporting), when in fact it may reflect genuine variation in the conditions that modulate the effect. The authors argued this is a category error: the assumption imports a specific mechanistic model of MMI that has no empirical support, then uses violations of that model as evidence against the phenomenon itself.4

Competing Non-Psi Explanations Considered

Nelson and colleagues explicitly addressed several non-psi explanations for the cumulative RNG meta-analytic results.4 (1) Publication bias / file-drawer: addressed via fail-safe N calculations and distribution-shape arguments (see above); assessed as implausible given the corpus size. (2) Low experimental quality: addressed by noting that Bösch et al. themselves rated the studies as generally high quality, and that quality ratings did not predict effect size in the direction expected under an artifact model. (3) Multiple comparisons: addressed by noting that the primary outcome measure (RNG deviation from chance expectation) was pre-specified in the experimental designs, limiting post-hoc flexibility. (4) Optional stopping: addressed by examining whether effect sizes correlated with study characteristics associated with optional stopping; no such correlation was found. The authors acknowledged that no single argument definitively eliminates all alternatives, but argued that the combination of high study quality, large fail-safe N, and the implausibility of the influence-per-bit assumption collectively makes selective reporting an insufficient explanation.

Statistical Methods, Stouffer Z and Cumulative Probability

The primary statistical method used in the Radin–Nelson meta-analyses was the Stouffer Z method for combining independent p-values across studies, which converts each study’s result to a standard normal deviate and sums them weighted by sample size. This approach is appropriate for combining studies with a common directional hypothesis (intention to increase or decrease RNG output) and is robust to heterogeneity in study design when the hypothesis is consistently operationalized. The 1989 meta-analysis1 and subsequent updates2 both reported cumulative Stouffer Z values far exceeding conventional significance thresholds. The 1990 Eastern Psychological Association proceedings5 presented an early version of this cumulative analysis. Effect sizes in the individual intention literature are small (typically in the range of r ≈ 0.02–0.05), meaning that the statistical significance of the cumulative result depends entirely on the large aggregate sample size across hundreds of studies, not on any single dramatic experiment.

Modern Context

Nelson’s use of meta-analytic aggregation to characterize cumulative evidence across hundreds of RNG and FieldREG studies sits within the mainstream meta-analysis methodology established by Borenstein, Hedges, Higgins, and Rothstein’s canonical textbook and the random-effects-model literature it consolidates. Mainstream meta-analysis distinguishes fixed-effect from random-effects models, requires explicit heterogeneity characterization (I², τ²), and treats publication-bias diagnostics as load-bearing.6 The publication-bias problem is particularly acute for small-effect, file-drawer-prone literatures: the Sterne et al. 2011 BMJ recommendations standardize funnel-plot interpretation and explicitly warn that asymmetry can reflect multiple competing causes (true publication bias, study-size×effect-size correlation, methodological heterogeneity), requiring triangulation across diagnostics rather than reliance on any single test.7 Duval and Tweedie’s trim-and-fill procedure provides one widely-deployed adjustment for funnel asymmetry, producing sensitivity bounds on the meta-analytic effect estimate under conservative assumptions about missing studies. Applied to the GCP/PEAR/PK corpus, these mainstream tools produce the residual-effect estimates that remain after the most aggressive publication-bias corrections — small but reportedly non-zero.8

Skeptical Critiques and Discussion

Critique 1: Claim: Effect-size heterogeneity in the RNG literature is best explained by selective reporting, rendering the psychokinesis hypothesis “not proven”

Skeptic source: Bösch, Steinkamp, and Boller’s (2006) meta-analysis, published in Psychological Bulletin, found that effect sizes across RNG studies were distributed heterogeneously and argued, supported by a Monte Carlo simulation, that this heterogeneity is most parsimoniously explained by selective reporting of favorable results rather than a genuine psychokinetic effect. Their verdict was “not proven” in the Scottish legal sense: evidence too strong to ignore but too weak to compel conviction.3

Response: Radin, Nelson, Dobyns, and Houtkooper’s published comment in the same issue of Psychological Bulletin accepted the empirical findings, statistical evidence for a PK effect, generally high study quality, heterogeneous effect sizes, but contested the interpretation. They argued that Bösch et al.’s Monte Carlo simulation rested on the influence-per-bit assumption, which predicts uniform effect sizes across all psychological contexts. Because this assumption is empirically unjustified (no human performance phenomenon is context-independent in this way), the heterogeneity is expected under a genuine effect model and does not constitute evidence for selective reporting. The specific artifact of experimenter cueing was partially addressed by the hardware RNG designs (which remove direct experimenter contact with the random process), though experimenter-participant interaction during task setup remains unresolved.3 4

Analysis. The heterogeneity-as-selective-reporting critique addresses whether the distribution of RNG effect sizes is most parsimoniously explained by selective reporting rather than a genuine effect. Both Bösch et al. (2006) and the Radin–Nelson–Dobyns–Houtkooper reply in the same Psychological Bulletin issue agree on the empirical facts (significant cumulative effect, high study quality, heterogeneous effect sizes); the dispute concerns the influence-per-bit assumption, which Bösch et al.’s Monte Carlo uses to predict uniform effect sizes across psychological contexts. Whether that uniformity assumption is empirically warranted — and what statistical model of heterogeneity should be considered the null — remains a theoretical commitment on which the meta-analytic literature has not converged.

Critique 2: Claim: The cumulative RNG meta-analytic results are artifacts of publication bias, with an implausibly large number of null studies suppressed in file drawers

Skeptic source: Multiple critics of the Radin–Nelson meta-analyses, including Schub (addressed in the 2006 Journal of Scientific Exploration response) and Ehm (2005), argued that the significant cumulative results are explained by selective reporting practices, that the field’s publication norms favor positive results, and that the true effect size is zero once unpublished null studies are accounted for. This critique treats the influence-per-bit assumption as valid and uses it to calculate expected effect sizes under a null hypothesis, finding that the observed distribution is consistent with selective reporting.4

Response: Radin, Nelson, Dobyns, and Houtkooper responded that publication bias is an implausible primary explanation for three reasons: (1) the fail-safe N, the number of unpublished null studies required to nullify the cumulative effect, is implausibly large relative to the field’s research output; (2) the specific artifact of publication bias (file-drawer suppression) would be expected to produce funnel-plot asymmetry, but the observed distribution does not straightforwardly match this pattern; and (3) the influence-per-bit assumption, which underlies the critics’ calculations, is not empirically supported. The multiple-comparisons artifact was partially addressed by the pre-specified nature of the primary outcome measure (RNG deviation from chance), though the authors acknowledged that across-study variation in secondary analyses remains a partially unresolved concern.4

Analysis. At issue is whether the cumulative RNG meta-analytic results are explained by file-drawer suppression of null studies. The Radin–Nelson–Dobyns–Houtkooper response cites a fail-safe N calculation large relative to the field’s research output, notes that the observed effect-size distribution does not straightforwardly match the funnel-plot asymmetry signature of publication bias, and argues that the underlying influence-per-bit assumption (which drives the critics’ null-effect calculations) is not empirically supported. The strength of the publication-bias inference depends on which heterogeneity model is adopted, and this remains an unresolved theoretical question across the meta-analytic literature.

References
  1. Radin, D. I., & Nelson, R. (1989). Evidence for consciousness-related anomalies in random physical systems. Foundations of Physics, 19(12), 1499–1514. https://doi.org/10.1007/bf00732509 R001 [Radin 1989] ↩︎
  2. Radin, D. I., & Nelson, R. D. (2003). Research on mind-matter interactions (MMI): Individual intention. In W. B. Jonas & C. C. Crawford (Eds.), Healing, intention, and energy medicine: Science, research methods, and clinical implications (pp. 39–48). Churchill Livingstone. https://doi.org/10.1016/b978-0-443-07237-6.50009-7 R002 [Radin 2003] ↩︎
  3. Radin, D. I., Nelson, R., Dobyns, Y., & Houtkooper, J. M. (2006). Reexamining psychokinesis: Comment on Bösch, Steinkamp, and Boller (2006). Psychological Bulletin, 132(4), 529–532. https://doi.org/10.1037/0033-2909.132.4.529 R003 [Radin 2006] ↩︎
  4. Radin, D. I., Nelson, R., Dobyns, Y., & Houtkooper, J. M. (2006). Assessing the Evidence for Mind-Matter Interaction Effects. Journal of Scientific Exploration, 20(3), 361–374. https://web.archive.org/web/20240712074323/https://www.scientificexploration.org/docs/20/jse_20_3_radin_1.pdf R004 [Radin 2006] ↩︎
  5. Nelson, R. D., & D. I. Radin (1990). Effects of human intention on random event generators: A meta-analysis. Proceedings & Abstracts of the Eastern Psychological Association. R005 [Nelson 1990] ↩︎
  6. Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to Meta-Analysis. John Wiley & Sons. https://doi.org/10.1002/9780470743386 R006 [Borenstein 2009] ↩︎
  7. Sterne, J. A. C., Sutton, A. J., Ioannidis, J. P. A., Terrin, N., Jones, D. R., Lau, J., et al. (2011). Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ, 343, d4002. https://doi.org/10.1136/bmj.d4002 R007 [Sterne 2011] ↩︎
  8. Duval, S., & Tweedie, R. (2000). Trim and fill: A simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics, 56(2), 455–463. https://doi.org/10.1111/j.0006-341x.2000.00455.x R008 [Duval 2000] ↩︎