Holmberg (2022)
Validating the GCP data hypothesis using internet search data
Holmberg, U. (2022). Validating the GCP data hypothesis using internet search data. Explore, 19(2), 228–237. https://doi.org/10.1016/j.explore.2022.07.007
AI Assessment
An econometric study testing the Global Consciousness Project (GCP) hypothesis indirectly: if mass-attention events really nudge the GCP’s random-number network, those same events should spike internet searches, so GCP data and Google Trends should correlate. Using time-series models, the author finds that monthly GCP aggregates do correlate positively with search indexes, improve model fit in 15 of 18 specifications, and can even sharpen out-of-sample forecasts (by up to 8.26%), though 2020 broke the pattern (attributed to COVID). The reproducibility is a real strength, both data sources are public. But the central inferential limit is that a correlation between GCP data and searches is equally consistent with both simply responding to real-world events by ordinary means, so it cannot by itself isolate a mind-matter mechanism, and the large specification space with no preregistration adds analytic latitude. This audit reports what the study did and found; it takes no position on whether the GCP hypothesis is true.
Provenance
DOI. 10.1016/j.explore.2022.07.007 · Explore 2022, volume 19, issue 2, pages 228 to 237 (available online 6 August 2022).
Study type. An observational time-series (econometric) analysis of two public archival data sources. It is not an experiment.
Author. Ulf Holmberg, an independent researcher based in Stockholm. The paper acknowledges comments from Roger D. Nelson, the GCP’s director.
Data availability. Both inputs are public: GCP data at noosphere.princeton.edu and Google Trends search data from Google, which makes the analysis reproducible in principle.
Source basis. Every figure below is taken from the article’s own Abstract and Results.
What the paper reports
The GCP runs a worldwide network of hardware random-number generators and hypothesizes that events drawing mass emotion or attention can shift their output away from chance. The author’s move is to test this indirectly: because people search the internet for information when engaging events occur, global search trends and GCP data should react to the same events, so they ought to correlate if the GCP hypothesis holds.1 The paper reports that they do correlate, that GCP data improves the statistical model’s in-sample fit, and that it can improve out-of-sample forecasts.
It is found that the GCP data significantly correlates with the indexes and can be used to improve the statistical model’s in-sample fit. Furthermore, it is found that out-of-sample forecasts can be made more accurate if the GCP data is used.
How it was run
- GCP aggregates. Second-by-second RNG data were bundled into 15-minute Stouffer Z-scores, reduced to daily Max and Average values, then averaged monthly (Max[Z] and Average[Z]).
- Search indexes. Google Trends data were combined into several indexes: a simple sum, a popularity-weighted sum, and a focused version (a binary indicator times a weight times the trend value). Some analyses used “boosted” indexes that added terms such as Earthquake, Hurricane, and Shooting.
- Models. ARMA models were fitted on first-differenced series with lagged GCP data as an exogenous predictor; models were compared by AIC, with heteroskedasticity-and-autocorrelation-consistent (HAC) standard errors for the regressions.
- Forecast test. Out-of-sample accuracy was measured as the reduction in root-mean-square error (RMSE) when conditioning on GCP data, using rolling three-year estimation windows forecasting one year ahead.
Results, as reported
| Metric | Result |
|---|---|
| Correlation with search indexes | monthly GCP aggregates (Max[Z] and Average[Z]) correlated significantly and positively across several index specifications |
| Which aggregate was more robust | Max[Z] was more consistently significant than Average[Z] |
| In-sample fit | adding GCP data improved AIC fit in 15 of 18 models tested |
| Boosted indexes | indexes adding Earthquake / Hurricane / Shooting terms showed stronger results (P < 0.01) |
| Subsample robustness | Max[Z] stayed highly significant (P < 0.01) in both subperiods; Average[Z] did not |
| Out-of-sample forecasting | RMSE reductions of up to 8.26% when conditioning on GCP data |
| Notable exception | 2020 showed anomalous worsening, attributed by the author to COVID-19 pandemic dynamics |
Values are reproduced from the article’s Abstract and Results. The pattern is one of generally supportive but uneven findings: the Max[Z] aggregate and the boosted indexes carry most of the signal, while the Average[Z] aggregate and the 2020 window do not.
Eleven-dimension audit
Pre-registration
Not preregistered. The analysis explores a sizable space of choices (three index constructions, boosted versus unboosted terms, two GCP aggregates, 18 model specifications), and the reported conclusions lean on the specifications that worked (Max[Z], boosted indexes). Without a registered plan, a reader cannot tell how much of the positive picture reflects specification search.
Randomization
Not applicable in the experimental sense: this is archival time-series analysis with no assignment of conditions. The “random” element is the GCP’s hardware RNG network, whose output is the object of study rather than a tool for allocation.
Sensory leakage
The analogue here is a common-cause confound rather than sensory leakage, and it is the study’s core inferential problem. Both GCP deviations (by hypothesis) and internet searches (obviously) respond to major world events, so a correlation between them is exactly what one expects whether or not consciousness affects the RNGs. The design cannot separate the psi mechanism from both series independently tracking newsworthy events.
Blinding
Not applicable; there are no participants or raters. Analyses are computational on fixed public datasets.
Optional stopping
Not applicable to archival data, but the analogous concern, model and specification selection across 18 models and multiple index types, is present and is not governed by a preregistered rule or a stated multiple-comparisons correction.
Outcome measure
Several measures are used: correlation significance, AIC-based in-sample fit, and out-of-sample RMSE reduction. Using out-of-sample forecasting is a genuine strength, because it is a harder test than in-sample fit; the weakness is that “up to 8.26%” reports the best case across specifications rather than a single pre-committed one.
Effect size
The reported effects are modest and specification-dependent: significant correlations, fit improvement in most (15 of 18) models, and forecast-error reductions up to a few percent. These are the kind of small, uneven effects that are suggestive but far from decisive.
Multiple comparisons
This is the second major concern. With three index constructions, boosted variants, two aggregates, subperiods, and 18 models, the analysis covers a large space, and the strongest results (Max[Z], boosted indexes) are highlighted. The paper does not describe a formal correction for this multiplicity, so the significance claims should be read with that latitude in mind.
Internal replication
There is a partial internal check: the subsample analysis shows Max[Z] holding across both subperiods while Average[Z] does not. That the more robust aggregate survives splitting is mildly reassuring, but the 2020 breakdown cuts the other way.
External replication
This is a novel indirect test rather than a replication of an established protocol, and it has not been independently replicated. It sits alongside, but is methodologically distinct from, the GCP’s own long-running event-based analyses.
Transparency
Strong on reproducibility: both data sources are public, and the modeling choices (aggregation, index construction, ARMA specification, HAC errors, rolling-window forecasting) are described in enough detail to be re-run. The main transparency gap is the absence of a preregistered specification and any stated multiplicity control.
The adversarial record
- Correlation does not validate the mechanism. The strongest objection is structural: GCP data and search trends should co-move if both react to big events, entirely without any mind-matter effect. A positive correlation is therefore consistent with the GCP hypothesis but does not distinguish it from the mundane alternative, so “validating” overstates what the design can show.
- Specification latitude. No preregistration, 18 models, multiple index constructions, hand-added “boost” terms, and no stated correction for multiple comparisons together give considerable room for favourable specifications to surface.
- Uneven results. The signal lives mainly in Max[Z] and the boosted indexes; Average[Z] is weaker, and the 2020 forecasting window worsened rather than improved. A critic will read the selective survival of certain specifications as a caution rather than a confirmation.
- What is genuinely valuable. The idea is clever and the execution is unusually reproducible: public inputs, an out-of-sample forecasting test that is harder to game than in-sample fit, and transparent methods. As a novel, checkable probe of the GCP data it has real merit, provided its correlational nature is kept firmly in view.
Sources
- Holmberg, U. (2022). Validating the GCP data hypothesis using internet search data. Explore, 19(2), 228–237. https://doi.org/10.1016/j.explore.2022.07.007 R001 [Holmberg 2022] ↩︎