Amorim Boyle (2025)

Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini

Amorim Boyle, B. J. (2025). Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini. Journal of Scientific Exploration, 39(3), 348–355. https://doi.org/10.31275/20253739

AI Assessment

A single 100-trial session in which an LLM beat chance on a forced-choice precognition task; statistically significant, but a small first test whose result hinges on an undocumented random generator. ChatGPT-4.1-mini selected the correct card 32 times in 100 five-choice trials (32% against a 20% baseline), significant on a pre-specified binomial test. The author is unusually candid that a proprietary server-side random generator, possible experimenter cueing, and the modest sample mean the result is suggestive, not conclusive.

Provenance

Source. Peer-reviewed Brief Report, open access (CC-BY-NC): Journal of Scientific Exploration, 39(3), pp. 348–355, 2025. Submitted 21 May 2025; accepted 1 June 2025; published 15 October 2025.

Study type. Single-session, 100-trial, double-blind, forced-choice precognition test of a large language model, with a five-alternative card task and a 20% chance baseline.

Funding. None stated. No human participants; the study was exempt from institutional review.

Data availability. The full trial-by-trial data are printed as Table 1 in the article’s appendix (all 100 trials: card selected, correct card, hit/miss).

Source basis. Figures confirmed against the published JSE article (pp. 348–355), Method and Results sections and the appendix data table.

What the paper reports

The sole “participant” was the language model ChatGPT-4.1-mini, prompted to pick which of five face-down cards concealed an image in PsiArcade’s “Find the Next Card” task; the target is drawn by the server only after the choice is registered, making this a precognition design with a 20% chance baseline. Across 100 trials the model scored 32 hits (32%), significantly above chance on the pre-specified exact binomial test (p = .005, two-tailed; Cohen’s h = 0.28).1

The author frames this as the first peer-reviewed, double-blind precognition test of an LLM and reports it with marked caution: a proprietary server-side random generator, possible experimenter cueing, a single small session, and ordinary statistical fluctuation are all offered as live non-psi explanations.

How it was run

Results, as reported

MetricResult
Hit rate32% (32/100) vs 20% chance
Primary significanceexact binomial p = .005 (two-tailed), Cohen’s h = 0.28
95% confidence interval (hit proportion)0.23 to 0.42 (Clopper-Pearson)
Learning trend across trialsnone (logistic regression, n.s.)

There was a single pre-specified outcome (the direct-hit rate against the 20% baseline). The effect size (h = 0.28) sits close to the small-to-moderate mean the author cites for human precognition work (Hedges g ≈ 0.20).

Eleven-dimension audit

Pre-registration

A single directional primary hypothesis (cumulative accuracy would exceed the 20% baseline) was stated in advance, but the author is explicit that no preregistration was filed. With one pre-specified test and no exploratory fishing this is less fraught than usual, but the design, sample size, and analysis were not independently locked before data collection.

Randomization

The target was drawn server-side by PsiArcade’s pseudo-random generator only after the model’s choice. The decisive caveat, which the author raises himself: the generator is proprietary and undocumented, so its entropy source and seeding cannot be inspected, and algorithmic predictability cannot be excluded as a non-psi explanation.

Sensory leakage

The target did not exist until after the choice was registered, so there was no concurrent sensory channel to it, and the devices were air-gapped. However, a human experimenter manually relayed every prompt and click, which the author notes could carry unconscious cueing (punctuation or timing); full automation would close this path.

Blinding

Double-blind in the sense that the model had no access to a target that did not yet exist, and scoring was an objective hit/miss against the server’s draw. There was no separate judge to blind. The experimenter, however, was in the loop for prompt delivery and data entry rather than running an automated pipeline.

Optional stopping

The session was a fixed 100 trials with no practice trials discarded and no indication of data-dependent stopping. Because nothing was preregistered, the fixed N rests on the author’s report rather than a binding pre-commitment; 100 trials is a small sample.

Outcome measure

The pre-specified primary measure was the cumulative direct-hit rate against the 0.20 baseline, tested with an exact binomial. It is objective and singular, with no alternative scoring substituted after the fact.

Effect size

Cohen’s h = 0.28 (small-to-moderate): a 12-percentage-point margin over chance. The author notes this is close to the mean human-precognition effect (Hedges g ≈ 0.20) cited in a recent meta-analysis.4

Multiple comparisons

A strength: there was a single pre-specified primary test (one binomial). A secondary logistic regression for a learning trend across trials was reported as non-significant. With one confirmatory test there is no multiplicity inflation to correct.

Internal replication

None. This is a single 100-trial session with one model variant; there is no within-study replication, and the author lists this as the leading limitation (results may not hold for other architectures or even other instantiations of the same model).

External replication

The author presents this as the first peer-reviewed double-blind precognition test of an LLM, so there is no direct prior replication. It sits within the contested human precognition literature, where positive forced-choice and “feeling the future” claims (Bem, 2011)2 are met by sustained statistical and methodological critique (Rouder and Morey, 2011),3 with a cumulative meta-analysis reporting a small overall effect.4

Transparency

High. The method, the verbatim prompts, the full 100-trial raw data table, and an explicit list of limitations and alternative explanations (random-generator opacity, experimenter cueing, statistical fluke, model memorisation) are all reported, and the paper is open access. The two transparency gaps are the absent preregistration and the proprietary generator.

The adversarial record

Sources
  1. Amorim Boyle, B. J. (2025). Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini. Journal of Scientific Exploration, 39(3), 348–355. https://doi.org/10.31275/20253739 R001 [Amorim Boyle 2025] ↩︎
  2. Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524 R002 [Bem 2011] ↩︎
  3. Rouder, J. N., & Morey, R. D. (2011). The future of precognition research: Statistical and methodological recommendations. Review of Philosophy and Psychology, 2, 161–168. R003 [Rouder & Morey 2011] ↩︎
  4. Tressoldi, P., & Paladino, P. (2024). Precognition research 1978–2023: A cumulative meta-analysis and assessment of evidential value. Journal of Parapsychology, 88(1), 45–66. R004 [Tressoldi & Paladino 2024] ↩︎