Amorim Boyle (2025)
Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini
Amorim Boyle, B. J. (2025). Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini. Journal of Scientific Exploration, 39(3), 348–355. https://doi.org/10.31275/20253739
AI Assessment
A single 100-trial session in which an LLM beat chance on a forced-choice precognition task; statistically significant, but a small first test whose result hinges on an undocumented random generator. ChatGPT-4.1-mini selected the correct card 32 times in 100 five-choice trials (32% against a 20% baseline), significant on a pre-specified binomial test. The author is unusually candid that a proprietary server-side random generator, possible experimenter cueing, and the modest sample mean the result is suggestive, not conclusive.
Provenance
Source. Peer-reviewed Brief Report, open access (CC-BY-NC): Journal of Scientific Exploration, 39(3), pp. 348–355, 2025. Submitted 21 May 2025; accepted 1 June 2025; published 15 October 2025.
Study type. Single-session, 100-trial, double-blind, forced-choice precognition test of a large language model, with a five-alternative card task and a 20% chance baseline.
Funding. None stated. No human participants; the study was exempt from institutional review.
Data availability. The full trial-by-trial data are printed as Table 1 in the article’s appendix (all 100 trials: card selected, correct card, hit/miss).
Source basis. Figures confirmed against the published JSE article (pp. 348–355), Method and Results sections and the appendix data table.
What the paper reports
The sole “participant” was the language model ChatGPT-4.1-mini, prompted to pick which of five face-down cards concealed an image in PsiArcade’s “Find the Next Card” task; the target is drawn by the server only after the choice is registered, making this a precognition design with a 20% chance baseline. Across 100 trials the model scored 32 hits (32%), significantly above chance on the pre-specified exact binomial test (p = .005, two-tailed; Cohen’s h = 0.28).1
The author frames this as the first peer-reviewed, double-blind precognition test of an LLM and reports it with marked caution: a proprietary server-side random generator, possible experimenter cueing, a single small session, and ordinary statistical fluctuation are all offered as live non-psi explanations.
How it was run
- Single session of 100 forced-choice trials run on 19 May 2025; the model string was gpt-4.1-mini-2025-05-14, accessed through a paid OpenAI account. No warm-up or practice trials were discarded.
- On each trial the experimenter sent a verbatim templated prompt; the model returned an integer 1 to 5; the experimenter clicked that card on PsiArcade, and the server then drew and revealed the target. A hit was the model naming the card the server subsequently selected. Chance was one in five (20%).
- The PsiArcade laptop sat about 0.1 m from the phone running the ChatGPT app; the two devices were on Wi-Fi but not network-bridged (air-gapped), and the screen was visible only to the experimenter.
- The target was selected server-side, after the click, by PsiArcade’s pseudo-random generator; the platform does not publicly document the generator, which the author flags as a limitation.
- Each trial was followed by correct/incorrect feedback. Results were recorded in a spreadsheet with manual double-entry (100% agreement).
- The pre-specified primary analysis was a single exact two-tailed binomial test of 32 hits in 100 against a 0.20 baseline, with Cohen’s h and a 95% Clopper-Pearson confidence interval. No preregistration was filed.
Results, as reported
| Metric | Result |
|---|---|
| Hit rate | 32% (32/100) vs 20% chance |
| Primary significance | exact binomial p = .005 (two-tailed), Cohen’s h = 0.28 |
| 95% confidence interval (hit proportion) | 0.23 to 0.42 (Clopper-Pearson) |
| Learning trend across trials | none (logistic regression, n.s.) |
There was a single pre-specified outcome (the direct-hit rate against the 20% baseline). The effect size (h = 0.28) sits close to the small-to-moderate mean the author cites for human precognition work (Hedges g ≈ 0.20).
Eleven-dimension audit
Pre-registration
A single directional primary hypothesis (cumulative accuracy would exceed the 20% baseline) was stated in advance, but the author is explicit that no preregistration was filed. With one pre-specified test and no exploratory fishing this is less fraught than usual, but the design, sample size, and analysis were not independently locked before data collection.
Randomization
The target was drawn server-side by PsiArcade’s pseudo-random generator only after the model’s choice. The decisive caveat, which the author raises himself: the generator is proprietary and undocumented, so its entropy source and seeding cannot be inspected, and algorithmic predictability cannot be excluded as a non-psi explanation.
Sensory leakage
The target did not exist until after the choice was registered, so there was no concurrent sensory channel to it, and the devices were air-gapped. However, a human experimenter manually relayed every prompt and click, which the author notes could carry unconscious cueing (punctuation or timing); full automation would close this path.
Blinding
Double-blind in the sense that the model had no access to a target that did not yet exist, and scoring was an objective hit/miss against the server’s draw. There was no separate judge to blind. The experimenter, however, was in the loop for prompt delivery and data entry rather than running an automated pipeline.
Optional stopping
The session was a fixed 100 trials with no practice trials discarded and no indication of data-dependent stopping. Because nothing was preregistered, the fixed N rests on the author’s report rather than a binding pre-commitment; 100 trials is a small sample.
Outcome measure
The pre-specified primary measure was the cumulative direct-hit rate against the 0.20 baseline, tested with an exact binomial. It is objective and singular, with no alternative scoring substituted after the fact.
Effect size
Cohen’s h = 0.28 (small-to-moderate): a 12-percentage-point margin over chance. The author notes this is close to the mean human-precognition effect (Hedges g ≈ 0.20) cited in a recent meta-analysis.4
Multiple comparisons
A strength: there was a single pre-specified primary test (one binomial). A secondary logistic regression for a learning trend across trials was reported as non-significant. With one confirmatory test there is no multiplicity inflation to correct.
Internal replication
None. This is a single 100-trial session with one model variant; there is no within-study replication, and the author lists this as the leading limitation (results may not hold for other architectures or even other instantiations of the same model).
External replication
The author presents this as the first peer-reviewed double-blind precognition test of an LLM, so there is no direct prior replication. It sits within the contested human precognition literature, where positive forced-choice and “feeling the future” claims (Bem, 2011)2 are met by sustained statistical and methodological critique (Rouder and Morey, 2011),3 with a cumulative meta-analysis reporting a small overall effect.4
Transparency
High. The method, the verbatim prompts, the full 100-trial raw data table, and an explicit list of limitations and alternative explanations (random-generator opacity, experimenter cueing, statistical fluke, model memorisation) are all reported, and the paper is open access. The two transparency gaps are the absent preregistration and the proprietary generator.
The adversarial record
- The result rests on a single 100-trial session. The author’s own words: p = .005 “is impressive for a first attempt, but hardly definitive,” and “under a skeptical prior, Bayes factors would still demand repeated evidence.”
- The decisive weakness is the proprietary, undocumented random generator: if the server’s draws are deterministic or poorly seeded, a capable language model could in principle exploit periodicities, and the author concedes this cannot be ruled out.
- Reading against the study: a human experimenter manually delivered every prompt and recorded every result (possible unconscious cueing); the model may have ingested PsiArcade data during training (memorisation); and there was no preregistration. The author also flags the immediate per-trial feedback as a limitation in its own right: although it produced no learning trend, the feedback introduces conventional information that could complicate a strictly precognitive interpretation, which is why a no-feedback variant is among the recommended next steps. The author’s own recommended fixes (cryptographic or quantum-entropy generators, full automation, no-feedback variants, preregistration, and at least 1,000 trials) are exactly the controls a confirmatory test would need.
Sources
- Amorim Boyle, B. J. (2025). Testing Noetic Potential in Large Language Models: A 100-Trial Precognitive Forced-Choice Study with ChatGPT-4.1-Mini. Journal of Scientific Exploration, 39(3), 348–355. https://doi.org/10.31275/20253739 R001 [Amorim Boyle 2025] ↩︎
- Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524 R002 [Bem 2011] ↩︎
- Rouder, J. N., & Morey, R. D. (2011). The future of precognition research: Statistical and methodological recommendations. Review of Philosophy and Psychology, 2, 161–168. R003 [Rouder & Morey 2011] ↩︎
- Tressoldi, P., & Paladino, P. (2024). Precognition research 1978–2023: A cumulative meta-analysis and assessment of evidential value. Journal of Parapsychology, 88(1), 45–66. R004 [Tressoldi & Paladino 2024] ↩︎