Mossbridge et al. (2024)

State, Trait, and Target Parameters Associated with Accuracy in Two Online Tests of Precognitive Remote Viewing

Mossbridge, J., Cameron, K., & Boccuzzi, M. (2024). State, Trait, and Target Parameters Associated with Accuracy in Two Online Tests of Precognitive Remote Viewing. Journal of Anomalous Experience and Cognition, 4(1), 88–121. https://doi.org/10.31156/jaex.24743

AI Assessment

A two-experiment report that finds significant free-response target precognition while its forced-choice experiment stays at or below chance, and that states its exploratory-to-confirmatory path and its own analysis missteps openly. The numbers in the abstract, results sections, and figure captions are internally consistent, and the paper discloses its pre-registration mismatch, its bot and inattention exclusions, and its experimenter belief levels. The confirmatory predictions were filed with a granting agency rather than a public registry, experiment 1 ran in an uncontrolled online environment by design, and the experiment 2 scoring produced non-target results the authors themselves describe as difficult to interpret.

Provenance

DOI. 10.31156/jaex.24743 · Open access, Copyright 2024 The Author(s), CC-BY License. Article page at the journal: journals.lub.lu.se/jaex/article/view/24743.

Authors. Julia Mossbridge (University of San Diego), Kirsten Cameron (TILT), and Mark Boccuzzi (Windbridge Institute), as printed in the article byline. Correspondence is addressed to Julia Mossbridge, Ph. D.

Study type. Two online precognitive remote viewing (PRV) experiments examining trait, state, and target parameters: a forced-choice, uncontrolled-time, self-judged task (experiment 1) and a free-response, controlled-time, independently judged task (experiment 2). Neither experiment pre-screened participants for precognition ability.

Funding. The authors thank the Bial Foundation for funding the first author (Bial grants 2014_260, 97_16, 369/20). Theresa Cheung funded the website that allowed data gathering for the first analysis set, and The Windbridge Institute built and maintained the website and software used to gather data for both analysis sets.

Data availability. The paper states for experiment 1 that “Raw data are available upon request”; appendix figures showing the target sets are linked from the text. No data-availability statement is given for experiment 2.

Editorial note. The article’s title footnote ends: “See also Letter to the Editor in this issue.” The paper itself does not describe that letter’s contents.

Source basis. Figures confirmed against the primary article (publisher PDF, Journal of Anomalous Experience and Cognition, 4(1), 88–121).

What the paper reports

Across two online experiments, the paper examines how age, sex-at-birth, gender, anxiety, unconditional love, and target interestingness relate to accuracy on precognitive remote viewing tasks.1 The report follows a stated exploration-to-confirmation discovery process modeled on the “SEARCH” strategy of a previous examination of online psi task performance.2 In experiment 1, a forced-choice, uncontrolled-time, self-judged task on the website ThePremonitionCode.com/tester, 682 unpaid participants contributed 5,432 trials across two data batches. There was no significant target precognition: batch-1 practice trials reached a proportion correct of 0.522 (binomial p < .10) and test trials 0.50 (p > .99), while batch-2 test trials scored 0.45, significantly below the 0.50 chance level (p < .01), a small expectation-opposing effect. Batch-1 participants showed what the authors call a massive bias toward selecting Profile A, the option shown in the left or top screen position (proportion 0.58, binomial p < 3×10−16). Targets most likely to be correctly predicted were rated more interesting than targets most likely to be incorrectly predicted by independent raters not informed about the experiment (batch 1: t(126) = 3.49, p < .0007; batch 2, pre-registered confirmation: t(48) = 5.03, p < .000008), and participants reporting female sex-at-birth scored below chance (t(287) = 2.68, p < .008) while males did not differ from chance, a significant difference favoring men (t(444) = 2.10, p < .04). In experiment 2, a two-minute, single-trial, free-response task, 307 paid participants each contributed one trial, and significant target precognition was evident: both judges agreed the transcript matched the target image on 35% of trials versus 25% chance expectation (p < .0002, h = .22). Agreed matches to the non-target image also exceeded chance (42% vs 25%, p < .00001, h = .36), and judge disagreements fell below chance (23% vs 50%, p < .00001, h = −.57). Greater feelings of unconditional love were associated with better accuracy (t(192) = 2.50, p < .01), confirming the paper’s first prediction and a prior exploratory finding by the first author’s team,3 greater anxiety was marginally associated with greater rather than lower accuracy (t(217) = 1.990, p < .048), weakly opposing the second prediction, and target interestingness again correlated with accuracy (r(84) = .21, p < .05), confirming the fifth prediction.

In experiment 2, agreed matches to the non-target images (42%) exceeded agreed matches to the targets (35%), and the authors state it remains difficult to tease out whether the non-target images were also precognized at a rate above chance.

How it was run

Results, as reported

MetricResult
Trials, experiment 1 (forced-choice)5,432 trials by 682 unpaid participants (3,003 in batch 1; 2,429 in batch 2)
Overall accuracy, batch 1Practice 0.522 vs .50 chance (binomial p < .10); test 0.50 (binomial p > .99)
Overall accuracy, batch 2Practice 0.49 (binomial p > .79); test 0.45 vs .50 chance (binomial p < .01, below chance); practice vs test χ²(1, N = 2429) = 3.9, p < .05
Profile A selection biasBatch 1: 0.58 vs .50 chance (binomial p < 3×10−16); batch 2: 0.51 (p > .17)
Interestingness, top-correct vs top-incorrect targets (across-responses)Batch 1: t(126) = 3.49, p < .0007; batch 2 (pre-registered): t(48) = 5.03, p < .000008
Interestingness vs incorrect-to-correct trial ratioBatch 1: r(14) = −0.45, p < .08; batch 2: r(8) = −0.70, p < .03
Age and accuracy (experiment 1)r(443) = −.003, p > .95
Sex-at-birth and accuracy (experiment 1)Female below chance: t(287) = 2.68, p < .008; male at chance: t(156) = 0.63, p = .53; difference favoring men: t(444) = 2.10, p < .04
Trials, experiment 2 (free-response)307 paid participants, one trial each; 18 apparent bots removed and replaced
Agreed target matches (experiment 2)35% vs 25% chance, p < .0002, h = .22
Agreed non-target matches (experiment 2)42% vs 25% chance, p < .00001, h = .36
Judge disagreements (experiment 2)23% vs 50% chance, p < .00001, h = −.57
Control judging (true target never shown)Yoked targets 26% vs 25% chance (p > .69, h = .03); yoked non-targets 41% vs 25% chance (p < .00001, h = .34); no-agreement 33% vs 50% chance (p < .00001, h = .35)
Original vs control judging distributionsχ²(1, N = 614) = 8.5, p < .015
Unconditional love and accuracy (prediction 1)t(192) = 2.50, p < .01; χ²(2, N = 194) = 10.0, p < .007
Anxiety and accuracy (prediction 2, direction opposed)t(217) = 1.990, p < .048; χ²(2, N = 219) = 4.2, p > .12
Anxiety and accuracy by gender (prediction 3)Women: χ²(2, N = 116) = 8.0, p < .02; men: χ²(2, N = 95) = 2.80, p < .10
Reproductive hormones (prediction 4)Women taking vs not taking hormones: t(154) = .82, p > .41; χ²(2, N = 156) = .68, p > .71
Interestingness and accuracy, experiment 2 (prediction 5)r(84) = .21, p < .05

The paper reports proportions, t, r, and chi-squared statistics with p values throughout, and effect sizes (h) for the experiment 2 binomial comparisons. It reports no standardized effect sizes for the experiment 1 comparisons and no confidence intervals for any comparison, so none are stated here.

Eleven-dimension audit

Pre-registration

The interestingness analysis of experiment 1 was pre-registered before the second batch of data was downloaded, and the paper states that four formal predictions for experiment 2 were submitted to a granting agency before the experiment, with a fifth confirmatory prediction added after the analysis of experiment 1 was complete. The paper does not name a public registry entry for either experiment. It also discloses that the pre-registration was completed before the authors discovered they had failed to remove data from the website developer and the experimenter, so “the numerical results presented here do not match the pre-registration document approved prior to discovering this oversight.” One prediction was changed before data collection: the plan originally submitted to the granting agency targeted pregnant women, and IRB constraints required the population to become women taking reproductive hormones.

Randomization

Experiment 1 selected targets with calls to the PHP random_int() function, which returns cryptographically secure random numbers through the Linux getrandom(2) system call; the hosting server’s /dev/urandom “gathers environmental noise from device drivers and other sources into an entropy pool” and passes the National Institute of Standards and Technology tests of randomness. Experiment 2 is described as using “a true random-number generator drawing from network traffic (the same type as in experiment 1)”; the paper does not reconcile that phrase with the device-driver entropy description given for experiment 1. For the experiment 2 judging, transcript, target, and non-target images were presented in an order randomized via a hashed version of the participant ID, with the target placed equiprobably on the left or the right.

Sensory leakage

Both tasks are precognitive: the target did not exist as a selection until after the participant’s response was locked, so no ordinary sensory path to the target identity was available at response time. In experiment 2 the uploaded transcript could not be altered, and in the two cases where a crumpled or blurry image forced a re-upload, the originals were compared with the re-uploads and no discrepancies were found. In experiment 1, participants saw the descriptor graphs of both candidate targets before choosing, by design; the authors tested whether this content foreknowledge produced the interestingness effect by re-running the analysis on content-matched subsets (no text-only targets, no human or animal subjects, single-element targets only) and report the effect largely persisted, while conceding the single-element analysis was “especially underpowered” and that they “could not entirely put to rest” a number-of-elements explanation. Both experiments were run online and unsupervised, so compliance with the task procedure rests on the software’s enforced step sequence rather than observation.

Blinding

In experiment 1 the participant’s own forced choice was the response, scored automatically against the software’s later target selection, so no human judging was involved. In experiment 2 the two transcript judges did not know which image was the target, were not told the purpose of the experiment, were not given each other’s identities, and judged separately; in the later control judging they were not told that the true target was absent, and they were debriefed afterward. The interestingness raters in both experiments were Amazon Mechanical Turk workers who were not informed about the experiment or its hypotheses. The paper reports each researcher’s belief in a positive outcome on a 1-to-5 scale: the first author and third author at 5 in both experiments, the second author neutral at 3 in experiment 2. The first author also interacted with experiment 1 participants by email and newsletter, including a newsletter that told subscribers about the Profile A bias between batches.

Optional stopping

Experiment 1 had no fixed trial count by design: participants performed as many trials as they liked, and the two data downloads were cut at calendar dates chosen because they “corresponded relatively well to the data download times used in a related experiment published previously.” Experiment 2 set recruitment goals of 150 men, 125 women not taking hormones, and 30 women taking hormones in advance, and the paper states the experiment “ended just after we reached these goals due to difficulties replacing apparent bots with actual respondents,” closing at 151, 125, and 31. No other stopping rule is reported.

Outcome measure

Experiment 1 used proportion of correct forced choices against a 0.50 chance level. Experiment 2 used a two-judge agreement score per trial: 1 for agreed target matches (chance expectation .25), 0 for agreed non-target matches (.25), and −1 for disagreements (.50), which the authors note is more conservative than counting the proportion of judges selecting the actual target because it discounts transcripts that fail to gain judge agreement. Per-target accuracy scores were computed as sums rather than averages, deliberately weighting targets that were presented more often and scored more consistently. The unconditional love measure was reduced to the single question about love for the survey device after the other three items proved essentially collinear with the anxiety items; the paper reports the correlational structure that motivated this choice.

Effect size

The experiment 2 binomial comparisons carry effect sizes: h = .22 for agreed target matches, h = .36 for agreed non-target matches, and h = −.57 for judge disagreements, with h = .03, h = .34, and h = .35 for the corresponding control judging comparisons. The paper reports no standardized effect sizes for the experiment 1 accuracy, bias, sex-at-birth, or interestingness comparisons, and no confidence intervals for any estimate in either experiment.

Multiple comparisons

The authors state their policy plainly: “We thus made no attempts to correct for multiple comparisons during the exploratory phases or to make ourselves appear prescient in retrospect during the confirmation/replication phases.” The exploratory layer is large: per-batch accuracy by trial type, position bias, trial-number effects, age, sex-at-birth, target selection and presentation rates by sex, and interestingness comparisons run by two methods across four content subsets. The confirmatory layer consists of the pre-registered batch-2 interestingness analysis and the five experiment 2 predictions, reported without correction across the five.

Internal replication

The target interestingness effect appears three times across the report: exploratory in batch 1, confirmed in the pre-registered batch-2 analysis (t(48) = 5.03, p < .000008 across-responses), and confirmed again in experiment 2 as prediction 5 (r(84) = .21, p < .05). The unconditional love prediction confirmed a prior exploratory finding from the same research line. Other effects did not carry across: the batch-1 above-chance performance over participants’ first five trials did not recur in batch 2 (p > .52), the experiment 1 sex-at-birth accuracy difference had no counterpart in experiment 2 accuracy by gender, and the anxiety result ran against its prediction, with the authors calling the relation weak and therefore potentially spurious.

External replication

The free-response task was previously shown to produce large effect sizes in an earlier experiment by the first author’s team, and the unconditional love result confirms exploratory analyses from that study. The expectation-opposing pattern in experiment 1 is described as “a result apparently common for online forced-choice precognition tasks,” citing the earlier smartphone-based study. The sex and gender findings are framed as echoing gender effects reported in prior forced-choice precognition tasks by other groups. The paper reports no independent replication of these two specific experiments by other laboratories.

Transparency

The article is open access under a CC-BY license, names its funding sources and grant numbers, gives both IRB approval numbers, reports each researcher’s belief level and participant-contact role, quotes both questionnaires verbatim, and describes its bot-detection tests and the exclusion of 76 survey-inattentive participants with the reasoning for each. It discloses the failure to remove developer and experimenter data from the original batch-1 analysis and the resulting mismatch with the pre-registration document. Raw data for experiment 1 “are available upon request,” and target-set figures are linked from the text; no code or data repository is named. The title footnote directs readers to a Letter to the Editor in the same issue without describing its contents.

The adversarial record

Sources
  1. Mossbridge, J., Cameron, K., & Boccuzzi, M. (2024). State, Trait, and Target Parameters Associated with Accuracy in Two Online Tests of Precognitive Remote Viewing. Journal of Anomalous Experience and Cognition, 4(1), 88–121. https://doi.org/10.31156/jaex.24743 R001 [Mossbridge 2024] ↩︎
  2. Mossbridge, J., & Radin, D. (2021). Psi performance as a function of demographic and personality factors in smartphone-based tests: Using a “SEARCH” approach. Journal of Anomalous Experience and Cognition, 1(1-2), 78–113. https://doi.org/10.31156/jaex.23419 R002 [Mossbridge & Radin 2021] ↩︎
  3. Mossbridge, J. A., Nisam, M., & Crabtree, A. (2021). Can hypnotic suggestion induce feelings of unconditional love and supernormal performance? Spirituality in Clinical Practice, 8(1), 30–50. https://doi.org/10.1037/scp0000239 R003 [Mossbridge et al. 2021b] ↩︎
  4. Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Lowe, B. (2006). A brief measure for assessing generalized anxiety disorder. Archives of Internal Medicine, 166, 1092–1097. https://doi.org/10.1001/archinte.166.10.1092 R004 [Spitzer 2006] ↩︎