How Studies Are Audited
Every study-audit page on ESP-Nexus is graded independently by four AI systems — Anthropic’s Claude, OpenAI’s GPT-5, Perplexity, and Google’s Gemini — against the single fixed checklist on this page. A study page is published as audited only when all four return their top grade (A+) with zero source-fidelity mismatches, and that page is then locked from further change. The audit checks one thing above all: that every number, citation, and claim on the page matches the study’s primary source, and that the authors’ own stated limitations are reported. It takes no position on whether psi is real.
The four-vendor audit
Each study-audit page is sent, as published HTML, to four independent large-language-model auditors, together with the verbatim text of the study’s primary source. Each auditor grades the page against the closed checklist below and returns a structured verdict. A page reaches its locked state only on a unanimous A+ across all four vendors with no fidelity mismatch; the per-page grades and the date of the audit are recorded in the footer of each study page. The checklist is deliberately closed — the auditors may not introduce criteria beyond the seven items — because an open-ended “what could be better?” rubric produces grades that drift run-to-run on an unchanged page.
The audit prompt
Current version: study-audit-v2.0 — locked 2026-06-05. This is the exact instrument given to each of the four auditors. The study’s citation, its DOI, and the verbatim text of its primary source are inserted where the {...} placeholders appear.
AUDIT_PROMPT_VERSION: study-audit-v2.0 | locked: 2026-06-05
# External audit (v2, closed checklist) — ESP-Nexus STUDY-AUDIT page: {short}
You are an independent external auditor. ESP-Nexus (esp-nexus.org) is a curated parapsychology reference site whose core commitment is that **every figure on a page is verified against its primary source** and that **methodological limitations are reported honestly**. The site takes no position on whether psi is real.
The page under audit (full raw HTML embedded below via --fetch-pages) is a **Study Audit**: it takes ONE published study and reports what the paper claims, what it shows, and where the two diverge. The verbatim primary article text is also embedded below; treat it as ground truth.
## The study under audit
{citation}
- **DOI:** {doi}
## How to grade — READ THIS CAREFULLY
You grade the page ONLY against the seven binary checklist items below. **Do not introduce any criterion beyond this list.** Each item is PASS or FAIL.
- **C1 - Fidelity.** Every statistic, sample size, p-value, effect size, date, journal, volume, page range, DOI, and author name on the page matches the embedded primary. (FAIL only with a specific `page_value -> correct_value` mismatch.)
- **C2 - Headline accuracy.** The page's headline conclusion reflects the paper's actual primary finding, neither overstated nor understated.
- **C3 - Confirmatory vs exploratory.** The paper's pre-specified primary outcome is labelled distinctly from any secondary, post-hoc, or exploratory analysis, consistent with how the paper itself frames them.
- **C4 - Citation integrity.** Every reference is a real work with the correct journal/venue and DOI; no fabricated, wrong-journal, or wrong-author citations.
- **C5 - Neutral register.** No advocacy for, and no debunking of, the effect; the page takes no position on whether psi is real.
- **C6 - Authors' own limitations present.** Every limitation the PAPER'S OWN AUTHORS explicitly acknowledge is represented on the page. (Bounded: a limitation counts as missing ONLY if the authors themselves raise it and the page omits it. Do NOT fail this for secondary sub-analyses or caveats the paper treats as minor, or for analyses you personally would have added.)
- **C7 - No contradiction.** No statement in the eleven-dimension audit contradicts the primary.
### Grade mapping (binding)
- **A+** = all seven items PASS and there are zero fidelity mismatches. **If that holds you MUST return A+.** Do NOT withhold A+ for stylistic preferences, wording you would phrase differently, or additional analyses you would have included.
- **A / A-** = all items pass but you are noting borderline judgement on C2/C3/C6 wording (use sparingly; if you cannot name a failing item, the grade is A+).
- **B+ or lower** = one or more checklist items FAIL; you must name which item and why.
There is no "completeness" score beyond C6, and C6 is bounded to the authors' own stated limitations. Anything you would *like* added that is not a checklist failure goes in `optional_suggestions`, which **MUST NOT affect the grade**.
## Output format (return ONLY this JSON)
```json
{
"overall_grade": "<A+|A|A-|B+|B|B-|C+|C|D|F>",
"checklist": {"C1_fidelity":"pass|fail","C2_headline":"pass|fail","C3_confirmatory_exploratory":"pass|fail","C4_citations":"pass|fail","C5_register":"pass|fail","C6_authors_limitations":"pass|fail","C7_no_contradiction":"pass|fail"},
"fidelity_mismatches": ["<page_value -> correct_value (location)>"],
"failed_items": ["<C#: specific reason>"],
"optional_suggestions": ["<non-grading nice-to-have; does NOT lower the grade>"],
"synthesis": "<120 words max: pass/fail summary, citing any failing item>"
}
```
---
## PRIMARY SOURCE (verbatim extracted text; ground truth)
[ The verbatim text of the study’s primary-source PDF is inserted here, per page. ]
Version history
The audit prompt is versioned and dated. When it is revised, a new row is added below with the new version number, its date, and a summary of what changed; earlier versions remain on record, and any correction that supersedes an earlier version is noted in its row. The version under which each study page was graded is shown in that page’s footer.
| Version | Date | Notes |
|---|---|---|
| study-audit-v2.0 | 2026-06-05 | First published version. A closed, binary seven-item checklist (C1 fidelity, C2 headline accuracy, C3 confirmatory-vs-exploratory, C4 citation integrity, C5 neutral register, C6 authors’ own limitations, C7 no contradiction) with a binding rule — all seven pass and zero fidelity mismatches means the auditor must return A+ — plus a separate, non-grading “optional suggestions” field. This replaced an earlier open-ended rubric whose grades drifted between A and A+ on an unchanged page. |