How accurate is FAKTS?
Here are the receipts.
We run FAKTS against a fixed set of claims with known answers and publish the score. No cherry-picking, the full test set is below. Anyone can inspect it. Anyone can try to break it.
Every claim. Expected vs got.
0 active benchmark claims. This is the whole test set.
How the number is computed.
The harness. Every benchmark claim runs through the same FAKTS harness the public product uses: an independent panel, an adversarial round when the panel is not unanimous, calibrated consensus weighted by web-grounding, per-domain reliability, and each model's confidence, and a source-grounded self-check.
Catch rate. Of the claims whose expected answer is false or misleading, the share that FAKTS flagged as anything other than verified. This is the number that actually matters for the "does it stop hallucinations" question.
Overall accuracy. Share of all benchmark claims where the final verdict matched the expected verdict.
Per-domain and per-model. Same comparison, sliced by claim domain and by individual model verdict. Per-model reliability feeds back into the calibration weights, so the next run of the harness weights each model by how it actually performed here.
Caveats. This is a fixed public test set, not a claim of universal accuracy. Verdicts reflect the sources available to the models at run time. Some claims sit on genuinely open questions and will move over time.
Found something it misses?
Submit a claim you think FAKTS will get wrong. Verified misses get added to the test set, and credited.