Every number, and the run that produced it.

This page is generated from docs/proof/claims.toml and the evidence files it points at. It is not written by hand. The same generator runs in the project's health check, so a number that drifts from its evidence fails the build rather than reaching a reader.

Two claims are deliberately unpublished and say so in the ledger, and several are marked read by a human: they are historical or rounded statements that no automatic check can settle, and pretending otherwise would be a false green.
ClaimValueEvidence
Machine time for one deployed run of the whole lifecycle
machine time only; the human approval is timed separately and excluded
55.3 s of machine timespikes/judge_run/evidence.json
The single human approval the same run required47.7 s human approvalspikes/judge_run/evidence.json
The budget one deployed run must finish inside130sspikes/judge_run/evidence.json
The same work done by hand, timed
author-timed, not practitioner-reviewed
663.5 sspikes/manual_baseline/evidence.json
Steps in the hand-done walkthrough
author-timed, not practitioner-reviewed
20 stepsspikes/manual_baseline/evidence.json
Steps of that walkthrough the run removes
measured against an author-timed baseline, not a practitioner-reviewed one
19 of which the run removesspikes/judge_run/evidence.json
Business days the two cases span
the lifecycle clock is compressed, and the compression is disclosed on screen
380 simulated business daysspikes/judge_run/evidence.json
Contracts re-verified at capture timenine contractsspikes/core_contracts/evidence.json
Graded domain cases, all passing
a deterministic pass metric; re-run the suite before citing it
24/24spikes/domain_evals/evidence.json
Mean score across the graded domain suite
a deterministic pass metric; re-run the suite before citing it
100%spikes/domain_evals/evidence.json
Fields entered into the ERP without a human retyping them22 fieldsspikes/judge_run/evidence.json
Days a non-compliant supplier was actually held from purchasing
claimed only where both the hold and its release executed in the ERP
5 enforced hold daysspikes/judge_run/evidence.json
Decisions policy requires a human to make1 policy-required interventionspikes/judge_run/evidence.json
Duplicate ERP writes after a retry0 duplicate writes after a retryspikes/judge_run/evidence.json
Contract and unit tests re-executed by the run that reports them
re-executed by the harness, not quoted from a previous run
549 passedspikes/judge_run/evidence.json
Gross running cost for the whole project, month to date
before credits, measured from the billing account rather than from published rates
$16.63spikes/cost_posture/evidence.json
Cloud credit remaining
console-only; no API exposes the balance, so it is passed in when the measurement is taken
$138.09spikes/cost_posture/evidence.json
Measured daily uptime of the screening VM5.6 h/dayspikes/cost_posture/evidence.json
The ERP hold was written minutes before the rejection committedread the evidence filespikes/hitl_reject/evidence.json
The graded suite passed 8/8 when the generation pin was measuredread the evidence filespikes/domain_evals/evidence.json
Reasoning tokens fell to effectively none and the timed sequence shortenedread the evidence filespikes/gemini_37_eval/evidence.json
The build piece states the graded suite passed 8/8 with reasoning pinned offread the evidence filespikes/domain_evals/evidence.json
The build piece states a two-supplier run dropped from 85 to 57 secondsread the evidence filespikes/gemini_37_eval/evidence.json