Naive compiler baseline → rocq-tools, mean over the three difficulty bins (dev set). Accuracy should rise; cost per solved proof and time per attempt should drop. The interface helps most where the model is weakest, and still pays at the frontier tier.
| policy | accuracy (mean pass@1) | cost ($/solved proof) | time (s/attempt) |
|---|---|---|---|
| haiku (weak) | 0.33 → 0.55 ▲66% | $0.37 → $0.16 ▼57% | 123s → 68s ▼45% |
| sonnet (strong) | 0.89 → 0.93 ▲5% | $0.14 → $0.09 ▼33% | 83s → 54s ▼35% |
| fable (frontier) | 0.97 → 0.98 ▲2% | $0.21 → $0.20 ▼3% | 54s → 49s ▼11% |
Each step was A/B-tested against its predecessor on the same 60 problems (20 per difficulty bucket), same weak policy (claude-haiku-4-5), ≥2 repetitions; kept only if the per-bucket numbers improved. Lines show pass@1 per bucket; reverted steps are part of the record but not of the shipped tool. Hover any point for exact values.
| step | what changed | agent-facing tools | verdict | measured outcome | easy | medium | hard |
|---|---|---|---|---|---|---|---|
| 0 naive | whole-file compile per call (control) | check | control | the deliberately-naive starting point | 0.44 | 0.25 | 0.30 |
| 1 session | persistent prover; sentence steps, O(1) undo | steprollbackstate | KEPT | medium/hard solve +30 %, output tokens −80 % | 0.47 | 0.33 | 0.33 |
| 2 +try | k candidate tactics tested in one call | steprollbackstatetry | KEPT | easy +37 %, medium/hard +15 % | 0.65 | 0.38 | 0.38 |
| 3 compact | token-thrifty goal rendering | steprollbackstatetry | REVERTED | −2.5 pp everywhere — agents re-fetch what you elide | 0.62 | 0.35 | 0.35 |
| 4 search | a Search tool the agent can call | steprollbackstatetrysearch | REVERTED | heavily used, rescued nothing — pull wastes turns | 0.60 | 0.35 | 0.40 |
| 5 +hints | errors carry Lean→Rocq rewrite hints | steprollbackstatetry | KEPT | medium +27 % | 0.60 | 0.47 | 0.38 |
| 6 +auto_close | server-side finisher portfolio | steprollbackstatetryauto_close | KEPT | every bucket up (+8/+16/+13 %) | 0.65 | 0.55 | 0.42 |
| 7 +did-you-mean | unknown names get near-miss suggestions | steprollbackstatetryauto_close | KEPT | easy +8 %, medium +9 % — push beats pull | 0.70 | 0.53 | 0.42 |
| 8 draft-first | whole-proof-first prompt style at haiku | steprollbackstatetryauto_close | REVERTED | far below the incremental winner at this policy | 0.40 | 0.17 | 0.23 |
| 9a fix | false-winner bug fix (A22) | steprollbackstatetryauto_close | KEPT | numbers unchanged; correctness fix | 0.65 | 0.45 | 0.45 |
| 9b synthesis | server synthesizes goal-specific hint terms | steprollbackstatetryauto_close | KEPT | medium +28 % — new haiku best | 0.65 | 0.57 | 0.47 |
| 10 universal | style-agnostic check + neutral prompt + atlas fixes | checksteprollbackstatetryauto_close | RECOMMENDED | best worst-case across BOTH policies (A24) | 0.66 | 0.57 | 0.40 |
The experiment's other two objectives. Each chart divides the per-attempt mean by pass@1: the expected cost (or wall-clock, or model-output tokens) to obtain ONE solved proof, failed attempts included. Same ladder steps and buckets as above — solve rate rose while $ and time per proof fell. Hover for values; hard-bucket spikes at early steps reflect near-zero solve rates there.
| step | easy $ | medium $ | hard $ | easy wall | medium wall | hard wall | easy tok | medium tok | hard tok |
|---|---|---|---|---|---|---|---|---|---|
| 0 naive | 0.078 | 0.111 | 0.148 | 90 | 122 | 157 | 7.3k | 12.1k | 16.1k |
| 1 session | 0.050 | 0.062 | 0.062 | 51 | 47 | 49 | 1.8k | 2.5k | 2.9k |
| 2 +try | 0.045 | 0.062 | 0.060 | 43 | 51 | 49 | 1.7k | 3.0k | 2.9k |
| 3 compact | 0.047 | 0.063 | 0.061 | 50 | 52 | 54 | 1.7k | 2.9k | 3.2k |
| 4 search | 0.042 | 0.059 | 0.057 | 42 | 53 | 52 | 1.7k | 2.7k | 3.0k |
| 5 +hints | 0.047 | 0.061 | 0.067 | 50 | 45 | 53 | 1.9k | 2.7k | 3.4k |
| 6 +auto_close | 0.049 | 0.056 | 0.063 | 51 | 48 | 56 | 1.7k | 2.5k | 3.3k |
| 7 +did-you-mean | 0.043 | 0.058 | 0.060 | 46 | 45 | 47 | 1.6k | 2.6k | 2.8k |
| 8 draft-first | 0.063 | 0.088 | 0.090 | 67 | 85 | 84 | 4.3k | 8.6k | 8.7k |
| 9a fix | 0.043 | 0.061 | 0.060 | 44 | 46 | 48 | 1.6k | 2.8k | 3.0k |
| 9b synthesis | 0.053 | 0.054 | 0.059 | 49 | 42 | 48 | 1.8k | 2.5k | 3.1k |
| 10 universal | 0.067 | 0.089 | 0.090 | 59 | 71 | 74 | 2.7k | 4.5k | 5.1k |
One server, one neutral prompt, measured at both a weak and a strong policy. The universal config beats the naive interface in every bucket at sonnet (first substrate config to do so) and is best-or-tied at haiku except hard, where the haiku-tuned variant keeps a .075 edge (within ½σ). pass@1, dev60.
accuracy — pass@1
cost — expected $ per solved proof (failures included)
latency — mean wall-clock seconds per attempt
| config | easy $/solve | medium $/solve | hard $/solve | easy wall | medium wall | hard wall |
|---|---|---|---|---|---|---|
| naive @ haiku | $0.18 | $0.44 | $0.49 | 90s | 122s | 157s |
| universal @ haiku | $0.10 | $0.16 | $0.23 | 59s | 71s | 74s |
| naive @ sonnet | $0.09 | $0.15 | $0.18 | 61s | 77s | 112s |
| universal @ sonnet | $0.07 | $0.09 | $0.13 | 36s | 42s | 85s |
github.com/LLM4Rocq/rocq-mcp, measured under the identical harness, gate, problems, and policies (first contact only after our design freeze; REPORT §SOTA). Bars: pass@1. Table: all three dimensions — $/solve = expected cost per solved proof (failures included), wall = mean s/attempt. Numbers from the FAIR rerun (Jul 8): an audit found the original runs started with rocq-mcp's server still connecting in 97-98 % of attempts (our integration artifact — retracted); an instant-handshake proxy fixed it, 120/120 connected. Fair verdict: near accuracy-parity at sonnet, rocq-mcp edges easy at haiku; rocq-tools leads weak-policy medium/hard and costs about half per solved proof at sonnet. See REPORT for the full audit story.
| config | easy pass@1 | medium pass@1 | hard pass@1 | easy $/solve | medium $/solve | hard $/solve | easy wall | medium wall | hard wall |
|---|---|---|---|---|---|---|---|---|---|
| rocq-mcp @ haiku | 0.68 | 0.35 | 0.33 | $0.08 | $0.22 | $0.31 | 65s | 76s | 97s |
| naive @ haiku | 0.44 | 0.25 | 0.30 | $0.18 | $0.44 | $0.49 | 90s | 122s | 157s |
| universal @ haiku | 0.66 | 0.57 | 0.40 | $0.10 | $0.16 | $0.23 | 59s | 71s | 74s |
| rocq-mcp @ sonnet | 0.95 | 0.93 | 0.80 | $0.11 | $0.19 | $0.26 | 50s | 64s | 105s |
| naive @ sonnet | 0.93 | 0.95 | 0.80 | $0.09 | $0.15 | $0.18 | 61s | 77s | 112s |
| universal @ sonnet | 0.95 | 1.00 | 0.85 | $0.07 | $0.09 | $0.13 | 36s | 42s | 85s |
Each point is one full run of a fixed 24-problem stratified batch (8 per bucket) executed with N = 1, 2, 4, 8 concurrent agents — eight runs total (4 per interface), all in a single night window (Jul 3–4) after two earlier measurement artifacts were caught and disclosed (REPORT §5). Both interfaces scale healthily (flat wall per attempt, ~72–80 % parallel efficiency at N=8); the winner's advantage is its ~2.7× per-attempt speed, compounding to ≈6× solved-proofs-per-hour at N=8. The local prover is never the bottleneck (CPU ≤ 24 % of a 14-core laptop).
| interface | N agents | attempts/h | wall s/attempt | solved | peak RSS MB* | CPU % |
|---|---|---|---|---|---|---|
| naive baseline | 1 | 28.1 | 128.0 | 3/24 | 1140.9 | 4.6 |
| naive baseline | 2 | 50.5 | 140.0 | 4/24 | 1487.3 | 4.4 |
| naive baseline | 4 | 91.5 | 135.1 | 3/24 | 4929.2 | 13.3 |
| naive baseline | 8 | 161.3 | 133.4 | 4/24 | 4025.2 | 12.3 |
| session winner | 1 | 74.7 | 47.9 | 11/24 | 962.7 | 5.3 |
| session winner | 2 | 142.6 | 49.7 | 10/24 | 1720.0 | 8.4 |
| session winner | 4 | 272.3 | 49.7 | 9/24 | 3319.8 | 16.8 |
| session winner | 8 | 478.3 | 51.3 | 8/24 | 6391.7 | 23.8 |
Cross-policy annexes, SOTA comparison, team experiments,
in-project probes, held-out — each analyzed in docs/REPORT.md. Raw per-run
numbers below; reproduce any row with python3 harness/report.py
<run>.
| run | policy | attempts | pass@1 e/m/h |
|---|---|---|---|
| smoke_baseline_1 | claude-haiku-4-5 | 5 | 0.80 / – / – |
| smoke_session_1 | claude-haiku-4-5 | 5 | 0.80 / – / – |
| session_try_dev150 | claude-haiku-4-5 | 300 | 0.45 / 0.33 / 0.35 |
| session_try_hints_minif2f_valid | claude-haiku-4-5 | 488 | 0.32 / 0.13 / 0.00 |
| session_try_hints_v2_minif2f_valid | claude-haiku-4-5 | 488 | 0.57 / 0.30 / 0.06 |
| sweep_session_try_hints_auto_N1 | claude-haiku-4-5 | 24 | 0.38 / 0.62 / 0.38 |
| sweep_session_try_hints_auto_N2 | claude-haiku-4-5 | 24 | 0.38 / 0.50 / 0.38 |
| sweep_session_try_hints_auto_N4 | claude-haiku-4-5 | 24 | 0.25 / 0.62 / 0.25 |
| sweep_session_try_hints_auto_N8 | claude-haiku-4-5 | 24 | 0.25 / 0.50 / 0.25 |
| baseline_sonnet_dev60 | claude-sonnet-5 | 120 | 0.93 / 0.95 / 0.80 |
| session_try_hints_auto_sonnet_dev60 | claude-sonnet-5 | 120 | 0.93 / 0.82 / 0.70 |
| solo_hard70 | claude-haiku-4-5 | 140 | – / – / 0.40 |
| team_k3_hard70 | claude-haiku-4-5 | 140 | – / – / 0.32 |
| unified_sonnet_dev60 | claude-sonnet-5 | 120 | 0.93 / 0.85 / 0.75 |
| FINAL_minif2f_test | claude-haiku-4-5 | 488 | 0.52 / 0.13 / 0.04 |
| sonnet_native_dev60 | claude-sonnet-5 | 120 | 0.95 / 0.85 / 0.75 |
| rocq_mcp_smoke | claude-haiku-4-5 | 3 | 0.33 / – / – |
| rocq_mcp_dev60 | claude-haiku-4-5 | 120 | 0.45 / 0.23 / 0.23 |
| winner_ctx_full_inproject60 | claude-haiku-4-5 | 120 | – / 0.72 / – |
| winner_ctx_lean_inproject60 | claude-haiku-4-5 | 120 | – / 0.60 / – |
| rocq_mcp_sonnet_dev60 | claude-sonnet-5 | 120 | 0.82 / 0.72 / 0.72 |
| sonnet_native_auto2_dev60 | claude-sonnet-5 | 60 | 1.00 / 0.85 / 0.85 |
| solo_decomposable | claude-haiku-4-5 | 54 | 0.67 / 0.75 / 0.59 |
| team_decomposable | claude-haiku-4-5 | 54 | 0.33 / 0.25 / 0.18 |
| winner_ctx_lean_mathcomp | claude-haiku-4-5 | 35 | – / 0.07 / – |
| sweep_baseline_N1 | claude-haiku-4-5 | 24 | 0.12 / 0.00 / 0.25 |
| sweep_baseline_N2 | claude-haiku-4-5 | 24 | 0.12 / 0.12 / 0.25 |
| sweep_baseline_N4 | claude-haiku-4-5 | 24 | 0.12 / 0.00 / 0.25 |
| sweep_baseline_N8 | claude-haiku-4-5 | 24 | 0.12 / 0.25 / 0.12 |
| universal_sonnet_dev60 | claude-sonnet-5 | 120 | 0.95 / 1.00 / 0.85 |
| winner_ctx_lean_ssr_mathcomp | claude-haiku-4-5 | 35 | – / 0.07 / – |
| universal_fable_dev60 | claude-fable-5 | 60 | 0.95 / 1.00 / 1.00 |
| baseline_fable_dev60 | claude-fable-5 | 60 | 0.95 / 1.00 / 0.95 |
| ctx_full_fable_mathcomp | claude-fable-5 | 35 | – / 0.93 / – |
| ctx_lean_fable_mathcomp | claude-fable-5 | 35 | – / 0.93 / – |
| winner_ctx_lean_ex_mathcomp | claude-haiku-4-5 | 35 | – / 0.07 / – |
| winner_ctx_lean_mc_mathcomp | claude-haiku-4-5 | 35 | – / 0.00 / – |
| mcp_timeout_probe | claude-haiku-4-5 | 1 | 0.00 / – / – |
| rocq_mcp_fair_dev60 | claude-haiku-4-5 | 120 | 0.68 / 0.35 / 0.33 |
| rocq_mcp_fair_sonnet_dev60 | claude-sonnet-5 | 120 | 0.95 / 0.93 / 0.80 |