AI-native Rocq tooling — final results

experiment complete (Jun 30 – Jul 7, 2026) · generated 2026-07-08 14:44:54 · full analysis: docs/REPORT.md · per-decision rationale: docs/DESIGN.md · repo README for install & try

The three objectives, across three policies

Naive compiler baseline → rocq-tools, mean over the three difficulty bins (dev set). Accuracy should rise; cost per solved proof and time per attempt should drop. The interface helps most where the model is weakest, and still pays at the frontier tier.

policyaccuracy (mean pass@1)cost ($/solved proof)time (s/attempt)
haiku (weak)0.33 → 0.55 ▲66%$0.37 → $0.16 ▼57%123s → 68s ▼45%
sonnet (strong)0.89 → 0.93 ▲5%$0.14 → $0.09 ▼33%83s → 54s ▼35%
fable (frontier)0.97 → 0.98 ▲2%$0.21 → $0.20 ▼3%54s → 49s ▼11%

The ladder — one measured change at a time

Each step was A/B-tested against its predecessor on the same 60 problems (20 per difficulty bucket), same weak policy (claude-haiku-4-5), ≥2 repetitions; kept only if the per-bucket numbers improved. Lines show pass@1 per bucket; reverted steps are part of the record but not of the shipped tool. Hover any point for exact values.

easymediumhard
00.250.50.7510naive1session2+try3compactreverted4searchreverted5+hints6+auto_close7+did-you-mea8draft-firstreverted9afix9bsynthesis10universal0 naive · easy: pass@1 0.4381 session · easy: pass@1 0.4752 +try · easy: pass@1 0.6503 compact · easy: pass@1 0.6254 search · easy: pass@1 0.6005 +hints · easy: pass@1 0.6006 +auto_close · easy: pass@1 0.6507 +did-you-mean · easy: pass@1 0.7008 draft-first · easy: pass@1 0.4009a fix · easy: pass@1 0.6509b synthesis · easy: pass@1 0.65010 universal · easy: pass@1 0.6620.660 naive · medium: pass@1 0.2501 session · medium: pass@1 0.3252 +try · medium: pass@1 0.3753 compact · medium: pass@1 0.3504 search · medium: pass@1 0.3505 +hints · medium: pass@1 0.4756 +auto_close · medium: pass@1 0.5507 +did-you-mean · medium: pass@1 0.5258 draft-first · medium: pass@1 0.1759a fix · medium: pass@1 0.4509b synthesis · medium: pass@1 0.57510 universal · medium: pass@1 0.5750.570 naive · hard: pass@1 0.3001 session · hard: pass@1 0.3252 +try · hard: pass@1 0.3753 compact · hard: pass@1 0.3504 search · hard: pass@1 0.4005 +hints · hard: pass@1 0.3756 +auto_close · hard: pass@1 0.4257 +did-you-mean · hard: pass@1 0.4258 draft-first · hard: pass@1 0.2259a fix · hard: pass@1 0.4509b synthesis · hard: pass@1 0.47510 universal · hard: pass@1 0.4000.40
step-by-step table (tools available to the agent at each step)
stepwhat changedagent-facing toolsverdictmeasured outcomeeasymediumhard
0 naivewhole-file compile per call (control)checkcontrolthe deliberately-naive starting point0.440.250.30
1 sessionpersistent prover; sentence steps, O(1) undosteprollbackstateKEPTmedium/hard solve +30 %, output tokens −80 %0.470.330.33
2 +tryk candidate tactics tested in one callsteprollbackstatetryKEPTeasy +37 %, medium/hard +15 %0.650.380.38
3 compacttoken-thrifty goal renderingsteprollbackstatetryREVERTED−2.5 pp everywhere — agents re-fetch what you elide0.620.350.35
4 searcha Search tool the agent can callsteprollbackstatetrysearchREVERTEDheavily used, rescued nothing — pull wastes turns0.600.350.40
5 +hintserrors carry Lean→Rocq rewrite hintssteprollbackstatetryKEPTmedium +27 %0.600.470.38
6 +auto_closeserver-side finisher portfoliosteprollbackstatetryauto_closeKEPTevery bucket up (+8/+16/+13 %)0.650.550.42
7 +did-you-meanunknown names get near-miss suggestionssteprollbackstatetryauto_closeKEPTeasy +8 %, medium +9 % — push beats pull0.700.530.42
8 draft-firstwhole-proof-first prompt style at haikusteprollbackstatetryauto_closeREVERTEDfar below the incremental winner at this policy0.400.170.23
9a fixfalse-winner bug fix (A22)steprollbackstatetryauto_closeKEPTnumbers unchanged; correctness fix0.650.450.45
9b synthesisserver synthesizes goal-specific hint termssteprollbackstatetryauto_closeKEPTmedium +28 % — new haiku best0.650.570.47
10 universalstyle-agnostic check + neutral prompt + atlas fixeschecksteprollbackstatetryauto_closeRECOMMENDEDbest worst-case across BOTH policies (A24)0.660.570.40

Cost and time — expected spend per solved proof

The experiment's other two objectives. Each chart divides the per-attempt mean by pass@1: the expected cost (or wall-clock, or model-output tokens) to obtain ONE solved proof, failed attempts included. Same ladder steps and buckets as above — solve rate rose while $ and time per proof fell. Hover for values; hard-bucket spikes at early steps reflect near-zero solve rates there.

easymediumhard
cost $ / solved proof$0.00$0.25$0.500123456789a9b100 naive · easy: $0.181 session · easy: $0.102 +try · easy: $0.073 compact · easy: $0.084 search · easy: $0.075 +hints · easy: $0.086 +auto_close · easy: $0.087 +did-you-mean · easy: $0.068 draft-first · easy: $0.169a fix · easy: $0.079b synthesis · easy: $0.0810 universal · easy: $0.100 naive · medium: $0.441 session · medium: $0.192 +try · medium: $0.173 compact · medium: $0.184 search · medium: $0.175 +hints · medium: $0.136 +auto_close · medium: $0.107 +did-you-mean · medium: $0.118 draft-first · medium: $0.509a fix · medium: $0.139b synthesis · medium: $0.0910 universal · medium: $0.160 naive · hard: $0.491 session · hard: $0.192 +try · hard: $0.163 compact · hard: $0.174 search · hard: $0.145 +hints · hard: $0.186 +auto_close · hard: $0.157 +did-you-mean · hard: $0.148 draft-first · hard: $0.409a fix · hard: $0.139b synthesis · hard: $0.1210 universal · hard: $0.23wall s / solved proof02625240123456789a9b100 naive · easy: 2071 session · easy: 1082 +try · easy: 673 compact · easy: 804 search · easy: 705 +hints · easy: 836 +auto_close · easy: 787 +did-you-mean · easy: 658 draft-first · easy: 1689a fix · easy: 689b synthesis · easy: 7610 universal · easy: 890 naive · medium: 4871 session · medium: 1442 +try · medium: 1373 compact · medium: 1504 search · medium: 1525 +hints · medium: 956 +auto_close · medium: 877 +did-you-mean · medium: 868 draft-first · medium: 4889a fix · medium: 1029b synthesis · medium: 7310 universal · medium: 1240 naive · hard: 5241 session · hard: 1512 +try · hard: 1313 compact · hard: 1554 search · hard: 1315 +hints · hard: 1426 +auto_close · hard: 1337 +did-you-mean · hard: 1118 draft-first · hard: 3729a fix · hard: 1079b synthesis · hard: 10110 universal · hard: 184output tokens / solved proof0k27k54k0123456789a9b100 naive · easy: 17k1 session · easy: 4k2 +try · easy: 3k3 compact · easy: 3k4 search · easy: 3k5 +hints · easy: 3k6 +auto_close · easy: 3k7 +did-you-mean · easy: 2k8 draft-first · easy: 11k9a fix · easy: 2k9b synthesis · easy: 3k10 universal · easy: 4k0 naive · medium: 48k1 session · medium: 8k2 +try · medium: 8k3 compact · medium: 8k4 search · medium: 8k5 +hints · medium: 6k6 +auto_close · medium: 5k7 +did-you-mean · medium: 5k8 draft-first · medium: 49k9a fix · medium: 6k9b synthesis · medium: 4k10 universal · medium: 8k0 naive · hard: 54k1 session · hard: 9k2 +try · hard: 8k3 compact · hard: 9k4 search · hard: 7k5 +hints · hard: 9k6 +auto_close · hard: 8k7 +did-you-mean · hard: 7k8 draft-first · hard: 39k9a fix · hard: 7k9b synthesis · hard: 7k10 universal · hard: 13k
per-attempt table (cost $, wall s, tokens out — mean per attempt)
stepeasy $medium $hard $easy wallmedium wallhard walleasy tokmedium tokhard tok
0 naive0.0780.1110.148901221577.3k12.1k16.1k
1 session0.0500.0620.0625147491.8k2.5k2.9k
2 +try0.0450.0620.0604351491.7k3.0k2.9k
3 compact0.0470.0630.0615052541.7k2.9k3.2k
4 search0.0420.0590.0574253521.7k2.7k3.0k
5 +hints0.0470.0610.0675045531.9k2.7k3.4k
6 +auto_close0.0490.0560.0635148561.7k2.5k3.3k
7 +did-you-mean0.0430.0580.0604645471.6k2.6k2.8k
8 draft-first0.0630.0880.0906785844.3k8.6k8.7k
9a fix0.0430.0610.0604446481.6k2.8k3.0k
9b synthesis0.0530.0540.0594942481.8k2.5k3.1k
10 universal0.0670.0890.0905971742.7k4.5k5.1k

Policy-neutrality — the headline result (A24)

One server, one neutral prompt, measured at both a weak and a strong policy. The universal config beats the naive interface in every bucket at sonnet (first substrate config to do so) and is best-or-tied at haiku except hard, where the haiku-tuned variant keeps a .075 edge (within ½σ). pass@1, dev60.

naive whole-file interfaceuniversal (recommended)

accuracy — pass@1

claude-haiku-4-5 (weak policy)easyclaude-haiku-4-5 (weak policy) · naive · easy: 0.4380.44claude-haiku-4-5 (weak policy) · universal · easy: 0.6620.66mediumclaude-haiku-4-5 (weak policy) · naive · medium: 0.2500.25claude-haiku-4-5 (weak policy) · universal · medium: 0.5750.57hardclaude-haiku-4-5 (weak policy) · naive · hard: 0.3000.30claude-haiku-4-5 (weak policy) · universal · hard: 0.4000.40claude-sonnet-5 (strong policy)easyclaude-sonnet-5 (strong policy) · naive · easy: 0.9250.93claude-sonnet-5 (strong policy) · universal · easy: 0.9500.95mediumclaude-sonnet-5 (strong policy) · naive · medium: 0.9500.95claude-sonnet-5 (strong policy) · universal · medium: 1.0001.00hardclaude-sonnet-5 (strong policy) · naive · hard: 0.8000.80claude-sonnet-5 (strong policy) · universal · hard: 0.8500.85

cost — expected $ per solved proof (failures included)

claude-haiku-4-5easyclaude-haiku-4-5 · naive · easy: $0.18$0.18claude-haiku-4-5 · universal · easy: $0.10$0.10mediumclaude-haiku-4-5 · naive · medium: $0.44$0.44claude-haiku-4-5 · universal · medium: $0.16$0.16hardclaude-haiku-4-5 · naive · hard: $0.49$0.49claude-haiku-4-5 · universal · hard: $0.23$0.23claude-sonnet-5easyclaude-sonnet-5 · naive · easy: $0.09$0.09claude-sonnet-5 · universal · easy: $0.07$0.07mediumclaude-sonnet-5 · naive · medium: $0.15$0.15claude-sonnet-5 · universal · medium: $0.09$0.09hardclaude-sonnet-5 · naive · hard: $0.18$0.18claude-sonnet-5 · universal · hard: $0.13$0.13

latency — mean wall-clock seconds per attempt

claude-haiku-4-5easyclaude-haiku-4-5 · naive · easy: 90s90sclaude-haiku-4-5 · universal · easy: 59s59smediumclaude-haiku-4-5 · naive · medium: 122s122sclaude-haiku-4-5 · universal · medium: 71s71shardclaude-haiku-4-5 · naive · hard: 157s157sclaude-haiku-4-5 · universal · hard: 74s74sclaude-sonnet-5easyclaude-sonnet-5 · naive · easy: 61s61sclaude-sonnet-5 · universal · easy: 36s36smediumclaude-sonnet-5 · naive · medium: 77s77sclaude-sonnet-5 · universal · medium: 42s42shardclaude-sonnet-5 · naive · hard: 112s112sclaude-sonnet-5 · universal · hard: 85s85s
cost per solve & wall per attempt
configeasy $/solvemedium $/solvehard $/solveeasy wallmedium wallhard wall
naive @ haiku$0.18$0.44$0.4990s122s157s
universal @ haiku$0.10$0.16$0.2359s71s74s
naive @ sonnet$0.09$0.15$0.1861s77s112s
universal @ sonnet$0.07$0.09$0.1336s42s85s

Comparison with SOTA (rocq-mcp) — fair integration

github.com/LLM4Rocq/rocq-mcp, measured under the identical harness, gate, problems, and policies (first contact only after our design freeze; REPORT §SOTA). Bars: pass@1. Table: all three dimensions — $/solve = expected cost per solved proof (failures included), wall = mean s/attempt. Numbers from the FAIR rerun (Jul 8): an audit found the original runs started with rocq-mcp's server still connecting in 97-98 % of attempts (our integration artifact — retracted); an instant-handshake proxy fixed it, 120/120 connected. Fair verdict: near accuracy-parity at sonnet, rocq-mcp edges easy at haiku; rocq-tools leads weak-policy medium/hard and costs about half per solved proof at sonnet. See REPORT for the full audit story.

rocq-mcp (SOTA)naive whole-fileuniversal (ours)
claude-haiku-4-5easyclaude-haiku-4-5 · rocq-mcp (SOTA) · easy: 0.6750.68claude-haiku-4-5 · naive · easy: 0.4380.44claude-haiku-4-5 · universal · easy: 0.6620.66mediumclaude-haiku-4-5 · rocq-mcp (SOTA) · medium: 0.3500.35claude-haiku-4-5 · naive · medium: 0.2500.25claude-haiku-4-5 · universal · medium: 0.5750.57hardclaude-haiku-4-5 · rocq-mcp (SOTA) · hard: 0.3250.33claude-haiku-4-5 · naive · hard: 0.3000.30claude-haiku-4-5 · universal · hard: 0.4000.40claude-sonnet-5easyclaude-sonnet-5 · rocq-mcp (SOTA) · easy: 0.9500.95claude-sonnet-5 · naive · easy: 0.9250.93claude-sonnet-5 · universal · easy: 0.9500.95mediumclaude-sonnet-5 · rocq-mcp (SOTA) · medium: 0.9250.93claude-sonnet-5 · naive · medium: 0.9500.95claude-sonnet-5 · universal · medium: 1.0001.00hardclaude-sonnet-5 · rocq-mcp (SOTA) · hard: 0.8000.80claude-sonnet-5 · naive · hard: 0.8000.80claude-sonnet-5 · universal · hard: 0.8500.85
all three dimensions
configeasy pass@1medium pass@1hard pass@1easy $/solvemedium $/solvehard $/solveeasy wallmedium wallhard wall
rocq-mcp @ haiku0.680.350.33$0.08$0.22$0.3165s76s97s
naive @ haiku0.440.250.30$0.18$0.44$0.4990s122s157s
universal @ haiku0.660.570.40$0.10$0.16$0.2359s71s74s
rocq-mcp @ sonnet0.950.930.80$0.11$0.19$0.2650s64s105s
naive @ sonnet0.930.950.80$0.09$0.15$0.1861s77s112s
universal @ sonnet0.951.000.85$0.07$0.09$0.1336s42s85s

Scalability — N parallel agents on one machine

Each point is one full run of a fixed 24-problem stratified batch (8 per bucket) executed with N = 1, 2, 4, 8 concurrent agents — eight runs total (4 per interface), all in a single night window (Jul 3–4) after two earlier measurement artifacts were caught and disclosed (REPORT §5). Both interfaces scale healthily (flat wall per attempt, ~72–80 % parallel efficiency at N=8); the winner's advantage is its ~2.7× per-attempt speed, compounding to ≈6× solved-proofs-per-hour at N=8. The local prover is never the bottleneck (CPU ≤ 24 % of a 14-core laptop).

naive baselinesession winner
throughput — attempts / hour0239478N=1N=2N=4N=8naive baseline · N=1: 28 (3/24 solved)28naive baseline · N=2: 50 (4/24 solved)50naive baseline · N=4: 92 (3/24 solved)92naive baseline · N=8: 161 (4/24 solved)161session winner · N=1: 75 (11/24 solved)75session winner · N=2: 143 (10/24 solved)143session winner · N=4: 272 (9/24 solved)272session winner · N=8: 478 (8/24 solved)478wall-clock per attempt (s)070140N=1N=2N=4N=8naive baseline · N=1: 128 (3/24 solved)128naive baseline · N=2: 140 (4/24 solved)140naive baseline · N=4: 135 (3/24 solved)135naive baseline · N=8: 133 (4/24 solved)133session winner · N=1: 48 (11/24 solved)48session winner · N=2: 50 (10/24 solved)50session winner · N=4: 50 (9/24 solved)50session winner · N=8: 51 (8/24 solved)51
full table (incl. resources; *RSS is a machine-wide upper bound)
interfaceN agentsattempts/hwall s/attemptsolvedpeak RSS MB*CPU %
naive baseline128.1128.03/241140.94.6
naive baseline250.5140.04/241487.34.4
naive baseline491.5135.13/244929.213.3
naive baseline8161.3133.44/244025.212.3
session winner174.747.911/24962.75.3
session winner2142.649.710/241720.08.4
session winner4272.349.79/243319.816.8
session winner8478.351.38/246391.723.8

Everything else measured

Cross-policy annexes, SOTA comparison, team experiments, in-project probes, held-out — each analyzed in docs/REPORT.md. Raw per-run numbers below; reproduce any row with python3 harness/report.py <run>.

all 40 non-ladder runs
runpolicyattemptspass@1 e/m/h
smoke_baseline_1claude-haiku-4-550.80 / – / –
smoke_session_1claude-haiku-4-550.80 / – / –
session_try_dev150claude-haiku-4-53000.45 / 0.33 / 0.35
session_try_hints_minif2f_validclaude-haiku-4-54880.32 / 0.13 / 0.00
session_try_hints_v2_minif2f_validclaude-haiku-4-54880.57 / 0.30 / 0.06
sweep_session_try_hints_auto_N1claude-haiku-4-5240.38 / 0.62 / 0.38
sweep_session_try_hints_auto_N2claude-haiku-4-5240.38 / 0.50 / 0.38
sweep_session_try_hints_auto_N4claude-haiku-4-5240.25 / 0.62 / 0.25
sweep_session_try_hints_auto_N8claude-haiku-4-5240.25 / 0.50 / 0.25
baseline_sonnet_dev60claude-sonnet-51200.93 / 0.95 / 0.80
session_try_hints_auto_sonnet_dev60claude-sonnet-51200.93 / 0.82 / 0.70
solo_hard70claude-haiku-4-5140– / – / 0.40
team_k3_hard70claude-haiku-4-5140– / – / 0.32
unified_sonnet_dev60claude-sonnet-51200.93 / 0.85 / 0.75
FINAL_minif2f_testclaude-haiku-4-54880.52 / 0.13 / 0.04
sonnet_native_dev60claude-sonnet-51200.95 / 0.85 / 0.75
rocq_mcp_smokeclaude-haiku-4-530.33 / – / –
rocq_mcp_dev60claude-haiku-4-51200.45 / 0.23 / 0.23
winner_ctx_full_inproject60claude-haiku-4-5120– / 0.72 / –
winner_ctx_lean_inproject60claude-haiku-4-5120– / 0.60 / –
rocq_mcp_sonnet_dev60claude-sonnet-51200.82 / 0.72 / 0.72
sonnet_native_auto2_dev60claude-sonnet-5601.00 / 0.85 / 0.85
solo_decomposableclaude-haiku-4-5540.67 / 0.75 / 0.59
team_decomposableclaude-haiku-4-5540.33 / 0.25 / 0.18
winner_ctx_lean_mathcompclaude-haiku-4-535– / 0.07 / –
sweep_baseline_N1claude-haiku-4-5240.12 / 0.00 / 0.25
sweep_baseline_N2claude-haiku-4-5240.12 / 0.12 / 0.25
sweep_baseline_N4claude-haiku-4-5240.12 / 0.00 / 0.25
sweep_baseline_N8claude-haiku-4-5240.12 / 0.25 / 0.12
universal_sonnet_dev60claude-sonnet-51200.95 / 1.00 / 0.85
winner_ctx_lean_ssr_mathcompclaude-haiku-4-535– / 0.07 / –
universal_fable_dev60claude-fable-5600.95 / 1.00 / 1.00
baseline_fable_dev60claude-fable-5600.95 / 1.00 / 0.95
ctx_full_fable_mathcompclaude-fable-535– / 0.93 / –
ctx_lean_fable_mathcompclaude-fable-535– / 0.93 / –
winner_ctx_lean_ex_mathcompclaude-haiku-4-535– / 0.07 / –
winner_ctx_lean_mc_mathcompclaude-haiku-4-535– / 0.00 / –
mcp_timeout_probeclaude-haiku-4-510.00 / – / –
rocq_mcp_fair_dev60claude-haiku-4-51200.68 / 0.35 / 0.33
rocq_mcp_fair_sonnet_dev60claude-sonnet-51200.95 / 0.93 / 0.80