Published copy. Evidence paths refer to the repository branch insights/generator-benchmark-b2.
Insights Goal B2: generator benchmark with subscription-backed CLI arms (2026-09-07)
Question: is the SQL generator the lever? Same code, same prompts, same fixture, same sealed holdout and scorer for every arm; only the generator (provider, model and, for the new arms, the CLI transport) varies.
The control arm (glm-5.2), the glm-5.3 arm and the frozen noise floor are Goal B's records (revision ff3e669e469d), inherited unchanged; the arms numbered 6 onwards ran on this freeze (revision e528a201e05c) through the Claude Code and Codex CLIs on the owner's subscriptions. This is a quality comparison: CLI latency is advisory.
Execution deviated from the frozen sequential-arms protocol: the CLI arms ran concurrently on the host, one driver per arm, with repeats sequential within an arm (details under Run inventory and Method). Quality metrics are compared; CLI wall time is advisory only and is not compared.
These results compare configured generator-and-transport combinations (provider, model, thinking setting, CLI or API path); they do not isolate model weights from reasoning settings or CLI behaviour.
Rendered 2026-09-08 01:28 UTC from the records listed under Evidence; every number is in results/<arm>.json, results/comparison.json and results/inventory.json.
Decision
Answer: yes by the frozen rule for gpt-6-astra-codex, on the precision test only: precision 75.5% against the control's 50.8%, paired difference 24.8% [4.3% to 46.2%], above the 9.0% noise floor. It fails the first-turn error test in the wrong direction (errors 73.0 against 55.7 per 100) and completes fewer answers correctly (26.7 against 35.3 per 250); the section below states the trade-off. Arms whose first-turn error gain exceeds the noise floor: claude-fable-5-1-cli, claude-opus-5-cli, gpt-5.6-sol-codex. Arms whose precision gain exceeds the noise floor: gpt-6-astra-codex. glm-5.3 was served by the same reported model as the control (glm-5.3), so that comparison measures request-configuration noise, not a second generator. Arms that could not run: claude-fable-5-1, claude-opus-5, claude-sonnet-5, glm-5.2-general (reasons in the Arms table). The planner failure class stays the largest first-turn cause in the control; it is about 31 of 100 first turns there.
Rule: an arm wins if it halves first-turn errors, or lifts precision above 70% with the lower paired bound above the control, and the gain exceeds the control noise floor (the control's own max-minus-min spread across its three repeats: first-turn errors 3.0, precision 9.0%).
glm-5.3: verdict NO. First-turn errors: 4.0 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: no; halves first-turn errors: no. Precision: 2.0% points against the control; noise floor 9.0%; gain exceeds the noise floor: no; above 70% with the lower paired bound above the control: no (paired difference 2.3% [-5.1% to 10.3%]).
claude-fable-5-1-cli: verdict NO. First-turn errors: -9.7 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: yes; halves first-turn errors: no. Precision: 2.3% points against the control; noise floor 9.0%; gain exceeds the noise floor: no; above 70% with the lower paired bound above the control: no (paired difference 2.4% [-11.6% to 15.5%]).
claude-opus-5-cli: verdict NO. First-turn errors: -17.0 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: yes; halves first-turn errors: no. Precision: 4.4% points against the control; noise floor 9.0%; gain exceeds the noise floor: no; above 70% with the lower paired bound above the control: no (paired difference 4.5% [-7.7% to 16.7%]).
claude-sonnet-5-cli: verdict NO. First-turn errors: 4.7 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: no; halves first-turn errors: no. Precision: -0.7% points against the control; noise floor 9.0%; gain exceeds the noise floor: no; above 70% with the lower paired bound above the control: no (paired difference -0.7% [-14.7% to 13.2%]).
gpt-5.6-sol-codex: verdict NO. First-turn errors: -9.7 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: yes; halves first-turn errors: no. Precision: -4.0% points against the control; noise floor 9.0%; gain exceeds the noise floor: no; above 70% with the lower paired bound above the control: no (paired difference -4.0% [-18.6% to 11.1%]).
gpt-6-astra-codex: verdict YES. First-turn errors: 17.3 per 100 against the control (negative means fewer errors); noise floor 3.0; gain exceeds the noise floor: no; halves first-turn errors: no. Precision: 24.7% points against the control; noise floor 9.0%; gain exceeds the noise floor: yes; above 70% with the lower paired bound above the control: yes (paired difference 24.8% [4.3% to 46.2%]).
What each arm changed against the control
The rule scores two metrics. The others moved too, and a reader deciding what to run next needs them beside the noise floor: every number below is the mean across repeats, with the paired family-cluster interval where the metric is paired and the control's own spread as the noise floor.
glm-5.3: correct first turns 14.3 against 15.0 per 100 (paired difference -0.7 [-3.3 to 2.3]; noise floor 2.0); silent wrong 30.7 against 34.3 per 250 (paired difference -3.7 [-10.7 to 3.0]; noise floor 9.0); correct completed 34.7 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 13.0 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 2.7 against 2.3 per 50 (paired difference 0.3 [-0.7 to 1.7]; noise floor 1.0). First-turn errors rose by 4.0 per 100 (paired interval 4.0 [-0.7 to 8.3]), beyond the noise floor in the wrong direction, though the paired interval includes zero.
claude-fable-5-1-cli: correct first turns 19.7 against 15.0 per 100 (paired difference 4.7 [-1.7 to 12.0]; noise floor 2.0); silent wrong 45.0 against 34.3 per 250 (paired difference 10.7 [-3.0 to 24.3]; noise floor 9.0); correct completed 51.0 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 20.7 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 5.3 against 2.3 per 50 (paired difference 3.0 [0.3 to 6.3]; noise floor 1.0). First-turn errors fell by 9.7 per 100 (paired interval -9.7 [-17.7 to -2.0]), outside the noise floor but short of the halving the rule requires (27.8 or fewer per 100). The observed silent-wrong increase (10.7 per 250, paired interval 10.7 [-3.0 to 24.3]) exceeds the control spread, but its paired interval includes zero, so an increase is not established. No precision change is established either: its difference (2.3% points) is within the control spread and its paired interval includes zero. These aggregate results do not determine the error rate of the newly executing first turns. Full conversations rose by 3.0 per 50 with the paired interval above zero and beyond the noise floor; the rule does not score this metric either.
claude-opus-5-cli: correct first turns 26.0 against 15.0 per 100 (paired difference 11.0 [4.3 to 18.0]; noise floor 2.0); silent wrong 49.7 against 34.3 per 250 (paired difference 15.3 [-0.3 to 30.7]; noise floor 9.0); correct completed 61.3 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 13.0 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 5.3 against 2.3 per 50 (paired difference 3.0 [0.3 to 6.3]; noise floor 1.0). First-turn errors fell by 17.0 per 100 (paired interval -17.0 [-24.7 to -9.7]), outside the noise floor but short of the halving the rule requires (27.8 or fewer per 100). The observed silent-wrong increase (15.3 per 250, paired interval 15.3 [-0.3 to 30.7]) exceeds the control spread, but its paired interval includes zero, so an increase is not established. No precision change is established either: its difference (4.4% points) is within the control spread and its paired interval includes zero. These aggregate results do not determine the error rate of the newly executing first turns. Correct first turns rose by 11.0 per 100 with the paired interval above zero and beyond the noise floor; the rule does not score this metric. Full conversations rose by 3.0 per 50 with the paired interval above zero and beyond the noise floor; the rule does not score this metric either.
claude-sonnet-5-cli: correct first turns 15.7 against 15.0 per 100 (paired difference 0.7 [-3.7 to 6.0]; noise floor 2.0); silent wrong 32.7 against 34.3 per 250 (paired difference -1.7 [-14.0 to 10.0]; noise floor 9.0); correct completed 32.7 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 13.3 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 3.0 against 2.3 per 50 (paired difference 0.7 [0.0 to 1.7]; noise floor 1.0). First-turn errors rose by 4.7 per 100 (paired interval 4.7 [-4.3 to 13.7]), beyond the noise floor in the wrong direction, though the paired interval includes zero.
gpt-5.6-sol-codex: correct first turns 15.7 against 15.0 per 100 (paired difference 0.7 [-4.0 to 6.0]; noise floor 2.0); silent wrong 46.0 against 34.3 per 250 (paired difference 11.7 [-5.0 to 28.7]; noise floor 9.0); correct completed 40.3 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 28.0 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 2.7 against 2.3 per 50 (paired difference 0.3 [0.0 to 1.0]; noise floor 1.0). First-turn errors fell by 9.7 per 100 (paired interval -9.7 [-17.0 to -2.3]), outside the noise floor but short of the halving the rule requires (27.8 or fewer per 100). The observed silent-wrong increase (11.7 per 250, paired interval 11.7 [-5.0 to 28.7]) exceeds the control spread, but its paired interval includes zero, so an increase is not established. No precision change is established either: its difference (-4.0% points) is within the control spread and its paired interval includes zero. These aggregate results do not determine the error rate of the newly executing first turns.
gpt-6-astra-codex: correct first turns 12.7 against 15.0 per 100 (paired difference -2.3 [-6.0 to 1.3]; noise floor 2.0); silent wrong 8.7 against 34.3 per 250 (paired difference -25.7 [-37.7 to -15.3]; noise floor 9.0); correct completed 26.7 against 35.3 per 250 (not a paired metric; noise floor 6.0); refusals 8.0 against 14.0 per 250 (not a paired metric; noise floor 6.0); full conversations 2.0 against 2.3 per 50 (paired difference -0.3 [-1.0 to 0.0]; noise floor 1.0). The verdict rests on the precision test alone. Precision rose because silent wrong answers fell (25.7 fewer per 250), not because correct answers rose (correct completed fell by 8.7 per 250), and first-turn errors rose by 17.3 per 100 (paired interval 17.3 [10.7 to 24.7]), outside the noise floor in the wrong direction. Under the frozen rule this arm wins; what it delivers is a generator that answers less often and is wrong less often when it answers, not one that answers more questions correctly.
Arms
API arms record the model identifier the provider returned per call. The Claude CLI arms record the model Claude Code billed per call (its modelUsage block); the Codex CLI arms record the model the CLI's own banner reports, which echoes the request and is not a server-side identifier.
The glm-5.2 alias question. The zai coding endpoint returned glm-5.3 for a glm-5.2 request at this freeze. Its Anthropic-compatible coding path returned glm-5.3. The general endpoint (api.z.ai/api/paas/v4) refused the account: RateLimitError: Error code: 429 - {'error': {'code': '1113', 'message': 'Insufficient balance or no resource package. Please recharge.'}}. No served model id could be read there, so the question is not closed at that endpoint. Both successfully probed coding paths returned glm-5.3; the general endpoint remains unverified. The control arm's own calls were served as glm-5.3 (1456 calls).
#
arm
transport
model requested
model returned or reported
thinking
sampling and thinking control
per-call deadline
status
reason
1
glm-5.2
API, coding endpoint
glm-5.2
glm-5.3 (1456 calls)
disabled
T=0, seed=42: sent and accepted; thinking disabled requested
25 s
available (inherited)
2
glm-5.3
API, coding endpoint
glm-5.3
glm-5.3 (1471 calls)
disabled
T=0, seed=42: sent and accepted; thinking disabled requested
25 s
available (inherited)
3
claude-fable-5-1
API
claude-fable-5-1
not probed
adaptive_low
temperature and seed are not supported by the Messages API for this model family; not sent
25 s
unavailable (inherited)
No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
4
claude-opus-5
API
claude-opus-5
not probed
disabled
temperature and seed are not supported by the Messages API for this model family; not sent
25 s
unavailable (inherited)
No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
5
claude-sonnet-5
API
claude-sonnet-5
not probed
disabled
temperature and seed are not supported by the Messages API for this model family; not sent
25 s
unavailable (inherited)
No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
6
claude-fable-5-1-cli
claude CLI, slot 5
claude-fable-5-1
claude-fable-5-1 (1572 calls)
adaptive_low
no CLI sampling control; thinking: --effort low (adaptive thinking)
75 s
available
7
claude-opus-5-cli
claude CLI, slot 3
claude-opus-5
claude-opus-5 (1485 calls)
disabled
no CLI sampling control; thinking: MAX_THINKING_TOKENS=0
75 s
available
8
claude-sonnet-5-cli
claude CLI, slot 2
claude-sonnet-5
claude-sonnet-5 (1564 calls)
disabled
no CLI sampling control; thinking: MAX_THINKING_TOKENS=0
75 s
available
9
gpt-5.6-sol-codex
codex CLI
gpt-5.6-sol
gpt-5.6-sol (1341 calls)
adaptive_low
no CLI sampling control; reasoning effort low
75 s
available
10
glm-5.2-general
API, general endpoint
glm-5.2
none (probe refused)
disabled
T=0 and seed=42 requested; the probe was refused, so acceptance was not established. Thinking disabled requested
25 s
unavailable
Provider probe failed: RateLimitError: Error code: 429 - {'error': {'code': '1113', 'message': 'Insufficient balance or no resource package. Please recharge.'}}
11
gpt-6-astra-codex
codex CLI
gpt-6-astra
gpt-6-astra (1189 calls)
adaptive_low
no CLI sampling control; reasoning effort low
75 s
available
Run inventory
The frozen protocol says arms run sequentially. Execution deviated. The sequential driver started at 21:48:12 UTC and ran the claude-fable-5-1-cli smoke and repeat 1. At 22:21:37 UTC its loop was stopped while that repeat 1 kept running (it finished at 22:30:21). Per-arm drivers were then launched on a staggered schedule: claude-opus-5-cli 22:21:45, claude-sonnet-5-cli 22:23:02, gpt-5.6-sol-codex 22:24:19 (crashed in its smoke, relaunched 22:28:21), gpt-6-astra-codex 22:25:36, and claude-fable-5-1-cli for repeats 2 and 3 at 22:31:23 once repeat 1's receipt existed. Up to five CLI arms overlapped, with repeats sequential within each arm; the last repeat finished at 00:24:59 UTC. All fifteen B2 repeats used the B2 freeze (same manifest, snapshot, fixture, corpus, scorer and clocks); the inherited control and glm-5.3 records keep their Goal B freeze and ran alone on an otherwise idle host (1-minute load 0.4 to 1.2 at start and end).
Why: Elapsed time. The whole B2 execution took about 2 hours 37 minutes including the initial sequential phase; a fully sequential schedule was estimated at roughly eight hours for fifteen repeats (about 25 to 40 minutes each). Consequence: Quality metrics are compared on identical scoring inputs (scorer, corpus, clocks, fixture). Within the SQL-generator path B2 recorded zero provider-call failures, zero 75 s CLI deadline kills and zero 180 s job-deadline expirations. The auxiliary-generation traces (presentation calls made by the application after the SQL step, with their own 25 s budget) separately recorded 54 optional-stage cancellations (JobStoppedError: fable 14, opus 14, sonnet 12, sol 8, astra 6; the inherited control recorded 10 and glm-5.3 4 under sequential execution); their relationship to concurrency was not established. No contention effect on quality was identified, and equivalence to sequential execution was not tested. CLI wall time is advisory only: it was measured under up to five concurrent arms (1-minute load average up to 14.5 on 16 CPUs, against 0.4 to 1.2 for the inherited Goal B runs) and is not compared across arms or with the API arms.
States: complete = 250 attempts recorded and exit 0; incomplete by limits = stopped by the 12-hour subscription cap with the attempts recorded so far; crashed = restarted once, record kept beside the restart; missing = no record. Pauses are the subscription budget pauses recorded per run; served model is the identifier in the traces (Claude Code's billed model; the Codex banner echo; the zai response model). Load is the 1-minute average at the start and end of the run.
arm
repeat
state
attempts
exit
generator calls
failed calls
served model (calls)
pauses (total)
started (UTC)
finished (UTC)
1-min load start to end
port
driver
glm-5.2
1
complete
250
0
488
0
glm-5.3 (488)
0 (0 s)
2026-09-07T16:46:16
2026-09-07T17:12:09
0.9 to 0.5
8190
inherited (Goal B sequential runs)
glm-5.2
2
complete
250
0
476
0
glm-5.3 (476)
0 (0 s)
2026-09-07T17:22:54
2026-09-07T17:46:01
0.4 to 0.5
8190
inherited (Goal B sequential runs)
glm-5.2
3
complete
250
0
492
0
glm-5.3 (492)
0 (0 s)
2026-09-07T17:46:14
2026-09-07T18:10:49
0.7 to 0.6
8190
inherited (Goal B sequential runs)
glm-5.3
1
complete
250
0
492
0
glm-5.3 (492)
0 (0 s)
2026-09-07T18:11:17
2026-09-07T18:34:49
0.7 to 1.0
8190
inherited (Goal B sequential runs)
glm-5.3
2
complete
250
0
491
0
glm-5.3 (491)
0 (0 s)
2026-09-07T18:35:04
2026-09-07T18:59:31
1.1 to 1.1
8190
inherited (Goal B sequential runs)
glm-5.3
3
complete
250
0
489
1
glm-5.3 (488)
0 (0 s)
2026-09-07T18:59:46
2026-09-07T19:25:40
1.2 to 0.9
8190
inherited (Goal B sequential runs)
claude-fable-5-1
unavailable: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
0
n/a
0 (0 s)
none
claude-opus-5
unavailable: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
0
n/a
0 (0 s)
none
claude-sonnet-5
unavailable: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
port race: the per-arm drivers were started before commit fdc307fb (honour BENCH_PORT_START) and two of them chose the same runtime port, so the smoke's admission probe reached another arm's server
the driver was relaunched at 22:28:21Z with its own port range (8260 onwards); its smoke passed at 22:31:09Z (5 of 5 questions functioning) and all three repeats completed. Smoke runs are not part of the frozen runs; no frozen repeat crashed and no repeat was restarted.
The control arm ran 3 times on identical code in Goal B (frozen 2026-09-07T18:11:04 UTC, inherited here unchanged). 172 of 250 cases had the same outcome in every repeat; 78 did not. Stable cases by outcome: comparison_unverified 8, correct_completed 17, refusal_unverified 11, silent_wrong 12, task_error 124.
A difference between arms smaller than a metric's spread (max minus min across the control repeats) is not a finding.
The noise floor is the observed range of three control runs, not a confidence bound. Exceeding it describes the point estimate; it does not establish that the true gain exceeds that range. The paired interval resamples intent families after averaging repeats; it does not resample run-to-run variability and makes no adjustment for the number of arms compared.
metric
mean
min
max
spread (noise floor)
first_turn_execution_errors
55.67
54.00
57.00
3.00
correct_first_turns
15.00
14.00
16.00
2.00
answer_precision
50.81%
45.95%
54.93%
8.98%
silent_wrong
34.33
31.00
40.00
9.00
correct_completed
35.33
33.00
39.00
6.00
refusals
14.00
11.00
17.00
6.00
comparison_unverified
13.00
10.00
15.00
5.00
full_conversation_successes
2.33
2.00
3.00
1.00
Per-arm results
Denominators are fixed: 100 supported first turns, 250 attempts, 50 conversations. Refusals, blocked and comparison-unverified answers earn no completion credit. Each cell is the mean across the arm's complete repeats with the min and max beside it; repeats stopped by subscription limits are excluded from the means and listed under budget handling.
Reference only, not an arm: goal-3 original e479e81 (single E5 run, different code; reference only) scored first-turn errors 60/100, correct first turns 14/100, precision 53.7% (36/67), silent wrong 31/250, full conversations 1/50.
Headline metrics per arm: bar = mean across repeats, whisker = min to max. First-turn errors and correct first turns are per 100 supported first turns; precision is per answered attempt; silent wrong is per 250 attempts.Table view
metric
glm-5.2 mean [min to max]
glm-5.3 mean [min to max]
claude-fable-5-1-cli mean [min to max]
claude-opus-5-cli mean [min to max]
claude-sonnet-5-cli mean [min to max]
gpt-5.6-sol-codex mean [min to max]
gpt-6-astra-codex mean [min to max]
glm-5.3 vs control, paired difference [95% CI]
claude-fable-5-1-cli vs control, paired difference [95% CI]
claude-opus-5-cli vs control, paired difference [95% CI]
claude-sonnet-5-cli vs control, paired difference [95% CI]
gpt-5.6-sol-codex vs control, paired difference [95% CI]
gpt-6-astra-codex vs control, paired difference [95% CI]
First-turn execution errors / 100
55.7 [54.0 to 57.0]
59.7 [56.0 to 62.0]
46.0 [43.0 to 49.0]
38.7 [37.0 to 41.0]
60.3 [58.0 to 62.0]
46.0 [43.0 to 49.0]
73.0 [71.0 to 75.0]
4.0 [-0.7 to 8.3]
-9.7 [-17.7 to -2.0]
-17.0 [-24.7 to -9.7]
4.7 [-4.3 to 13.7]
-9.7 [-17.0 to -2.3]
17.3 [10.7 to 24.7]
Correct first turns / 100
15.0 [14.0 to 16.0]
14.3 [13.0 to 16.0]
19.7 [19.0 to 20.0]
26.0 [25.0 to 27.0]
15.7 [15.0 to 17.0]
15.7 [15.0 to 16.0]
12.7 [11.0 to 14.0]
-0.7 [-3.3 to 2.3]
4.7 [-1.7 to 12.0]
11.0 [4.3 to 18.0]
0.7 [-3.7 to 6.0]
0.7 [-4.0 to 6.0]
-2.3 [-6.0 to 1.3]
Precision = correct / (correct + silent wrong)
50.8% [45.9% to 54.9%]
52.8% [49.1% to 55.3%]
53.1% [51.0% to 54.7%]
55.3% [54.1% to 56.0%]
50.1% [47.7% to 54.0%]
46.8% [44.6% to 50.0%]
75.5% [72.2% to 80.0%]
2.3% [-5.1% to 10.3%]
2.4% [-11.6% to 15.5%]
4.5% [-7.7% to 16.7%]
-0.7% [-14.7% to 13.2%]
-4.0% [-18.6% to 11.1%]
24.8% [4.3% to 46.2%]
Silent wrong / 250
34.3 [31.0 to 40.0]
30.7 [29.0 to 34.0]
45.0 [43.0 to 47.0]
49.7 [48.0 to 51.0]
32.7 [29.0 to 35.0]
46.0 [40.0 to 51.0]
8.7 [7.0 to 10.0]
-3.7 [-10.7 to 3.0]
10.7 [-3.0 to 24.3]
15.3 [-0.3 to 30.7]
-1.7 [-14.0 to 10.0]
11.7 [-5.0 to 28.7]
-25.7 [-37.7 to -15.3]
Correct completed / 250
35.3 [33.0 to 39.0]
34.7 [28.0 to 42.0]
51.0 [49.0 to 52.0]
61.3 [60.0 to 63.0]
32.7 [31.0 to 34.0]
40.3 [40.0 to 41.0]
26.7 [26.0 to 28.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
Refusals / 250
14.0 [11.0 to 17.0]
13.0 [12.0 to 14.0]
20.7 [18.0 to 24.0]
13.0 [12.0 to 14.0]
13.3 [11.0 to 17.0]
28.0 [25.0 to 30.0]
8.0 [7.0 to 9.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
Blocked by parent / 250
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
Comparison unverified / 250
13.0 [10.0 to 15.0]
12.3 [10.0 to 15.0]
14.3 [13.0 to 16.0]
12.0 [11.0 to 13.0]
11.3 [9.0 to 13.0]
14.0 [12.0 to 15.0]
15.0 [15.0 to 15.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
Unnecessary clarification / 250
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
Full conversation success / 50
2.3 [2.0 to 3.0]
2.7 [2.0 to 4.0]
5.3 [5.0 to 6.0]
5.3 [5.0 to 6.0]
3.0 [2.0 to 4.0]
2.7 [2.0 to 3.0]
2.0 [2.0 to 2.0]
0.3 [-0.7 to 1.7]
3.0 [0.3 to 6.3]
3.0 [0.3 to 6.3]
0.7 [0.0 to 1.7]
0.3 [0.0 to 1.0]
-0.3 [-1.0 to 0.0]
Jobs at deadline / 250
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
not a paired metric
First-turn execution errors by cause (per 100 supported first turns)
Causes follow the goal-3 classification of the persisted terminal record and the retained generated SQL. Provider or transport failures (including CLI deadline kills and usage-limit errors) are listed separately and are counted.
cause
glm-5.2 mean [min to max]
glm-5.3 mean [min to max]
claude-fable-5-1-cli mean [min to max]
claude-opus-5-cli mean [min to max]
claude-sonnet-5-cli mean [min to max]
gpt-5.6-sol-codex mean [min to max]
gpt-6-astra-codex mean [min to max]
Generated SQL lacked schema qualification
7.0 [6.0 to 8.0]
8.3 [6.0 to 10.0]
4.7 [4.0 to 5.0]
0.0 [0.0 to 0.0]
9.7 [8.0 to 11.0]
3.3 [3.0 to 4.0]
0.0 [0.0 to 0.0]
Generated function used the wrong type
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.3 [0.0 to 1.0]
0.7 [0.0 to 2.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
Generated schema reference was wrong or absent
9.7 [8.0 to 13.0]
14.0 [12.0 to 16.0]
12.0 [11.0 to 13.0]
11.3 [9.0 to 13.0]
25.7 [24.0 to 28.0]
0.3 [0.0 to 1.0]
0.0 [0.0 to 0.0]
Generation returned no usable query
0.7 [0.0 to 1.0]
1.3 [0.0 to 2.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
Planner returned no usable source SQL
31.0 [28.0 to 35.0]
28.7 [27.0 to 31.0]
16.0 [14.0 to 18.0]
4.7 [3.0 to 7.0]
14.0 [12.0 to 16.0]
21.7 [19.0 to 25.0]
73.0 [71.0 to 75.0]
Required merge key was missing
4.7 [3.0 to 7.0]
6.0 [4.0 to 7.0]
10.0 [9.0 to 11.0]
20.0 [18.0 to 23.0]
9.3 [7.0 to 13.0]
20.7 [20.0 to 21.0]
0.0 [0.0 to 0.0]
Terminal boundary needs review
1.3 [1.0 to 2.0]
1.0 [1.0 to 1.0]
0.0 [0.0 to 0.0]
0.7 [0.0 to 1.0]
0.3 [0.0 to 1.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
Validation boundary needs review
1.3 [1.0 to 2.0]
0.3 [0.0 to 1.0]
3.0 [2.0 to 4.0]
1.3 [1.0 to 2.0]
1.3 [0.0 to 2.0]
0.0 [0.0 to 0.0]
0.0 [0.0 to 0.0]
Total first-turn execution errors
55.7 [54.0 to 57.0]
59.7 [56.0 to 62.0]
46.0 [43.0 to 49.0]
38.7 [37.0 to 41.0]
60.3 [58.0 to 62.0]
46.0 [43.0 to 49.0]
73.0 [71.0 to 75.0]
Regressions versus the control (stable cases only)
A case counts as lost only if it was correct in all control repeats and wrong or error in all candidate repeats. Gained cases follow the mirror rule.
API arms: latency is per SQL-generator provider call, measured inside the isolated server; cost is list price times the tokens the provider reported, summed over one 250-attempt run. CLI arms: latency is the child process's wall time per call, which carries the CLI's start-up and its own request pipeline. CLI wall times were measured while up to five arms ran concurrently on the host (1-minute load average up to 14.5 on 16 CPUs), while the inherited API arms ran alone (load 0.41 to 1.2); they are advisory only and are not compared, across arms or with the API arms.
Subscription arms report tokens only and no dollar figure: the Claude Code and Codex accounts are flat-rate subscriptions, so no per-call price is paid and there is no cost to report. Claude Code prints a list-price equivalent per call; it is retained in the traces as a CLI field but it is not money paid and is not shown. Token counts are what the CLI reports: Claude Code reports input (cache creation and cache reads included), output and thinking tokens per call. The Codex CLI printed one unsplit 'tokens used' total on stderr for every generator call; the frozen collector searched stdout only, so the provider traces carry no Codex token figure. The totals were recovered after the run from the retained per-call stderr transcripts (codex-tokens-recovered.json beside each repeat) and are shown as totals; input, output and reasoning splits are unavailable for Codex.
API-arm latency distributions, all completed calls of all repeats pooled. Box p25 to p75, line = median, whiskers p5 to p95. CLI wall-time observations appear only in the table for audit purposes; differing execution conditions (concurrent arms, CLI start-up) prevent comparisons across CLI arms or against API arms.Cost per 250-attempt run against precision, API arms only (list price times the tokens the provider reported). Subscription arms have no price and are absent. Small dots are single repeats; the large dot is the mean of repeats.Table view
arm
calls per run
failed calls per run
p50 ms
p95 ms
input tokens per run
output tokens per run
reasoning or thinking tokens per run
cost basis
glm-5.2
485 [476 to 492]
0 [0 to 0]
5362 [5170 to 5577]
10109 [9652 to 10652]
3689078 [3665204 to 3701821]
141303 [134396 to 145893]
59354 [54532 to 62615]
list $1.40 / $4.40 per 1M: 5.786 [5.723 to 5.824]
glm-5.3
491 [489 to 492]
0 [0 to 1]
5307 [5088 to 5561]
10607 [10130 to 10962]
3726734 [3710575 to 3748945]
143361 [142287 to 144163]
60266 [59411 to 61039]
list $1.40 / $4.40 per 1M: 5.848 [5.827 to 5.883]
claude-fable-5-1-cli
524 [517 to 528]
0 [0 to 0]
9263 [9261 to 9265] (CLI wall)
14366 [13869 to 15061] (CLI wall)
6513853 [6475086 to 6537380]
315658 [310632 to 318712]
141052 [137874 to 142695]
subscription: tokens only, no dollar figure (flat-rate plan, no per-call price is paid)
claude-opus-5-cli
495 [490 to 498]
0 [0 to 0]
5378 [5314 to 5441] (CLI wall)
7542 [7135 to 8124] (CLI wall)
6352488 [6313724 to 6373262]
172538 [168045 to 176827]
0 [0 to 0]
subscription: tokens only, no dollar figure (flat-rate plan, no per-call price is paid)
claude-sonnet-5-cli
521 [520 to 523]
0 [0 to 0]
5025 [4987 to 5098] (CLI wall)
7322 [7167 to 7454] (CLI wall)
6466809 [6447989 to 6487192]
173346 [172854 to 173718]
0 [0 to 0]
subscription: tokens only, no dollar figure (flat-rate plan, no per-call price is paid)
gpt-5.6-sol-codex
447 [446 to 449]
0 [0 to 0]
10362 [9543 to 11988] (CLI wall)
17353 [16082 to 19807] (CLI wall)
total 5113742 [4908869 to 5462646] (unsplit; recovered after the run from the retained CLI stderr)
in total
not reported by the CLI
subscription: tokens only, no dollar figure (flat-rate plan, no per-call price is paid)
gpt-6-astra-codex
396 [396 to 397]
0 [0 to 0]
7923 [6952 to 9733] (CLI wall)
11199 [9575 to 13922] (CLI wall)
total 4917527 [4577486 to 5475167] (unsplit; recovered after the run from the retained CLI stderr)
in total
not reported by the CLI
subscription: tokens only, no dollar figure (flat-rate plan, no per-call price is paid)
claude-fable-5-1
not run
claude-opus-5
not run
claude-sonnet-5
not run
glm-5.2-general
not run
Subscription budget handling
Each Claude arm was pinned to one account slot for all its repeats (slot 4, the daily driver, and slot 1, the orchestrating session, were never used). The slot's five-hour and weekly usage was read before each run and every 25 generator calls; a five-hour window above 90% or a rate-limit error paused new attempts (5-minute polls) until the window reset. Every pause is listed. A run that could not finish within 12 hours because of limits stopped and is kept as incomplete with its attempt count; nothing was rerun.
arm
slot
five-hour % before each run
usage checks per run
pauses
pause total (s)
rate-limited calls per run
deadline kills per run
complete repeats
incomplete repeats
tools disabled on every call
claude-fable-5-1-cli
5
3, 49, 24
22 [21 to 22]
0
0
0 [0 to 0]
0 [0 to 0]
3
none
yes
claude-opus-5-cli
3
0, 9, 1
20 [20 to 20]
0
0
0 [0 to 0]
0 [0 to 0]
3
none
yes
claude-sonnet-5-cli
2
5, 10, 12
21 [21 to 21]
0
0
0 [0 to 0]
0 [0 to 0]
3
none
yes
gpt-5.6-sol-codex
n/a (Codex)
n/a, n/a, n/a
0 [0 to 0]
0
0
0 [0 to 0]
0 [0 to 0]
3
none
n/a
gpt-6-astra-codex
n/a (Codex)
n/a, n/a, n/a
0 [0 to 0]
0
0
0 [0 to 0]
0 [0 to 0]
3
none
n/a
Per-family outcome matrix
Each cell is the mean count across repeats of correct / silent wrong / task error attempts in that intent family, out of the family's attempts.
Per-family mean counts per arm (darker = larger share of the family's attempts). Columns per arm: correct, silent wrong, task error. Denominator: the family's attempts (shown in the tooltip).Table view
intent family
glm-5.2
glm-5.3
claude-fable-5-1-cli
claude-opus-5-cli
claude-sonnet-5-cli
gpt-5.6-sol-codex
gpt-6-astra-codex
activity_deadline_before_creation
0.0 / 0.0 / 5.0 of 5
1.3 / 0.7 / 3.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
4.0 / 0.0 / 0.3 of 5
3.7 / 0.0 / 1.3 of 5
0.3 / 0.0 / 4.7 of 5
activity_duration_missing_due_time
2.7 / 0.3 / 2.0 of 5
2.7 / 0.0 / 2.3 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
4.0 / 0.0 / 0.7 of 5
4.7 / 0.0 / 0.0 of 5
0.7 / 0.0 / 4.3 of 5
all_commission_unknown
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 1.3 / 3.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.7 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
booked_without_settled_history
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 1.0 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
campaign_and_web_referral
3.0 / 0.0 / 0.0 of 5
3.0 / 0.3 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
2.3 / 0.7 / 0.0 of 5
2.7 / 0.0 / 0.3 of 5
3.0 / 0.0 / 0.0 of 5
corporate_registration_without_trade
3.0 / 0.7 / 0.0 of 5
3.0 / 0.7 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
1.0 / 2.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
1.3 / 0.0 / 1.7 of 5
customer_referrals_with_reward_event
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
deadline_actual_date_disagreement
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
deal_assignment_override
1.7 / 2.0 / 0.7 of 5
0.7 / 1.0 / 3.0 of 5
3.7 / 1.0 / 0.0 of 5
2.3 / 2.7 / 0.0 of 5
1.7 / 2.0 / 1.3 of 5
0.0 / 4.7 / 0.3 of 5
0.0 / 0.0 / 5.0 of 5
deal_person_organization_conflict
0.7 / 0.7 / 3.7 of 5
0.7 / 0.3 / 4.0 of 5
1.0 / 0.3 / 3.7 of 5
1.3 / 0.0 / 3.7 of 5
0.0 / 0.7 / 4.0 of 5
0.3 / 0.3 / 3.7 of 5
0.0 / 0.0 / 5.0 of 5
deal_stage_timestamp_order
0.7 / 0.0 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.0 / 4.0 of 5
2.0 / 0.0 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
deleted_booking_live_account
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.3 / 3.7 of 5
0.0 / 1.3 / 3.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
different_sales_ownership
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.7 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
first_booked_before_first_broker
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 1.3 / 3.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 3.7 of 5
0.0 / 0.0 / 5.0 of 5
funnel_incomplete_contact
0.0 / 4.7 / 0.0 of 5
0.3 / 4.7 / 0.0 of 5
1.3 / 3.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
2.3 / 1.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
fx_missing_on_settled
0.0 / 0.3 / 4.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 1.0 / 4.0 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 0.3 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
invoice_before_settlement
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
known_external_person_qualification
0.0 / 2.7 / 2.3 of 5
0.0 / 3.3 / 1.7 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 3.3 / 1.7 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 3.0 / 2.0 of 5
known_zero_booked_commission
1.7 / 0.3 / 0.0 of 5
1.3 / 1.7 / 0.0 of 5
1.3 / 0.0 / 0.0 of 5
2.7 / 0.0 / 0.0 of 5
1.3 / 2.0 / 0.0 of 5
1.7 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
london_signup_month_edge
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 0.0 of 5
0.0 / 0.0 / 3.0 of 5
marketing_registration_gap
0.0 / 0.3 / 3.3 of 5
0.0 / 0.0 / 4.0 of 5
0.0 / 0.7 / 3.3 of 5
0.0 / 2.7 / 1.3 of 5
0.0 / 0.0 / 3.7 of 5
0.0 / 2.0 / 2.0 of 5
0.0 / 0.0 / 4.0 of 5
multi_broker_activation
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 1.7 / 3.3 of 5
0.0 / 1.0 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.7 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
negative_realized_commission
1.3 / 0.7 / 0.3 of 5
1.7 / 0.7 / 0.0 of 5
2.0 / 0.0 / 0.0 of 5
2.3 / 0.0 / 0.0 of 5
1.7 / 1.0 / 0.0 of 5
2.7 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
note_author_deal_owner_difference
0.0 / 1.0 / 4.0 of 5
0.0 / 0.3 / 4.3 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 4.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
notified_without_verification
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
org_contact_without_website
0.7 / 2.0 / 2.0 of 5
0.7 / 1.7 / 2.7 of 5
1.0 / 2.3 / 1.7 of 5
1.0 / 0.3 / 3.7 of 5
0.0 / 0.0 / 4.0 of 5
0.0 / 1.0 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
organization_client_type_multiple
0.0 / 0.0 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 3.3 of 5
0.0 / 0.3 / 4.3 of 5
0.0 / 0.0 / 4.0 of 5
0.0 / 2.0 / 2.7 of 5
0.0 / 0.0 / 5.0 of 5
organization_employee_data_gap
0.0 / 0.7 / 4.3 of 5
0.3 / 1.0 / 3.3 of 5
1.0 / 0.3 / 3.3 of 5
1.7 / 0.0 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
organization_erp_multiple
0.0 / 0.0 / 4.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.7 / 2.7 of 5
0.0 / 1.0 / 4.0 of 5
0.0 / 0.3 / 4.0 of 5
0.0 / 2.0 / 2.7 of 5
0.0 / 0.0 / 5.0 of 5
organization_hedging_multiple
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 3.0 of 5
0.3 / 0.7 / 3.7 of 5
0.0 / 0.3 / 4.0 of 5
0.0 / 2.0 / 2.7 of 5
0.0 / 0.0 / 5.0 of 5
organization_introducer_link
0.0 / 0.7 / 4.0 of 5
0.7 / 0.0 / 4.3 of 5
0.7 / 0.3 / 2.0 of 5
1.0 / 0.0 / 4.0 of 5
0.0 / 0.0 / 3.3 of 5
0.0 / 1.3 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
organization_wallet_share_unknown
0.0 / 0.3 / 4.0 of 5
0.0 / 0.0 / 4.3 of 5
0.0 / 0.0 / 3.0 of 5
1.0 / 0.0 / 4.0 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.7 / 3.7 of 5
0.0 / 0.0 / 5.0 of 5
outdated_still_activated
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 1.0 / 3.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
overdue_expected_close
1.3 / 0.0 / 3.7 of 5
1.3 / 0.3 / 3.3 of 5
0.7 / 1.7 / 2.7 of 5
2.0 / 0.0 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
2.0 / 0.0 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
partner_paid_and_unpaid_statements
1.3 / 2.3 / 0.0 of 5
2.7 / 1.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
0.7 / 4.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
2.3 / 0.0 / 0.7 of 5
person_communication_preferences
2.0 / 0.7 / 1.7 of 5
1.3 / 1.0 / 2.0 of 5
0.7 / 2.0 / 2.0 of 5
3.3 / 1.0 / 0.0 of 5
0.0 / 0.0 / 5.0 of 5
0.7 / 3.7 / 0.7 of 5
0.0 / 0.0 / 5.0 of 5
person_multi_frequency_declarations
0.0 / 0.3 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.3 of 5
0.0 / 0.0 / 4.7 of 5
0.0 / 0.3 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
person_note_pinned_exclusively
0.0 / 0.0 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 2.0 / 2.7 of 5
0.0 / 0.0 / 5.0 of 5
person_organization_owner_gap
0.7 / 1.3 / 3.0 of 5
0.0 / 1.0 / 4.0 of 5
0.0 / 0.0 / 4.7 of 5
2.0 / 0.0 / 3.0 of 5
0.0 / 0.7 / 4.0 of 5
0.0 / 0.0 / 4.7 of 5
0.0 / 0.0 / 5.0 of 5
prebillable_eligible_stats
0.0 / 0.3 / 4.7 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 0.7 / 3.3 of 5
0.0 / 0.0 / 5.0 of 5
private_declared_currency_overlap
0.0 / 0.3 / 4.7 of 5
0.0 / 0.3 / 4.7 of 5
0.3 / 0.0 / 4.7 of 5
0.3 / 0.7 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
professional_zero_rate_referrals
0.0 / 4.7 / 0.3 of 5
0.0 / 4.0 / 1.0 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 5.0 / 0.0 of 5
0.3 / 1.7 / 3.0 of 5
0.0 / 5.0 / 0.0 of 5
0.0 / 0.7 / 4.3 of 5
qualified_without_deal_stage
1.3 / 0.0 / 3.7 of 5
0.3 / 0.0 / 4.3 of 5
2.0 / 0.0 / 3.0 of 5
2.0 / 0.0 / 3.0 of 5
0.7 / 0.0 / 3.7 of 5
1.7 / 0.3 / 3.0 of 5
0.0 / 0.0 / 5.0 of 5
quote_conversion_valid_link
0.3 / 1.3 / 3.3 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 2.3 / 2.7 of 5
0.0 / 3.0 / 2.0 of 5
0.0 / 2.3 / 2.7 of 5
0.0 / 1.3 / 3.0 of 5
1.0 / 0.0 / 4.0 of 5
quote_days_before_activation
0.7 / 0.3 / 4.0 of 5
1.0 / 0.0 / 4.0 of 5
1.0 / 0.7 / 3.3 of 5
2.0 / 0.0 / 3.0 of 5
1.0 / 0.0 / 4.0 of 5
1.0 / 0.0 / 3.3 of 5
1.0 / 0.0 / 4.0 of 5
quote_only_system_activity
1.0 / 0.7 / 3.3 of 5
0.3 / 0.3 / 3.7 of 5
0.3 / 2.0 / 2.0 of 5
0.0 / 1.7 / 2.3 of 5
1.0 / 1.3 / 2.7 of 5
1.0 / 0.0 / 3.7 of 5
1.0 / 0.0 / 4.0 of 5
trade_sell_not_declared
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.3 / 4.7 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 2.0 / 3.0 of 5
0.0 / 0.0 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
unknown_sti_with_broker
0.0 / 1.3 / 3.7 of 5
0.0 / 1.0 / 4.0 of 5
0.0 / 1.3 / 3.7 of 5
0.0 / 0.7 / 4.3 of 5
0.0 / 0.0 / 5.0 of 5
0.0 / 0.0 / 4.0 of 5
0.0 / 0.0 / 5.0 of 5
verified_after_registration
4.7 / 0.3 / 0.0 of 5
4.0 / 1.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
5.0 / 0.0 / 0.0 of 5
verified_private_rejected_registration
1.7 / 2.3 / 0.0 of 5
2.3 / 0.7 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
3.0 / 0.0 / 0.0 of 5
Method
Code: revision e528a201e05c on branch insights/generator-benchmark-b2; application manifest eabe3ef0b914aa43 (1690 files), frozen from Git blobs into an isolated snapshot before any run. Difference from the inherited Goal B snapshot (ff3e669e469d): docs/bi/benchmark-arms.json, docs/bi/benchmark.md, src/institutional_kb/eval_oracle/bench_cli_transport.py, src/institutional_kb/eval_oracle/bench_metrics.py, src/institutional_kb/eval_oracle/bench_providers.py, src/institutional_kb/eval_oracle/bench_subscription_budget.py; the frozen prompts, routing and guards do not import these files.
Fixture: insights_goal3_fixture (frozen goal-3 fixture, data hash 986553e1a09dc997); control-39 correction not applied.
Scorer hash 56e04e173b388e87; statistics 102ce65f618557db; evaluation runner 9a3c85ff027381cc (the runner gained an optional budget gate that only waits before submitting an attempt; scoring is untouched).
Clock: 2026-09-06T12:00:00Z through libfaketime 0.9.13, Europe/London calendar, real monotonic time.
Protocol as frozen: 3 repeats per arm, 2 workers, arms and repeats sequential, sampling requested temperature 0 and seed 42 where the provider accepts it, thinking disabled where the provider allows it (lowest effort otherwise). Per-call provider deadline: glm-5.2 25 s, glm-5.3 25 s, claude-fable-5-1 25 s, claude-opus-5 25 s, claude-sonnet-5 25 s, claude-fable-5-1-cli 75 s, claude-opus-5-cli 75 s, claude-sonnet-5-cli 75 s, gpt-5.6-sol-codex 75 s, glm-5.2-general 25 s, gpt-6-astra-codex 75 s; job deadline 180 s for every arm. Nothing was rerun to improve a number.
Method files re-hashed before every run: bench runtime 99888dd40c4c368b, adapters b4920151932ef18a, arm registry ab6d93300793cfc3.
Frozen at 2026-09-07T21:47:14.779180+00:00 on devbox (16 CPUs, load [2.94, 1.34, 0.73]). Other interactive sessions ran on the host; load averages are in every run receipt.
CLI transport: each CLI call used a child process, file-backed stdin, stdout and stderr (no pipes), an absolute per-call deadline with SIGKILL to the process group and reap, and an empty working directory. Claude calls disabled tools (proven per call: empty tool list, no tool-use block, canary not echoed) and pinned the account slot inside the child; account-slot usage checks and subscription-window monitoring applied to the Claude arms only. Codex calls used the CLI's read-only sandbox, which permits command execution and does not confine filesystem reads; command use was not counted. Claude CLI flags: 2.1.258 (Claude Code) with -p --output-format stream-json --tools "" --safe-mode --strict-mcp-config --no-session-persistence --system-prompt-file <app system prompt>, user prompt on stdin, account pinned by ccuse <slot> inside the child (a failed pin exits before the CLI starts). Codex CLI: codex-cli 0.153.4 exec -s read-only --ephemeral -m <model> with the control path's single-message framing on stdin and reasoning effort low.
Subscription budget: usage read before each run and every 25 calls; pause above 90% of the five-hour window or on a rate-limit error, polled every 300 s; run cap 12 h; permitted slots [2, 3, 5]; slots never used [4, 1].
Execution as run (deviation from the frozen protocol): The frozen protocol says arms run sequentially. Execution deviated. The sequential driver started at 21:48:12 UTC and ran the claude-fable-5-1-cli smoke and repeat 1. At 22:21:37 UTC its loop was stopped while that repeat 1 kept running (it finished at 22:30:21). Per-arm drivers were then launched on a staggered schedule: claude-opus-5-cli 22:21:45, claude-sonnet-5-cli 22:23:02, gpt-5.6-sol-codex 22:24:19 (crashed in its smoke, relaunched 22:28:21), gpt-6-astra-codex 22:25:36, and claude-fable-5-1-cli for repeats 2 and 3 at 22:31:23 once repeat 1's receipt existed. Up to five CLI arms overlapped, with repeats sequential within each arm; the last repeat finished at 00:24:59 UTC. All fifteen B2 repeats used the B2 freeze (same manifest, snapshot, fixture, corpus, scorer and clocks); the inherited control and glm-5.3 records keep their Goal B freeze and ran alone on an otherwise idle host (1-minute load 0.4 to 1.2 at start and end). Why: Elapsed time. The whole B2 execution took about 2 hours 37 minutes including the initial sequential phase; a fully sequential schedule was estimated at roughly eight hours for fifteen repeats (about 25 to 40 minutes each).
Noise floor frozen at 2026-09-07T18:11:04.085909+00:00 before any candidate arm ran.
Unverified
claude-fable-5-1 did not run: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
claude-opus-5 did not run: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
claude-sonnet-5 did not run: No ANTHROPIC_API_KEY in the private credential setup (.envrc.local, process environment, ~/.config). Claude Code setup and OAuth tokens are not API keys and were not used.
glm-5.2-general did not run: Provider probe failed: RateLimitError: Error code: 429 - {'error': {'code': '1113', 'message': 'Insufficient balance or no resource package. Please recharge.'}}
Whether a real glm-5.2 is still served anywhere could not be established: the general zai endpoint refuses this account and every reachable path reports glm-5.3. Goal-3 and Goal A runs recorded only the requested model.
Thinking control as observed in the traces: glm-5.2 (thinking disabled requested): about 59,354 reasoning tokens per run, about 122 per call; glm-5.3 (thinking disabled requested): about 60,266 reasoning tokens per run, about 123 per call; claude-fable-5-1-cli (adaptive_low): 141,052 thinking tokens per run, about 269 per call; claude-opus-5-cli (disabled): 0 thinking tokens per run, about 0 per call; claude-sonnet-5-cli (disabled): 0 thinking tokens per run, about 0 per call. The zai coding endpoint accepts thinking-disabled but still reports reasoning tokens for glm-5.3; those tokens are counted in cost and latency.
Temperature 0 and seed 42 were sent and accepted by zai; the control's own run-to-run spread shows they do not make the provider deterministic. The CLI arms have no sampling control at all.
Live latency and cost are advisory: other sessions ran on the host, list prices are the published rates on the day, and the CLI arms' wall time includes the CLI's own start-up and request pipeline.
The corrected control-39 fixture was not used; five holdout expected answers change under it.
Isolation: Claude tools were disabled and proven per call. Codex retained command capability under its read-only sandbox, and an empty working directory does not confine filesystem reads (the child inherited HOME), so isolation of the Codex arms from the repository, the sealed files and other arms' outputs is not proven; the retained Codex transcripts show no command or MCP-use markers, and Codex network and MCP isolation is not established by the CLI banner.
Output limits: the API adapters forward the application's max_tokens; both CLI adapters only record the requested value (requested_max_tokens in the traces) and do not forward it, so equivalent output-token limits across API and CLI arms were not verified.
Claude Code's print mode replaced its own system prompt with the application's (--system-prompt-file) and had no tools, project files, MCP servers or session, but it still ran its own request pipeline: the model saw the application's prompts through Claude Code, not through a bare API request.
Codex has no system-prompt control, so the prompt was framed as in the control path (one message) while the CLI kept its own instructions. Codex printed an unsplit token total on stderr for every generator call; the frozen collector searched stdout only, so the traces omit it and the totals shown were recovered afterwards from the retained stderr transcripts (input, output and reasoning splits remain unavailable). No price is reported. Its reported model id is the CLI banner's echo of the request, not a server-side identifier.
CLI arms ran with a 75 s per-call deadline (API arms 25 s) because the child process carries transport overhead; deadline kills are counted as provider failures and listed under budget handling. CLI calls whose wall time exceeded the 25 s API deadline (they completed within 75 s): claude-fable-5-1-cli 1, claude-opus-5-cli 0, claude-sonnet-5-cli 1, gpt-5.6-sol-codex 7, gpt-6-astra-codex 0; an equal-deadline comparison was not run.
Auxiliary calls: after the SQL step the application makes optional presentation calls with their own 25 s budget; their traces (auxiliary-provider-trace.jsonl) recorded cancellations (JobStoppedError) per arm: glm-5.2 10, glm-5.3 4, claude-fable-5-1-cli 14, claude-opus-5-cli 14, claude-sonnet-5-cli 12, gpt-5.6-sol-codex 8, gpt-6-astra-codex 6. Their relationship to concurrency was not established; they are not SQL-generator failures.
Concurrency: Quality metrics are compared on identical scoring inputs (scorer, corpus, clocks, fixture). Within the SQL-generator path B2 recorded zero provider-call failures, zero 75 s CLI deadline kills and zero 180 s job-deadline expirations. The auxiliary-generation traces (presentation calls made by the application after the SQL step, with their own 25 s budget) separately recorded 54 optional-stage cancellations (JobStoppedError: fable 14, opus 14, sonnet 12, sol 8, astra 6; the inherited control recorded 10 and glm-5.3 4 under sequential execution); their relationship to concurrency was not established. No contention effect on quality was identified, and equivalence to sequential execution was not tested. CLI wall time is advisory only: it was measured under up to five concurrent arms (1-minute load average up to 14.5 on 16 CPUs, against 0.4 to 1.2 for the inherited Goal B runs) and is not compared across arms or with the API arms. The effect of concurrent execution on quality remains unverified.
gpt-5.6-sol-codex: smoke incident at 2026-09-07T22:24:58 UTC, crashed before any question ran: RuntimeError: The server serves a different arm than requested (scripts/insights_bench.py admit()). Cause: port race: the per-arm drivers were started before commit fdc307fb (honour BENCH_PORT_START) and two of them chose the same runtime port, so the smoke's admission probe reached another arm's server. Resolution: the driver was relaunched at 22:28:21Z with its own port range (8260 onwards); its smoke passed at 22:31:09Z (5 of 5 questions functioning) and all three repeats completed. Smoke runs are not part of the frozen runs; no frozen repeat crashed and no repeat was restarted.
Labelling: Each repeat directory's run-manifest.json and every attempt_id carry 'repeat 1' because the bench passes the repeat index to the evaluation runner only through --label (goal-b-<arm>-r<n>) and the runner's own --repeat defaults to 1. The bench repeat index is in the directory name, the label and run-receipt.json. Scoring is unaffected; the same quirk is in the inherited Goal B records.
Evidence
docs/evidence/insights-goal-b2-2026-09-07/freeze-manifest.json: arms, hashes, probes (including the alias probes and the tools-disabled proof), protocol, inherited noise floor.
docs/evidence/insights-goal-b2-2026-09-07/noise-floor.json: per-case agreement and metric spread of the control repeats (copied from Goal B).
docs/evidence/insights-goal-b2-2026-09-07/arms/<arm>/repeat-<n>/: compact attempt records; provider traces (model identifier, exit code and wall time per call; Claude traces additionally carry token usage and tool availability and use; Codex token collection in the trace is incomplete and Codex tool use was not measured); auxiliary-provider-trace.jsonl (optional presentation calls); codex-tokens-recovered.json(l) (Codex totals parsed after the run from the retained stderr); run receipt with load averages and the budget summary; budget-pauses.jsonl and budget-usage.jsonl.
docs/evidence/insights-goal-b2-2026-09-07/results/<arm>.json and results/comparison.json: every number in this report; results/inventory.json: the run inventory.
docs/evidence/insights-goal-b2-2026-09-07/execution.json and logs/: how the runs were actually driven (sequential then per-arm drivers), the smoke incident, and the driver logs.
docs/evidence/insights-goal-b2-2026-09-07/smoke/<arm>/: five development questions per arm, retained separately.
docs/evidence/insights-goal-b-2026-09-07/: the inherited Goal B records for claude-fable-5-1, claude-opus-5, claude-sonnet-5, glm-5.2, glm-5.3 (control repeats, glm-5.3 repeats, noise floor).
docs/bi/benchmark.md: how to run one arm, add a model or a CLI arm, and render the report.