CentralGauge

Benchmark for LLMs on Microsoft Dynamics 365 Business Central AL code.

Updated 5d ago 2 attempts/model 95% paired-bootstrap CI Solve AUC@2 = (pass@1 + solve@2) / 2

Best overall

Gemini 3.8 Flash (OpenRouter) gemini google/gemini-3.8-flash
· 81.8

Best value · AUC ≥ 75

Gemini 3.8 Flash (OpenRouter) gemini google/gemini-3.8-flash

81.8 AUC · $0.06/solved

Fastest ≥ 75 AUC

Best open-weight

65.4 AUC

Leaderboard
#Model

Solve AUC@2

Skill score: full credit for a first-try solve, half for a retry solve. Not the solve rate.

Formula: AUC@2 = (pass@1 + solve@2) / 2

Use as the headline ranking metric. Rewards first-try correctness over fail-then-repair without ignoring the two-attempt protocol. De-saturates the headline that pass_at_n compresses. Significance via paired bootstrap (tier bands), not Wilson.

CI

Pass Rate 95% CI

95% Wilson confidence interval on the pass rate.

Formula: Wilson score interval: center ± half-width, where n = strict denominator (task_set_size or category/difficulty-scoped count when taskSetHash is provided; falls back to tasks_attempted_distinct for legacy callers).

Use to judge whether a lead over another model is statistically meaningful. Wide CIs indicate too few tasks to draw firm conclusions.

Avg cost / task

Average LLM cost of running one benchmark task once, in USD.

Formula: SUM(cost_usd) / COUNT(DISTINCT (run_id, task_id)) across all the model's results in scope.

Use to compare operating cost across models with similar pass rates. Does not account for quality. Combine with $/Pass for a cost-efficiency view.

Latency p95

95th-percentile per-task wall time. Captures tail latency.

Formula: 95th percentile of per-task duration_ms across all tasks in all runs.

Use p95 to understand worst-case latency. A low p95 means the model rarely stalls, relevant for automated pipelines with timeouts. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model.

Details
1
Gemini 3.8 Flash (OpenRouter) gemini google/gemini-3.8-flash
81.8 upstream unrecorded ±3.6$0.06N/A
2
GPT-5.6 Sol gpt gpt-5.6-sol
79.7 ±3.9$0.06N/A
3
Claude Opus 5 claude claude-opus-5
79.3 ⊘10 ±3.6$0.18N/A
4
GPT-5.6 Terra gpt gpt-5.6-terra
72.3 ±4.7$0.04N/A
5
Claude Sonnet 5 claude claude-sonnet-5
68.9 ±5.1$0.09N/A
6
GPT-5.6 Luna gpt gpt-5.6-luna
67.8 ±5.4$0.004N/A
7 n=265.4 upstream unrecorded ±5.7$0.19N/A
8
MiniMax M3 minimax minimax/minimax-m3
57.4 upstream unrecorded ±6.1$0.01N/A
9
Claude Haiku 4 5 claude claude-haiku-4-5
47.2 ±6.4$0.01N/A

Showing 9 of 9