CentralGauge
Benchmark for LLMs on Microsoft Dynamics 365 Business Central AL code.
Updated 5d ago 2 attempts/model 95% paired-bootstrap CI Solve AUC@2 = (pass@1 + solve@2) / 2
Best overall
Best value · AUC ≥ 75
81.8 AUC · $0.06/solved
Fastest ≥ 75 AUC
| # | Model | Solve AUC@2 Skill score: full credit for a first-try solve, half for a retry solve. Not the solve rate. Formula: Use as the headline ranking metric. Rewards first-try correctness over fail-then-repair without ignoring the two-attempt protocol. De-saturates the headline that pass_at_n compresses. Significance via paired bootstrap (tier bands), not Wilson. | CI Pass Rate 95% CI 95% Wilson confidence interval on the pass rate. Formula: Use to judge whether a lead over another model is statistically meaningful. Wide CIs indicate too few tasks to draw firm conclusions. | Avg cost / task Average LLM cost of running one benchmark task once, in USD. Formula: Use to compare operating cost across models with similar pass rates. Does not account for quality. Combine with $/Pass for a cost-efficiency view. | Latency p95 95th-percentile per-task wall time. Captures tail latency. Formula: Use p95 to understand worst-case latency. A low p95 means the model rarely stalls, relevant for automated pipelines with timeouts. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model. | Details |
|---|---|---|---|---|---|---|
| 1 | 81.8 upstream unrecorded | ±3.6 | $0.06 | N/A | ||
| 2 | 79.7 | ±3.9 | $0.06 | N/A | ||
| 3 | 79.3 ⊘10 | ±3.6 | $0.18 | N/A | ||
| 4 | 72.3 | ±4.7 | $0.04 | N/A | ||
| 5 | 68.9 | ±5.1 | $0.09 | N/A | ||
| 6 | 67.8 | ±5.4 | $0.004 | N/A | ||
| 7 | n=2 | 65.4 upstream unrecorded | ±5.7 | $0.19 | N/A | |
| 8 | 57.4 upstream unrecorded | ±6.1 | $0.01 | N/A | ||
| 9 | 47.2 | ±6.4 | $0.01 | N/A |
Showing 9 of 9