Pass rate
Tasks solved / tasks in scope, up to 2 attempts (strict per-set denominator).
Formula: (tasks_passed_attempt_1 + tasks_passed_attempt_2_only) / task_set_size, where both numerator terms are means across the runs.
Includes unattempted tasks as failures. Scope-aware; reflects active filters (set, category, difficulty). A model benched several times scores the average of its runs, so run count does not inflate it. Final assisted solve rate with up to 2 attempts; drill-down companion to Solve AUC@2.
Avg cost / task
Average LLM cost of running one benchmark task once, in USD.
Formula: SUM(cost_usd) / COUNT(DISTINCT (run_id, task_id)) across all the model's results in scope.
Use to compare operating cost across models with similar pass rates. Does not account for quality. Combine with $/Pass for a cost-efficiency view.
Latency p50
Median per-task wall time (LLM call + compile + test), in milliseconds.
Formula: 50th percentile of per-task duration_ms: LLM latency + compile time + test time.
Use p50 for a typical-case latency expectation. Unaffected by outlier slow tasks. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model.
Avg attempt score
Mean per-attempt score on a 0–100 point scale (partial credit). Drill-down only.
Formula: Mean of attempt scores across all results rows: SUM(score) / COUNT(*) over the results table. Each attempt earns 0–100 points based on compile + test outcomes.
Drill-down companion to pass_at_n. Rewards partial credit but not directly comparable to pass rate; use for within-model analysis.
All-runs pass rate
Fraction of tasks the model solved in every single run (strict consistency, also written pass^n).
Formula: tasks where ALL runs produced a passing result / tasks_attempted_distinct
Measures reliability under repetition. High value means the model is unlikely to regress on a re-run, important for CI and production use. Formal name in the literature: pass^n.
$/Pass
Average USD cost per solved task (any-attempt pass).
Formula: SUM(cost_usd) / number of passed (run, task) cells across all runs.
Best single cost-efficiency metric. Penalises expensive models that pass few tasks and rewards cheap models with high pass rates.
Latency p95
95th-percentile per-task wall time. Captures tail latency.
Formula: 95th percentile of per-task duration_ms across all tasks in all runs.
Use p95 to understand worst-case latency. A low p95 means the model rarely stalls, relevant for automated pipelines with timeouts. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model.
Overview
Claude Haiku 4 5 has run on 3 occasions, attempting 232 tasks with an average score of 57.6 / 100.
Settings
Generation parameters used across this model's runs. "varies" indicates the value differed between runs.
- Temperature
- varies
- Thinking budget
- varies
- Avg tokens / run (input + output)
- 1,464,574
- Consistency
- 88.8%
History
Cost
Failure modes
- AL0104 313 Syntax error, 'end' expected view all →
- AL0000 307 App generation failed view all →
- AL0132 287 'Record Customer' does not contain a definition for 'Preferred Contact Method' view all →
- AL0185 148 Interface 'INotificationChannel' is missing view all →
- AL0118 108 The name 'CreateSequentialGuid' does not exist in the current context. view all →
- AL0107 86 Syntax error, identifier expected. Provide a valid name (letters, digits, and underscores only). view all →
- AL0133 69 Argument 2: cannot convert from 'Text' to 'Boolean' view all →
- AL0111 65 Semicolon expected. Add a semicolon (;) to terminate the statement. view all →
- AL0126 65 No overload for method 'AddConstantValue' takes 3 arguments. Candidates: built-in method 'AddConstantValue(Joker, Integer)' view all →
- AL0105 42 Syntax error, identifier expected; 'key' is a keyword view all →
Shortcomings
AL concepts Claude Haiku 4 5 struggles with. Click a row for description, correct pattern, and observed error codes.
Shortcomings analysis queued
Queued for analysis. This section will populate once the run is processed.
Recent runs
| Started | Run | Model | Tasks | Score | Cost | Duration | Status |
|---|---|---|---|---|---|---|---|
| 9d ago | a85a82f8-784… | 123/232 | 57.6 / 100 | $3.22 | 1h 3m | completed | |
| 9d ago | 8c384b82-1d4… | 121/232 | 57.6 / 100 | $3.56 | 1h 4m | completed | |
| 9d ago | 1a73d36e-7b7… | 125/232 | 57.7 / 100 | $3.53 | 1h 4m | completed | |
| 9d ago | efcd43ae-089… Excluded | 105/232 | 46.0 / 100 | $3.46 | 56m 28s | completed | |
| 9d ago | 8e71584d-7ee… Excluded | 109/232 | 45.4 / 100 | $3.50 | 59m 35s | completed | |
| 9d ago | 20a0f409-7e1… Excluded | 129/231 | 59.6 / 100 | $3.23 | 1h 20m | completed |
Methodology
Pass rate counts unattempted tasks as failures (strict denominator). Avg score is the mean of every attempt the model produced — failed first tries that triggered a retry contribute one observation each, pulling the mean down. See the about page for the full breakdown and unit conventions.