Pass@N

Pass rate

Tasks solved / tasks in scope, up to 2 attempts (strict per-set denominator).

Formula: (tasks_passed_attempt_1 + tasks_passed_attempt_2_only) / task_set_size, where both numerator terms are means across the runs.

Includes unattempted tasks as failures. Scope-aware; reflects active filters (set, category, difficulty). A model benched several times scores the average of its runs, so run count does not inflate it. Final assisted solve rate with up to 2 attempts; drill-down companion to Solve AUC@2.

91.5%
Tasks pass 212.3/232
1st: 167 2nd: 45.3 Failed: 19.7
Avg cost / task

Avg cost / task

Average LLM cost of running one benchmark task once, in USD.

Formula: SUM(cost_usd) / COUNT(DISTINCT (run_id, task_id)) across all the model's results in scope.

Use to compare operating cost across models with similar pass rates. Does not account for quality. Combine with $/Pass for a cost-efficiency view.

$0.06
Latency p50

Latency p50

Median per-task wall time (LLM call + compile + test), in milliseconds.

Formula: 50th percentile of per-task duration_ms: LLM latency + compile time + test time.

Use p50 for a typical-case latency expectation. Unaffected by outlier slow tasks. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model.

0.0s
Avg score

Avg attempt score

Mean per-attempt score on a 0–100 point scale (partial credit). Drill-down only.

Formula: Mean of attempt scores across all results rows: SUM(score) / COUNT(*) over the results table. Each attempt earns 0–100 points based on compile + test outcomes.

Drill-down companion to pass_at_n. Rewards partial credit but not directly comparable to pass rate; use for within-model analysis.

83.2 / 100
All-runs pass rate

All-runs pass rate

Fraction of tasks the model solved in every single run (strict consistency, also written pass^n).

Formula: tasks where ALL runs produced a passing result / tasks_attempted_distinct

Measures reliability under repetition. High value means the model is unlikely to regress on a re-run, important for CI and production use. Formal name in the literature: pass^n.

85.8%
$/Pass

$/Pass

Average USD cost per solved task (any-attempt pass).

Formula: SUM(cost_usd) / number of passed (run, task) cells across all runs.

Best single cost-efficiency metric. Penalises expensive models that pass few tasks and rewards cheap models with high pass rates.

$0.0635
Latency p95

Latency p95

95th-percentile per-task wall time. Captures tail latency.

Formula: 95th percentile of per-task duration_ms across all tasks in all runs.

Use p95 to understand worst-case latency. A low p95 means the model rarely stalls, relevant for automated pipelines with timeouts. Shown as N/A for batch runs: a batch request waits in the provider queue, so no per-request model timing is recorded and the only duration left would be our own compile-and-test time, which describes the harness rather than the model.

0.0s

Overview

Gemini 3.8 Flash (OpenRouter) has run on 3 occasions, attempting 232 tasks with an average score of 83.2 / 100.

Settings

Generation parameters used across this model's runs. "varies" indicates the value differed between runs.

Temperature
varies
Thinking budget
varies
Avg tokens / run (input + output)
4,091,326
Consistency
78.0%

History

1
2
3
3 runs · oldest 7d ago · latest 7d ago

Cost

meanp95

Failure modes

  • AL0132 91 'FieldType' does not contain a definition for 'Enum' view all →
  • AL0000 88 App generation failed view all →
  • AL0104 86 Syntax error, ')' expected view all →
  • AL0107 34 Syntax error, identifier expected. Provide a valid name (letters, digits, and underscores only). view all →
  • AL0111 34 Semicolon expected. Add a semicolon (;) to terminate the statement. view all →
  • AL0105 31 Syntax error, identifier expected; 'key' is a keyword view all →
  • AL0360 24 Text literal was not properly terminated. Use the character ' to terminate the literal. view all →
  • AL0297 15 The application object identifier '50100' is not valid. It must be within the allowed ranges '[70000..89999]'. view all →
  • AL0135 11 There is no argument given that corresponds to the required formal parameter 'DaysInPeriod' of 'CalculateOrderFrequency(Integer, Integer)' view all →
  • AL0126 9 No overload for method 'Clear' takes 1 arguments. Candidates: 'Clear()' defined in Codeunit 'CG H054 Cache' by the extension CentralGauge_CG-AL-H054_1 by CentralGauge (1.0.0.0) view all →

Shortcomings

AL concepts Gemini 3.8 Flash (OpenRouter) struggles with. Click a row for description, correct pattern, and observed error codes.

Shortcomings analysis queued

Queued for analysis. This section will populate once the run is processed.

Recent runs

Runs
StartedRunModelTasksScoreCostDurationStatus
7d ago1f99aac6-4b2… 210/23282.8 / 100$13.3255m 56scompleted
7d ago53112fb9-651… 212/23282.5 / 100$13.7858m 10scompleted
7d ago9d1ac7e7-d75… 215/23284.3 / 100$13.3458m 10scompleted

See all 3 runs →

Methodology

Pass rate counts unattempted tasks as failures (strict denominator). Avg score is the mean of every attempt the model produced — failed first tries that triggered a retry contribute one observation each, pulling the mean down. See the about page for the full breakdown and unit conventions.