Cost / quality frontier
Campaign frontier-calibration · 6 models · 200 rollouts each · scored on public and hidden tests
● marks a model on the Pareto frontier, meaning no other model is both cheaper and better at the same time. Dimmed points are dominated by something above and to the left of them. Click a point for the per-model breakdown.
Leaderboard
pass@1 clean is measured on the public tests the solver can see. pass@1 adv adds the hidden and adversary-mined tests. The gap between the two numbers is how much a model overfits to what it can see.
| model | pass@1 clean | pass@1 adv | gap | p50 | p95 | $/attempt | $/solved |
|---|---|---|---|---|---|---|---|
| ●claude-opus-5 | 94% | 89% | 5% | 2400 ms | 5200 ms | $0.00950 | $0.0107 |
| ●claude-sonnet-5 | 90% | 83% | 7% | 1500 ms | 3400 ms | $0.00380 | $0.0046 |
| ●claude-haiku-4-5 | 78% | 66% | 12% | 700 ms | 1600 ms | $0.00070 | $0.0011 |
| gpt-oss-20b | 69% | 61% | 8% | 950 ms | 2200 ms | $0.00085 | $0.0014 |
| ●qwen3.5-27b | 72% | 58% | 14% | 1100 ms | 2600 ms | $0.00045 | $0.0008 |
| ●qwen3.5-9b | 55% | 40% | 15% | 600 ms | 1400 ms | $0.00015 | $0.0004 |