Model evidence
Good enough for which task?
Explore measured models and available catalog listings. Quality comes from published task evidence; models without task results are clearly marked.
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
Model evidence
Explore measured models and available catalog listings. Quality comes from published task evidence; models without task results are clearly marked.
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
| openai/gpt-6-astra:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-luna Unclassified | Not yet measured | — | — | OpenRouterTrustedRouter |
| openai/gpt-6-luna-pro Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-luna-pro:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-luna:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-sol Unclassified | Not yet measured | — | — | OpenRouterTrustedRouter |
| openai/gpt-6-sol-codex Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/gpt-6-sol-pro Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-sol-pro:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6-sol:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6.1-sol Unclassified | Not yet measured | — | — | OpenRouterTrustedRouter |
| openai/gpt-6.1-sol-pro Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6.1-sol-pro:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-6.1-sol:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-audio Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-audio-mini Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-chat-latest Unclassified | Not yet measured | — | — | OpenRouter |
| openai/gpt-image-2.5-flare Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/gpt-image-2.5-sunburst Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/gpt-oss-120b Unclassified | Not yet measured | — | — | OpenRouterTrustedRouter |
Cheapest passing counts tasks where the model clears the task’s published quality floor (80% when none is recorded) and has the lowest observed workload cost. Median quality summarizes those published task results; it is not a general model ranking.