Model evidence
Good enough for which task?
Explore measured models and available catalog listings. Quality comes from published task evidence; models without task results are clearly marked.
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
Model evidence
Explore measured models and available catalog listings. Quality comes from published task evidence; models without task results are clearly marked.
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
| Model | Evidence | Cheapest passing | Median quality | Providers |
|---|---|---|---|---|
| openai/o3:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/o4-mini Unclassified | Not yet measured | — | — | OpenRouterTrustedRouter |
| openai/o4-mini-deep-research Unclassified | Not yet measured | — | — | OpenRouter |
| openai/o4-mini-high Unclassified | Not yet measured | — | — | OpenRouter |
| openai/o4-mini-high:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/o4-mini:batch Unclassified | Not yet measured | — | — | OpenRouter |
| openai/sora-2 Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/sora-2-pro Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/text-embedding-3-large Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/text-embedding-3-small Unclassified | Not yet measured | — | — | TrustedRouter |
| openai/text-embedding-ada-002 Unclassified | Not yet measured | — | — | TrustedRouter |
| openbmb/MiniCPM-V-4_5 Unclassified | Not yet measured | — | — | TrustedRouter |
| openjev/structured-decisions Unclassified | Not yet measured | — | — | Price unavailable |
| openpipe/qwen3-14b-instruct Unclassified | Not yet measured | — | — | TrustedRouter |
| openrouter/free Unclassified | Not yet measured | — | — | OpenRouter |
| openrouter/owl-alpha Unclassified | Not yet measured | — | — | OpenRouter |
| paddlepaddle/paddleocr-vl Unclassified | Not yet measured | — | — | TrustedRouter |
| parasail/liberty-2.0 Unclassified | Not yet measured | — | — | TrustedRouter |
| pearl-ai/gemma-4-31b-it Unclassified | Not yet measured | — | — | TrustedRouter |
| perceptron/perceptron-mk1 Unclassified | Not yet measured | — | — | OpenRouter |
Cheapest passing counts tasks where the model clears the task’s published quality floor (80% when none is recorded) and has the lowest observed workload cost. Median quality summarizes those published task results; it is not a general model ranking.