Classification Router
Classify each request into the configured taxonomy with strict structured output.
Compare measured model quality, cost and latency, or ask how routes, execution plans and Playground work. No account is required to ask.
Describe your workload → inspect the evidence → create an API route
See routing in action
Choose a task and set your quality target. See which models qualify, which one is selected, and the fallback order—all from published benchmark evidence.
1 · Your task
Return RMA Agent
2 · Your requirements
90% minimum quality · lowest cost
3 · Model selection
No qualifying model
4 · Stable API endpoint
taskrouter/<your-route>
Lowest-cost policy. Quality is measured on the published evaluation set; it is not a guarantee for every request.
No model selected
Lower the quality target, relax the latency limit, or choose another task. A route will only be created when a model qualifies.
No qualifying fallbacks available.
| Model | Quality | p50 | Cost | Selection |
|---|---|---|---|---|
| Nemotron Nano 9B v2 | 13.0% | 6.37s | $26.945 | Below quality target |
| Qwen3 235B A22B | 38.0% | 8.13s | $84.634 | Below quality target |
| DeepSeek V3 | 43.0% | 6.68s | $119.464 | Below quality target |
| Mistral Large 2407 | 24.0% | 23.63s | $992.40 | Below quality target |
| Arcee Trinity Large Thinking | 25.0% | 5.16s | $117.475 | Below quality target |
| GPT-5.4 | 67.0% | 8.73s | $1,469.00 | Below quality target |
| Claude Opus 4.7 | 73.0% | 16.46s | $4,176.70 | Below quality target |
| Gemini 3.1 Pro Preview | 67.0% | 12.78s | $1,912.40 | Below quality target |
| Gemma 4 E4B IT | 31.0% | 10.96s | $8.727 | Below quality target |
| Granite 4.1 8B | 15.0% | 8.38s | $8.965 | Below quality target |
| Ministral 8B Instruct 2410 | 13.0% | 24.52s | $9.9993 | Below quality target |
| Qwen3 4B Instruct 2507 | 18.0% | 70.59s | $4.7823 | Below quality target |
Evidence preview, not a live request. Route creation checks current availability; your API alias stays stable as qualified models change.
Search is the fastest way to find a fit. The library exposes every published benchmark for comparison, evidence, and discovery.
Showing 20 of 263 matching tasks
Classify each request into the configured taxonomy with strict structured output.
Produce a concise structured summary of a business conversation and its follow-up actions.
Extract schema-conformant structured fields from business documents.
Answer questions from supplied documents while staying grounded in the provided evidence.
Validate requester, role, resource, approvals, and least privilege before granting or rejecting access.
Verify permitted identity evidence and complete or escalate account recovery.
Analyzes meeting notes to assign a "Health Score" (Green/Yellow/Red).
Records why a decision was made to prevent "we already tried that" syndrome later.
Rebook or cancel flights while applying fare, baggage, loyalty, and disruption policies.
Given anonymous employee questions, classify topic, merge duplicates, preserve critical wording, and flag content requiring private follow-up.
Convert Suspicious Activity Report (SAR) data into a narrative report.
Flags transactions that deviate from historical patterns for internal audit.
Find feasible slots and book, reschedule, or cancel an appointment under policy constraints.
Given a customer contract, performance obligations, delivery events, and invoices, determine the proposed recognition schedule and flag contract-data conflicts.
Given synthetic business emails, classify likely privilege status, identify attorney involvement and legal-purpose signals, and route uncertain cases for review.
Reviews authentication logic for common flaws (e.g., hardcoded keys).
Given a supplier spreadsheet sample and a target product schema, map source columns, propose type conversions, and identify unmapped required fields.
Investigate failed backups, rerun permitted steps, and verify recovery-point integrity.
Given bank transactions and general-ledger entries, match reconciling items and explain outstanding deposits, fees, timing differences, and unexplained variance.
Verify proof-of-claim amounts against the debtor's schedule data.
Explain the work in your own words or start from a task we already benchmark.
Compare task-specific quality, cost, latency, and availability—not a generic leaderboard.
Review the benchmark method, sample, pricing source, and provider before following an outbound link.
Follow relevant price moves, new providers, benchmark results, and better models for your task.
The benchmark catalog helps you compare models. Task Router turns that evidence into an OpenAI-compatible route with a stable alias, fallbacks, API keys, and prepaid credits.
Describe a workload, compare measured model tradeoffs, verify the evidence, and follow an available provider.
Create a route that selects the lowest-cost model meeting your benchmark-quality target, then adapts as model quality, pricing, and availability change. Use organization API keys and prepaid credits for production access. $0.15 per 1,000 routing decisions, plus inference at the provider’s list price with no markup.
Create a stable task alias that selects the lowest-cost qualified model, keeps operational fallbacks, and can adapt as benchmark quality, pricing, and availability change.