Neurometric · Task Intelligence

Tell us the job.
We’ll find the right model.

Compare measured model quality, cost and latency, or ask how routes, execution plans and Playground work. No account is required to ask.

Describe your workload → inspect the evidence → create an API route

Conversation
263 measured tasks3120 model resultsUp to 99.94% lower cost on an eligible measured taskQuality · cost · latency · availability

See routing in action

Your requirements choose the model.

Choose a task and set your quality target. See which models qualify, which one is selected, and the fallback order—all from published benchmark evidence.

  1. 1 · Your task

    Project Status Sync Agent

  2. 2 · Your requirements

    90% minimum quality · lowest cost

  3. 3 · Model selection

    No qualifying model

  4. 4 · Stable API endpoint

    taskrouter/<your-route>

Lowest-cost policy. Quality is measured on the published evaluation set; it is not a guarantee for every request.

No model selected

No model meets these requirements.

Lower the quality target, relax the latency limit, or choose another task. A route will only be created when a model qualifies.

No qualifying fallbacks available.

Measured inference cost for 100,000 project. Routing fees are additional.
ModelQualityp50CostSelection
Nemotron Nano 9B v212.0%11.84s$48.008Below quality target
Qwen3 235B A22B47.0%11.81s$141.878Below quality target
DeepSeek V343.0%9.88s$215.202Below quality target
Mistral Large 240741.0%27.21s$1,350.60Below quality target
Arcee Trinity Large Thinking20.0%5.90s$178.995Below quality target
GPT-5.463.0%10.92s$2,345.00Below quality target
Claude Opus 4.760.0%20.42s$6,111.60Below quality target
Gemini 3.1 Pro Preview2.0%15.43s$2,632.60Below quality target
Gemma 4 E4B IT32.0%16.83s$13.677Below quality target
Granite 4.1 8B17.0%14.61s$15.705Below quality target
Ministral 8B Instruct 241012.0%26.23s$13.2445Below quality target
Qwen3 4B Instruct 250719.0%78.27s$6.3722Below quality target

Evidence preview, not a live request. Route creation checks current availability; your API alias stays stable as qualified models change.

Existing task benchmarks

Browse measured task benchmarks.

Search is the fastest way to find a fit. The library exposes every published benchmark for comparison, evidence, and discovery.

Browse all benchmarks
Neurometric · task-optimizedImproved for the task through fine-tuning, prompting, or harnesses
SmallOff-the-shelf, under 100B parameters
Mid100B–1T parameters
FrontierFlagship general-purpose models

Showing 20 of 263 matching tasks

Classification Featured

Classification Router

Classify each request into the configured taxonomy with strict structured output.

Classification and routingPer request
Neurometric task-optimized
Hosted · Task-prompted
483.7 in · 55.4 out · 2.57s p50
$1.04 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0100 input · $0.1000 output /1M
100%
Best accuracy
Qwen3 4B Instruct
Hosted · Off-the-shelf
373.9 in · 62.6 out · 2.30s p50
$0.9999 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0100 input · $0.1000 output /1M
97%
Cheapest
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
GPT-5.4
Hosted · Off-the-shelf
293.1 in · 54.9 out · 1.33s p50
$156 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
100%
Best accuracy
Summarization Featured

Conversation Summary

Produce a concise structured summary of a business conversation and its follow-up actions.

Business operationsPer conversation
Neurometric task-optimized
Hosted · Task-prompted
435.133 in · 112.233 out · 4.21s p50
$1.56 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0100 input · $0.1000 output /1M
100%
Best accuracyCheapest
Granite 4.1 8B
Hosted · Off-the-shelf
343.6 in · 93.5 out · 1.39s p50
$2.65 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 10/2/2026
OpenRouter · $0.0500 input · $0.1000 output /1M
TrustedRouter · $0.0527 input · $0.1055 output /1M
0%
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
GPT-5.4
Hosted · Off-the-shelf
390.6 in · 100.633 out · 1.84s p50
$249 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
100%
Best accuracy
Extraction Featured

Document Structured Extraction

Extract schema-conformant structured fields from business documents.

Document processingPer document
Neurometric task-optimized
Hosted · Task-prompted
253.767 in · 75.09 out · 2.95s p50
$1.00 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0100 input · $0.1000 output /1M
98%
Best accuracyCheapest
Gemma 4 E4B IT
Hosted · Off-the-shelf
193.4 in · 82.9 out · 1.29s p50
$1.28 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0211 input · $0.1055 output /1M
20%
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
GPT-5.4
Hosted · Off-the-shelf
240.2 in · 70.943 out · 1.46s p50
$166 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
81%
Analysis Featured

Grounded Document QA

Answer questions from supplied documents while staying grounded in the provided evidence.

Document processingPer question
Neurometric task-optimized
Hosted · Task-prompted
485.16 in · 26.93 out · 1.70s p50
$0.7545 /100K
taskrouter-admin snapshot
View token rates · observed 9/25/2026
taskrouter-admin · $0.0100 input · $0.1000 output /1M
99%
Cheapest
Granite 4.1 8B
Hosted · Off-the-shelf
222 in · 27.2 out · 0.51s p50
$1.38 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 10/2/2026
OpenRouter · $0.0500 input · $0.1000 output /1M
TrustedRouter · $0.0527 input · $0.1055 output /1M
20%
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
GPT-5.4
Hosted · Off-the-shelf
451.733 in · 19.947 out · 1.02s p50
$143 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
100%
Best accuracy
Analysis

Access Request Agent

Validate requester, role, resource, approvals, and least privilege before granting or rejecting access.

IT and security operationsPer request
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
651 in · 748 out · 11.87s p50
$9.27 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
76%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
612 in · 898 out · 9.21s p50
$92.49 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.2200 input · $0.8800 output /1M
82%
GPT-5.4
Hosted · Off-the-shelf
575 in · 760 out · 7.75s p50
$1284 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
99%
Best accuracy
Analysis

Account Recovery Agent

Verify permitted identity evidence and complete or escalate account recovery.

Customer service and account operationsPer account
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
550 in · 634 out · 9.87s p50
$7.85 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
70%
DeepSeek V3
Hosted · Off-the-shelf
510 in · 647 out · 15.05s p50
$138 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
75%
GPT-5.4
Hosted · Off-the-shelf
482 in · 690 out · 6.92s p50
$1156 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
89%
Best accuracy
Analysis

Account-Pulse

Analyzes meeting notes to assign a "Health Score" (Green/Yellow/Red).

Customer SuccessPer account
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
540 in · 586 out · 9.33s p50
$7.32 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
93%
DeepSeek V3
Hosted · Off-the-shelf
505 in · 526 out · 4.57s p50
$118 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
97%
GPT-5.4
Hosted · Off-the-shelf
489 in · 605 out · 6.65s p50
$1030 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
99%
Best accuracy
Generation

ADR-Writer

Records why a decision was made to prevent "we already tried that" syndrome later.

People ManagementPer decision
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
365 in · 340 out · 5.22s p50
$4.36 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
98%
Best accuracy
Mistral Large 2407
123B · Off-the-shelf
399 in · 578 out · 15.22s p50
$640 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
97%
GPT-5.4
Hosted · Off-the-shelf
345 in · 408 out · 4.78s p50
$698 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
98%
Best accuracy
Analysis

Airline Itinerary Change Agent

Rebook or cancel flights while applying fare, baggage, loyalty, and disruption policies.

Customer service and account operationsPer itinerary
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
926 in · 986 out · 15.67s p50
$12.36 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
50%
Mistral Large 2407
123B · Off-the-shelf
1,019 in · 1,231 out · 36.28s p50
$1414 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
44%
GPT-5.4
Hosted · Off-the-shelf
817 in · 822 out · 8.54s p50
$1437 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
88%
Best accuracy
Classification

All-Hands Question Moderation

Given anonymous employee questions, classify topic, merge duplicates, preserve critical wording, and flag content requiring private follow-up.

HR & Internal OpsPer question
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Granite 4.1 8B
8B · Off-the-shelf
326 in · 105 out · 1.28s p50
$2.68 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 10/2/2026
OpenRouter · $0.0500 input · $0.1000 output /1M
TrustedRouter · $0.0527 input · $0.1055 output /1M
28%
Mistral Large 2407
123B · Off-the-shelf
376 in · 256 out · 6.24s p50
$343 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
45%
Claude Opus 4.7
Hosted · Off-the-shelf
553 in · 487 out · 5.66s p50
$1643 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $5.50 input · $27.50 output /1M
82%
Best accuracy
Generation

AML Narrative Writer

Convert Suspicious Activity Report (SAR) data into a narrative report.

Legal & CompliancePer case
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
1,448 in · 685 out · 11.14s p50
$10.28 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
81%
Mistral Large 2407
123B · Off-the-shelf
1,574 in · 879 out · 21.92s p50
$1263 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
89%
GPT-5.4
Hosted · Off-the-shelf
1,236 in · 833 out · 8.43s p50
$1559 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
95%
Best accuracy
Analysis

Anomaly-Sense

Flags transactions that deviate from historical patterns for internal audit.

Finance & AccountingPer transaction
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
309 in · 523 out · 8.11s p50
$6.17 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
63%
DeepSeek V3
Hosted · Off-the-shelf
271 in · 435 out · 4.53s p50
$88.80 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
88%
GPT-5.4
Hosted · Off-the-shelf
266 in · 533 out · 6.17s p50
$866 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
98%
Best accuracy
Analysis

Appointment Management Agent

Find feasible slots and book, reschedule, or cancel an appointment under policy constraints.

Customer service and account operationsPer appointment
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
685 in · 910 out · 14.85s p50
$11.05 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
25%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
652 in · 1,226 out · 18.13s p50
$122 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.2200 input · $0.8800 output /1M
37%
GPT-5.4
Hosted · Off-the-shelf
538 in · 631 out · 5.94s p50
$1081 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
76%
Best accuracy
Analysis

ASC 606 Revenue Recognition Review

Given a customer contract, performance obligations, delivery events, and invoices, determine the proposed recognition schedule and flag contract-data conflicts.

Finance & AccountingPer contract
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
1,024 in · 1,936 out · 32.44s p50
$22.59 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
5%
DeepSeek V3
Hosted · Off-the-shelf
905 in · 1,657 out · 14.13s p50
$331 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
30%
GPT-5.4
Hosted · Off-the-shelf
872 in · 1,997 out · 18.11s p50
$3214 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
28%
Classification

Attorney-Client Privilege Triage

Given synthetic business emails, classify likely privilege status, identify attorney involvement and legal-purpose signals, and route uncertain cases for review.

Legal & CompliancePer email
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
568 in · 30 out · 0.48s p50
$1.51 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
82%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
556 in · 55 out · 0.98s p50
$17.07 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.2200 input · $0.8800 output /1M
91%
Claude Opus 4.7
Hosted · Off-the-shelf
901 in · 86 out · 2.11s p50
$732 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $5.50 input · $27.50 output /1M
100%
Best accuracy
Analysis

Auth-Guard

Reviews authentication logic for common flaws (e.g., hardcoded keys).

EngineeringPer codebase
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
471 in · 918 out · 14.26s p50
$10.68 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
69%
DeepSeek V3
Hosted · Off-the-shelf
432 in · 1,056 out · 7.48s p50
$202 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
84%
GPT-5.4
Hosted · Off-the-shelf
418 in · 1,415 out · 13.82s p50
$2227 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
97%
Best accuracy
Analysis

B2B Catalog Schema Mapping

Given a supplier spreadsheet sample and a target product schema, map source columns, propose type conversions, and identify unmapped required fields.

E-commerce & LogisticsPer catalog
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
829 in · 1,690 out · 27.79s p50
$19.58 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
32%
DeepSeek V3
Hosted · Off-the-shelf
773 in · 1,470 out · 17.15s p50
$292 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.5800 input · $1.68 output /1M
56%
Claude Opus 4.7
Hosted · Off-the-shelf
1,225 in · 1,970 out · 20.51s p50
$6091 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $5.50 input · $27.50 output /1M
53%
Analysis

Backup Recovery Agent

Investigate failed backups, rerun permitted steps, and verify recovery-point integrity.

IT and security operationsPer backup
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
599 in · 919 out · 14.39s p50
$10.96 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
44%
Mistral Large 2407
123B · Off-the-shelf
630 in · 1,117 out · 29.55s p50
$1194 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
35%
GPT-5.4
Hosted · Off-the-shelf
498 in · 1,248 out · 13.34s p50
$1997 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
92%
Best accuracy
Analysis

Bank-to-GL Reconciliation Explanation

Given bank transactions and general-ledger entries, match reconciling items and explain outstanding deposits, fees, timing differences, and unexplained variance.

Finance & AccountingPer statement
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
827 in · 1,532 out · 24.56s p50
$17.91 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
29%
Mistral Large 2407
123B · Off-the-shelf
881 in · 1,408 out · 38.28s p50
$1532 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $3.00 input · $9.00 output /1M
44%
GPT-5.4
Hosted · Off-the-shelf
687 in · 1,434 out · 12.84s p50
$2323 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
65%
Analysis

Bankruptcy Proof of Claim Auditor

Verify proof-of-claim amounts against the debtor's schedule data.

Legal & CompliancePer claim
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Gemma 4 E4B IT
E4B · Off-the-shelf
326 in · 357 out · 5.53s p50
$4.45 /100K
Current price · TrustedRouter
View token rates · observed 10/2/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
96%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
309 in · 305 out · 3.69s p50
$33.64 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $0.2200 input · $0.8800 output /1M
100%
Best accuracy
GPT-5.4
Hosted · Off-the-shelf
286 in · 186 out · 2.32s p50
$351 /100K
LiteLLM snapshot
View token rates · observed 10/2/2026
LiteLLM · $2.50 input · $15.00 output /1M
99%

From question to evidence

A model recommendation you can verify.

01

Describe the task

Explain the work in your own words or start from a task we already benchmark.

02

Get the right model

Compare task-specific quality, cost, latency, and availability—not a generic leaderboard.

03

Check the evidence

Review the benchmark method, sample, pricing source, and provider before following an outbound link.

04

Track what changes

Follow relevant price moves, new providers, benchmark results, and better models for your task.

One task-based platform

Explore now. Route next.

The benchmark catalog helps you compare models. Task Router turns that evidence into an OpenAI-compatible route with a stable alias, fallbacks, API keys, and prepaid credits.

Free · available now

Benchmark catalog

Describe a workload, compare measured model tradeoffs, verify the evidence, and follow an available provider.

Production API · available now

Task Router

Create a route that selects the lowest-cost model meeting your benchmark-quality target, then adapts as model quality, pricing, and availability change. Use organization API keys and prepaid credits for production access. $0.15 per 1,000 routing decisions, plus inference at the provider’s list price with no markup.

Task Router · Production routes

Benchmarks show the evidence.
Task Router puts it to work.

Create a stable task alias that selects the lowest-cost qualified model, keeps operational fallbacks, and can adapt as benchmark quality, pricing, and availability change.