Published task benchmark

Tool Selection and Argument Validation

Given conversation history and OpenAI-compatible tool definitions, call every required tool with valid supplied arguments or explain missing or invalid parameters without fabricating a call. Available model: neurometric/tool-choice.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Engineering & ProductGenerationEvaluated 9/9/2026Pricing data observed October 2, 2026

Have a neighboring workload?

Describe it in Task Router. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task