Published task benchmark

Interest-Based Math Problem Generation

Given a math skill, difficulty, and student interest, generate solvable word problems with verified answers and no irrelevant assumptions.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Education & EdTechGenerationEvaluated 9/9/2026Pricing data observed October 2, 2026

Have a neighboring workload?

Describe it in Task Router. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task