We ran 200 prompts across four task categories through Claude Sonnet 4.5, GPT-4o, and Gemini 2.5 Pro. This is not a benchmark — benchmarks optimize for a single number. This is a routing guide. The question is not which model wins overall, but which model wins for your specific task type so you can make a defensible decision about where your traffic should go.

Methodology

200 prompts, four categories (coding, reasoning and math, creative writing, document summarization), 50 prompts each. Scored on three axes: accuracy, instruction-following precision, and output format adherence. We ran each prompt three times per model and averaged scores to reduce sampling variance. All models accessed at default temperature.

Important caveat: Model quality changes with every deployment. These results reflect performance as of September 2026. Treat the rankings as directional, not absolute — but the task-type patterns have been consistent across the six months we have been running this.

Coding (50 prompts)

Winner: Claude Sonnet 4.5

Claude produces cleaner diffs, follows existing code style more consistently, and handles multi-file context better than the alternatives. When we passed in a legacy TypeScript codebase and asked for a targeted refactor, Claude was the only model that touched only what was asked without introducing unrequested changes.

GPT-4o is a strong second. It is capable, but occasionally adds boilerplate that was not requested — extra comments, wrapper functions, imports that are not needed. Gemini 2.5 Pro is solid on greenfield code but struggles with large existing codebases where context about established patterns matters.

Reasoning and Math (50 prompts)

Winner: DeepSeek R2 (among extended set); Gemini 2.5 Pro (among the three main models)

We included DeepSeek R2 and o3 mini as reference points. Both significantly outperform all three main models on pure logic chains and formal math. If your use case is reasoning-heavy, neither Claude, GPT-4o, nor Gemini is the right primary choice — you should be on a dedicated reasoning model.

Among the three main models, Gemini 2.5 Pro edges ahead on structured multi-step problems. It handles nested conditionals and constraint satisfaction problems more reliably. Claude is close. GPT-4o is noticeably weaker on pure logic tasks, though it remains competitive on problems that blend reasoning with world knowledge.

Creative Writing (50 prompts)

Winner: Claude Sonnet 4.5

Claude is measurably better at tone consistency and stylistic nuance over long outputs. When given a voice reference and asked to continue a piece in that style, Claude maintained the register across 2,000 words. GPT-4o drifted after approximately 800 words. Gemini 2.5 Pro produced competent prose but defaulted to generic structure — three-act, five-paragraph — even when asked for something unconventional.

Claude's advantage in creative tasks is not raw fluency. It is the ability to hold a constraint — a character's voice, a tonal register, a structural rule — across a long output without reverting to defaults.

Document Summarization (50 prompts)

Winner: Gemini 2.5 Pro

Gemini's 1M token context window is a real operational advantage here, not a marketing number. When we passed in 200-page PDFs and 80,000-word technical specifications, Gemini handled them without chunking. Claude handles up to 200K tokens well but required document splitting for the largest inputs. GPT-4o struggled noticeably when documents exceeded 60K tokens — extraction accuracy dropped, and key details from later sections were frequently omitted.

Summary table

Model Best at Avoid for Input cost / 1M tok
Claude Sonnet 4.5 Coding Creative writing Very long docs \$3.00
GPT-4o Vision General Q&A Long-context docs, pure logic \$2.50
Gemini 2.5 Pro Summarization Long docs Creative voice consistency \$1.25
DeepSeek R2 Math Reasoning Creative tasks, vision \$0.27
Claude Haiku 3.5 Speed Simple tasks Complex reasoning, long output \$0.25

Routing rules we would recommend

On Yoosh, you can create one Card per routing tier, set a monthly SOL cap on each, and change one environment variable to start routing. The routing logic lives in your code. The budget enforcement lives in your Cards.