Most developers treat AI APIs like electricity — one provider, one rate, consumed without thought. The problem is that flagship models cost 40–100x more than fast lightweight models, and most tasks do not need flagship capability. If your application sends every prompt to Claude Sonnet or GPT-4o, you are almost certainly overpaying by a significant margin.
The Routing Principle
Intelligent cost management starts with a single discipline: match model to task complexity. Every prompt your application sends falls somewhere on a spectrum from trivial to demanding. Treating that spectrum as flat is where the money goes.
Tasks suited for fast, cheap models
- Simple classification and intent detection
- Short-form question answering with well-scoped context
- Summarization of structured or semi-structured text
- Single-turn conversational responses
- Code completion within a known pattern
Models like Claude Haiku 3.5, Gemini Flash 2.5, and Llama 3.3 70B handle these workloads reliably at a fraction of the cost. For these tasks, spending more buys you no measurable quality improvement.
Tasks that justify flagship models
- Multi-step reasoning across long or ambiguous context
- Full code review with architectural feedback
- Long-form generation requiring coherence over thousands of tokens
- Synthesis across conflicting sources
Here, models like Claude Sonnet 4.5 and DeepSeek R2 earn their premium. The quality delta is real and measurable. Reserve them for exactly these cases.
A Routing Decision Table
The table below gives you a concrete starting point. Costs are approximate and reflect current market pricing per one million tokens. Use this as a reference when assigning models to prompt categories in your codebase.
| Task Type | Recommended Model | Cost / 1M tokens | vs Flagship |
|---|---|---|---|
| Simple Q&A | Claude Haiku 3.5 | \$0.25 | 12x cheaper |
| Summarization | Gemini Flash 2.5 | \$0.075 | 40x cheaper |
| Code completion | Llama 3.3 70B | \$0.05 | 50x cheaper |
| Complex reasoning | DeepSeek R2 | \$0.27 | 11x cheaper |
| Full code review | Claude Sonnet 4.5 | \$3.00 | Baseline |
Yoosh Cards as a Routing Enforcement Layer
Knowing which model to use is the analytical step. Enforcing it in production is the operational one. Yoosh Cards give you a practical mechanism to lock that discipline in place at the infrastructure level, not just in your application code.
One card per model tier
Create a dedicated card for each tier of your routing strategy. A cheap-tier card covers fast models — Haiku, Flash, Llama. A flagship card covers Sonnet or GPT-4o. Separate cards mean separate spend visibility. You see immediately how much each tier is consuming, and you can verify your routing logic is actually working.
Monthly caps as a hard backstop
Set a monthly spend cap on each card. The cheap card should handle roughly 80% of your traffic volume. The flagship card cap should be proportionally smaller. When a routing bug sends expensive prompts to the wrong tier — and eventually it will — the cap prevents a billing surprise at the end of the month. The card stops spending before the damage compounds.
Routing logic in code is optimistic. Caps on cards are pessimistic. You need both. The card is the safety net below the code.
Your 3-Step Action Plan
Step 1 — Audit your current prompt mix
Pull your API logs for the last 30 days. Categorize each call by task type: Q&A, summarization, code, reasoning, generation. Assign each category to the routing table above. Calculate what you would have paid if those categories had used the recommended model. That delta is your savings opportunity.
Step 2 — Implement a router in your application layer
Add a lightweight routing function before every API call. It reads the prompt intent — from metadata your application already knows, or from a fast classifier — and selects the appropriate model. Start with hard rules based on prompt type. Refine with observed quality metrics over the following weeks.
Step 3 — Create tiered Yoosh Cards and set caps
Issue a cheap-tier card and a flagship card on Yoosh. Wire each card to the corresponding model calls in your router. Set monthly caps that reflect your expected traffic split — aim for 80% of volume on the cheap card. Review spend weekly for the first month to confirm the routing split matches your expectations.