Model Routing: Haiku, Sonnet, Opus
Model Routing: Haiku, Sonnet, Opus
A feature run can spawn 50 agent invocations. All Opus = 20× cost, 5× latency, no marginal benefit on the easy ones. All Haiku = silent failures on the hard ones. Route each task to the smallest model that's correct.
Tiers
flowchart LR
Task --> R{What kind?}
R -->|lookup, grep, read| H[Haiku]
R -->|docs, tests, bulk edit, straight impl| S[Sonnet]
R -->|design, complex impl, review| O[Opus]
R -->|security, SEC-*| F[Fable]
| Tier | Best for | Worst at |
|---|---|---|
| Haiku | lookups, grep, mechanical reads, listings | judgement, design |
| Sonnet | well-specified impl, tests, docs, bulk edits | open-ended design |
| Opus | spec writing, plan/design, complex impl, adversarial review | throughput, cost |
| Fable | security review, threat modelling, SEC-* tasks |
cost; refuses much of the work it's pointed at — see below |
Use the next tier up when in doubt.
The Fable tier
@fable is the security tier: SEC-* tasks, threat modelling, the security dimension
of adversarial review. Model ID claude-fable-5. It is the top of the capability
ladder — the strongest reasoning available, thinking always on, effort dialable
low → max.
Three things to know before you route anything here.
It refuses a large share of security work by design. Fable 5 runs classifiers over
incoming requests targeting most cybersecurity content, and Anthropic states plainly
that the model is not intended for that domain. Its documented bug-finding gains
exclude security-focused analysis, which is precisely where these classifiers apply.
Benign security tooling trips them as a false positive. A refusal is not an error —
you get an HTTP 200 with stop_reason: "refusal" and stop_details.category: "cyber",
so a harness that reads content[0] without checking stop_reason first sees an empty
result rather than a failure.
Budget for it. $10 / $50 per MTok — 2× Opus 5's $5 / $25. A tier that both costs
double and rejects part of its own workload is a deliberate trade, not a free upgrade.
It requires 30-day data retention. Fable 5 is unavailable under zero data retention:
an org configured below 30 days gets 400 invalid_request_error on every Fable request,
valid payload or not. Check the org's retention setting before assuming a request is
malformed.
When a @fable task comes back refused
In the harness (@fable tags dispatched by /speckit.implement), there is no per-tag
fallback parameter — the tier tag is a routing suggestion the harness reads, nothing more.
The recovery is the one this page already prescribes for every tier: re-run the task at
@opus. Treat a cyber refusal as a tier mis-route, not a task failure.
In an API-driven pipeline, opt into server-side fallbacks so the refusal is re-served
in the same call rather than surfacing to you:
client.beta.messages.create(
model="claude-fable-5",
max_tokens=16000,
betas=["server-side-fallback-2026-07-01"],
fallbacks="default", # routes by refusal category; cyber → claude-opus-4-8
messages=[...],
)
fallbacks: "default" picks the recommended substitute per refusal category, so you never
maintain a model list. Pin one explicitly with the array form
(betas=["server-side-fallback-2026-06-01"], fallbacks=[{"model": "claude-opus-4-8"}])
— note the earlier beta date gates the array form; pairing either header with the other
form is a 400. Server-side fallbacks are Claude API only; on Bedrock, Vertex, or Foundry
register the SDK's client-side refusal-fallback middleware instead.
Either way the destination is Opus. Which is the honest summary of this tier: Fable is
the strongest model in the ladder and the weakest choice for the tasks routed to it, and
Opus catches what it drops.
Where it's encoded
Per phase — frontmatter on each command (model: opus). The harness reads it and routes the invocation.
Per task — @haiku / @sonnet / @opus / @fable tag on each line in tasks.md. /speckit.implement dispatches subagents at the suggested tier.
The tag is a routing suggestion, not a guarantee. Validate what comes back: if a @haiku task returns something that clearly needed judgement, re-run it a tier up rather than accepting it. Trusting tier tags blindly is how subtle semantic failures get through.
@fable needs one extra check the others don't: an empty or non-committal result may be a
cyber refusal rather than a finding of "nothing to report". On a security task, no
findings and declined to look are indistinguishable downstream — and they mean opposite
things. Confirm which one you got before you record a clean pass.
Worked example
Feature: add /health endpoint returning {status, deps[]}.
| Task | Tier | Why |
|---|---|---|
| Locate router setup | @haiku | Pure lookup |
| Find dependency-check helpers | @haiku | grep-style |
| Design the response shape | @opus | Design decision; affects callers |
| Write the model struct | @sonnet | Mechanical given the design |
| Implement the handler | @sonnet | Straight impl |
| Wire the route | @sonnet | Mechanical |
| Write the integration test | @sonnet | Behaviour fully specified |
Auth check on the endpoint (SEC-001) |
@fable | Security task |
| Adversarial review | @opus | Judgement-heavy |
Two Opus invocations, one Fable, six Sonnet/Haiku. ~15% of tokens at Opus tier. ~4× faster wall-clock than all-Opus, same correctness.
SEC-001 is the one to watch. If it comes back refused, re-run it at @opus — and note
that on a /health endpoint the security question ("should this leak dependency names to
an unauthenticated caller?") is exactly the kind of benign review that trips the cyber
classifiers as a false positive.
Heuristics
Output shape. Listing/extraction → Haiku. Specified function body → Sonnet. New design or cross-file → Opus.
Judgement required. None → Haiku. Bounded by the spec → Sonnet. About the spec, or spanning files → Opus.
Security-shaped. SEC-*, threat modelling, authz review → Fable, with @opus as the
standing fallback. Budget for a meaningful refusal rate rather than treating each one as an
incident.
Anti-patterns
- All Opus: cheap to do, expensive to run. Wall-clock balloons on trivial subtasks.
- All Haiku: cheap and wrong. Failures look like The One-Shot Problem — surface passes, semantics off.
- Trust the tag blindly: tags are hypotheses. Flaky @sonnet → bump to @opus. Trivial @opus → demote.
- Reading a
@fablerefusal as a clean bill of health: the worst failure on this page. A declined security review that gets logged as "no findings" is a silent gap with a green tick on it — checkstop_reasonbefore you believe the result. @fableas a general "hardest tasks" tier: it's scoped to security here. Non-security work that genuinely needs the top of the ladder is a separate routing question — don't let the tag drift into meaning "important".
Use the smallest model that's correct, not the biggest model you can afford.