I have the $200 Claude Max plan. I use Claude for my business and pretty much everything business-related, and when I start my new job I'm planning to drop to the $100 tier, because I won't be in it all day every day anymore.
$100 a month is still a lot of money.
And I kept catching myself using Fable, the most expensive model, for everything. Strategy work, sure. But also brainstorming, routine drafts, things a cheaper model could probably handle. Every time, the same two questions: do I actually need Fable for this? And when DO I need it?
Claude doesn't tell you. You get a model picker (Haiku, Sonnet, Opus, Fable) and an effort dial (low up through max), the price differences are real, and nobody explains how to choose. Not the product, and not really anyone else either. It's a gamble.
So I did some research, put what I found into a system, and turned the system into a Claude Skill so Claude makes the call for me. The full skill is at the bottom if you want to copy it.
Quality at a good value
The way I think about it now is the way I shop. When I buy a nice piece of clothing, I want high quality at a good value. I'm not hunting through the cheapest store for the cheapest version, but there's also no reason to spend $10,000 on a shirt (not that I have $10,000 to spend on a shirt anyway, but I digress). That's the goal with models too: give the task all the intelligence it needs, and if you'd get the same result with fewer tokens, pare it down.
One idea from the research changed how I look at the cheap option, so it goes first.
Total cost = what the model burns + the time you spend fixing what it gave you.
Because your time has a value too. How much is your time worth to you? Probably a lot. You have a job, or a job search. Kids, a dog, friends. There are better ways to spend an evening than fixing AI drafts.
Say Sonnet writes you an acceptable article draft and you spend 25 minutes editing it. Fable costs more to run, writes a stronger draft, and you spend five minutes. On tokens, Sonnet won. On getting to a finished article, Fable did, because your 20 minutes are worth more than the token difference.
It cuts the other way too. For a routine summary, both models hand you the same thing, so the cheap one wins outright. The comparison is always cost to finished work, not cost of the first response.
How to think about the models
Not bad, better, best. More like: which jobs each one is for.
Sonnet is the starting point. Clearly specified, routine work: summaries, standard emails, rewrites, routine content, simple research questions, high-volume generation. Sonnet is very capable. The reason to start here is value. If a bigger model costs a lot more and produces basically the same output, why pay for it?
Opus is for when reasoning matters. Competitive analysis, synthesizing research, difficult troubleshooting, decisions where the first obvious answer isn't enough. The question to ask: would more reasoning power materially change the quality of this answer? Important doesn't automatically mean yes. Plenty of important tasks are easy tasks.
Fable is for when the ceiling is worth it. The truly hard reasoning, long difficult workflows, decisions you'll act on without double-checking. And one category that surprised me: writing that has to sound like you. More on that in a second.
Effort is a second bill
Model choice is half the spend. The effort dial is the other half, and it's easy to set it high once and never think about it again.
The short version: low or medium for routine work where the path to the answer is obvious. High as the default for real work (writing, analysis, synthesis). Xhigh and max only when deeper reasoning is actually likely to change the answer: very hard problems, lots of interacting constraints, something that already failed at high.
The rule that helped me most: important and difficult are not the same thing. A deliverable can matter a lot to your business and still be an easy task. Max effort on an easy task buys you a bigger bill and a slower answer, not a better one.
How to think about writing
Before you route a writing task, ask what the writing has to accomplish. Is it for a search engine or for a human? Does the reader need to act on it, or does it just need to move information from point A to point B? Does it have to sound like you, or does it just have to be clear?
Commodity writing is the everyday stuff that doesn't have to be beautiful. It just has to make the point. SEO pages, product descriptions, summaries, standard emails, turning one article into five posts. Clear and correct is the whole job, and the cheap model and the expensive model land in basically the same place. Sonnet.
Authored writing is for humans, and it has a job to do. The reader needs to be engaged, or convinced, or moved to act, and the piece has to sound like YOU. I write with custom Claude Skills loaded with my voice rules, examples, and banned phrases, and following all of that well means the model is juggling a lot at once: inferring the principles behind the examples, resolving conflicts between rules, structuring the argument, and not reading like it worked through a checklist.
Honestly, I still think most of my writing needs Fable. The same writing skills produce noticeably better drafts there than on Sonnet. Less correcting, less rewriting. I can't point you to a benchmark for that. It's just what keeps happening in my work, and the editing time I get back covers the extra cost.
One related lesson: a skill is instructions, not capability. The model still has to interpret everything in it. When a smaller model follows a complicated skill badly, the fix isn't always a longer skill. Sometimes it's the same skill on a stronger model.
The decision, as questions
This is the framework I put together from the research. Ask in order, stop at the first yes.
- Is it mechanical? Formatting, extraction, triage, bulk scanning. Haiku or Sonnet at low.
- Is it routine and clearly specified? Standard email, summary, rewrite, commodity content. Sonnet at medium, or high if it's a real draft.
- Does it need real reasoning or judgment? Analysis, strategy, synthesis, a genuinely hard problem. Opus at high. Xhigh if it has a lot of moving parts or already failed once.
- Is it voice-critical writing, a decision you'll act on without checking, or truly hard? Fable at high. Save xhigh and max for the rare monsters.
And one check on the way out: whatever you picked, would the next model down give you basically the same result? If yes, go down.
Then I made Claude do it for me
Knowing the system didn't change my behavior. I'd still open a session and run whatever was selected. So I put the system into a Claude Skill that runs on every substantive task.
Before Claude starts working, it names the routing in one line, model and effort both: "Routing: Fable at high. Right for this." If my setup is wrong, it tells me what to switch to and why before doing the work: "Model recommendation: Sonnet at medium is enough here. Fable won't materially improve this."
The part I care about most is that it recommends downgrades as readily as upgrades. I don't need a tool that always says use the expensive one. I need the one that says you're overpaying for this task. And it only pushes a switch when the difference is meaningful. Nobody needs a model debate on every prompt.
Copy the skill
Here it is in full. To install: in Claude, create a new skill named model-router and paste this in as the instructions (or just ask Claude to save it as a skill for you). In Claude Code, save it as .claude/skills/model-router/SKILL.md.
# Model Router Match every task to the model + effort configuration with the lowest **total effective cost** at the quality the task requires, not the lowest token cost. **Total effective cost = model cost + latency + reprompting + correction + the user's editing time + risk of a bad output they act on.** A cheap model that needs three prompts and 25 minutes of editing is more expensive than a strong model getting it right the first time. A top-tier pass that produces what the mid-tier model would have produced is pure waste. Both directions are routing failures. Model lineups and pricing change every few months. If a decision hinges on exact pricing or an unfamiliar model, check the model picker or search before asserting. ## Always-on advisor (runs before every substantive task) Before starting any substantive task (skip trivial chat, quick lookups, and mid-task continuations already routed), evaluate: 1. What model + effort is this session running? If the effort setting isn't visible, say so and state the target instead of guessing. 2. What would the framework below choose for this task? 3. Would switching meaningfully change output quality, reliability, iteration count, editing time, latency, or cost? **Then state a one-line routing readout naming BOTH the model and the effort level, on every substantive task, even when nothing needs to change.** Never name the model without the effort level. Formats: - Setup already right: "Routing: Fable at high. Right for this, proceeding." - Upgrade: "Model recommendation: switch to Fable at high effort. This depends on voice, synthesis, and first-draft quality, and the stronger model will save you real editing time." - Downgrade: "Model recommendation: Sonnet at high is enough here. You're on Fable, but the extra capability won't change this result enough to justify the burn." - Effort only: "Effort recommendation: stay on this model, but high is sufficient. Max adds cost and latency with no expected quality gain." If the user must switch manually, tell them what to switch to and why. **Rules:** the readout is mandatory; the change recommendation is reserved for meaningful gaps. No gratuitous switching; a configuration that's theoretically 5% better gets a confirmation line, not a switch pitch. Flag a mismatch once, respect the user's call, move on. Recommend **downgrades as readily as upgrades**; preventing wasted top-tier and max-effort spend is half this skill's job. The downgrade test: is the expected final result meaningfully worse on the cheaper configuration? If not, recommend it. ## The two dials **Dial 1: effort.** The primary cost lever. Max effort can consume several times the tokens of low effort on the same model. Effort and model are separate decisions: never auto-pair the strongest model with the highest effort. **Dial 2: model tier.** Tier sets the ceiling; effort decides how much of it gets used. Move up a tier only when the task needs capability the current model lacks at any effort: deeper judgment, better instruction-following under many constraints, longer autonomous horizon, higher cost of being confidently wrong. Rule of thumb: **tune effort first, switch models second.** Exceptions: judgment-heavy work and authored writing, where effort on a smaller model doesn't substitute for tier. Go straight to the right tier there. ## The routing questions (in order) 1. **What's the cost of a wrong answer?** Acted on unverified (a decision, a negotiation, a public post): route up. Reviewed and edited anyway: downward pressure. 2. **Is this judgment or execution?** Judgment (diagnose, decide, prioritize, synthesize under conflicting constraints) wants tier. Execution (format, extract, apply a template, follow explicit steps) wants a smaller model at lower effort. 3. **Is it writing?** Classify with the writing check below. 4. **What's the iteration economics?** Estimate the user's total work under each option. If the cheaper model's draft will need substantial reprompting or editing the stronger model would avoid, the stronger model is the cheaper choice. If outputs would be effectively equivalent, the cheaper model wins outright. 5. **How long is the autonomous run?** Long agentic runs compound small errors; use a stronger model as the driver even when each individual step looks easy. ## Effort (decided after the model) - **Low/medium:** obvious tasks, simple rewrites, routine repetitive work, little ambiguity. - **High (the normal default for substantive work):** standard analysis, normal professional writing, research synthesis, moderately complex reasoning. - **Xhigh/max (reserved):** extremely complicated reasoning, large interacting constraint sets, very long agentic workflows, difficult debugging, problems that already failed at high. **Never select max because something is important. Importance and computational difficulty are not the same thing.** Never leave max on as a session default. ## The writing check (run for every writing request) First ask what the writing has to accomplish: search engine or human reader, action or information transfer, the user's voice or just clarity. Then classify: 1. **Commodity content: Sonnet.** Everyday writing that just has to be clear and correct: rewrites, summaries, repurposing, variants, SEO pages, routine copy. A stronger model adds little. 2. **Normal professional writing: Sonnet or Opus** depending on how much synthesis is involved. 3. **Authored/voice-critical/high-value writing: strongly consider Fable.** Original articles, essays, argument-driven pieces, anything that must sound distinctly like the user, anything they historically spend real time editing. Why: high-quality writing is a complex instruction-following and judgment problem (inferring principles from examples, resolving conflicting style rules, structuring an argument, deciding emphasis, avoiding AI patterns, all at once). That is tier-limited, not effort-limited. When the user invokes a substantial personal writing skill, give extra weight to whether the stronger model will interpret it better and cut their editing time. ## Skills and model capability A skill is instructions, not capability. A well-skilled mid-tier model often beats an unskilled bigger one, but a stronger model extracts more value from the exact same skill, especially complex writing and strategy skills. If a sophisticated skill executes poorly, consider whether model capability is the limit before making the skill longer. The right fix is often the same skill on a stronger model. ## Learned performance (observed evidence outranks generic guidance) The user's repeated experience with a task class is routing evidence. When they find one model consistently better or worse for a category, record it here and route on it. Entry format: - [Task class]: [winning model + effort], because [observed difference]. Added [date]. ## Starter task map (customize this section first) - **Bulk mechanical work** (scanning, triage, formatting, extraction): Haiku, low. - **Routine professional tasks and commodity content:** Sonnet, medium/high. - **Standard professional writing:** Sonnet, high. - **Difficult analysis, strategy, multi-document synthesis:** Opus, high (xhigh when truly hard). - **Voice-critical authored writing:** Fable, high. - **Exceptionally difficult reasoning or long agentic workflows:** Fable or Opus, xhigh/max as justified. - **Decisions the user will act on without checking:** Fable, high. Unmapped tasks: run the routing questions. True coin flips: name both options and the tradeoff in one line and let the user pick. ## Anti-patterns - Skipping the routing readout, or naming a model without its effort level. - Recommending a bigger model as a hedge because you're unsure. - Classifying all writing as "just generation" and defaulting to the cheap model. - Routing judgment or authored writing down "to save tokens" when the user acts on it unverified or pays in editing time. Redone work is the most expensive token spend there is. - Treating token cost as the whole cost. The user's time counts. - Leaving max effort on as a session default; pairing the strongest model with the highest effort by reflex. - Assuming skills fully substitute for model capability. - Switch-nagging over marginal differences; a fine configuration gets a confirmation line, not a pitch.
Make it yours
The skill is written so anyone can run it as-is, but the starter task map is a starting point, not an answer. A developer will route more toward the strong coding models. A high-volume content operation will lean on Sonnet much harder. If most of your Claude use is admin work, you may almost never need the top tier.
The learned performance section is how it gets better. When you notice one model consistently winning or losing on a category of your work, have Claude add an entry, and let your own results override everything in this article. That includes my Fable-for-writing lane. Try it on one real piece with your own instructions and see what the editing time tells you.
LMK what your routing ends up looking like. I'm curious whether other people's maps land anywhere near mine.
The skills stay free. New ones go out in my weekly newsletter, along with real marketing job leads and my thinking on careers and AI.