Skip to content

Choosing the right Claude model for the job

By SunnyKumar Jonwal 9 min read

Every model family ships in tiers. There's a top model that's the most capable and the most expensive, a middle model that's the everyday workhorse, and a small model that's fast and cheap. Anthropic has organized Claude this way for a long time, and other providers do the same. The names and exact lineup change with each release, so check the current models page before you copy anything from a blog post, this one included.

The question that outlives any particular lineup is how to pick. Defaulting to the biggest model feels safe, and it's usually the wrong call for most of what you'll build. This post gives a way to decide that rests on measurement.

What actually differs between tiers

Three things move together as you go up the ladder.

Capability. Larger models generally handle harder reasoning, longer chains of steps, subtle instructions, and unfamiliar tasks better. On easy tasks the gap can be tiny or invisible.

Cost. Price per token climbs steeply. The top tier can cost several times what the middle one does and an order of magnitude more than the smallest. Since costs multiply in agents, the difference compounds.

Speed. Smaller models respond faster and stream faster. For interactive features, latency shapes how the product feels, and a response in one second and one in eight seconds are different experiences.

Other differences can matter: context window size, support for certain features like extended thinking, image input, or tool use behaviors, and output limits. Look at the documentation for the specific models you're weighing, since these details shift between versions.

The default that saves money

The stance I'd suggest is to start in the middle, then move in whichever direction the data points. A mid-tier model is usually strong enough to build the whole feature and see where it struggles. If it works well, try the smaller tier and see whether quality holds. If it fails on specific cases, try the larger one and see whether those cases get fixed.

Starting at the top hides problems. You never learn how much of the task was easy, and you ship an expensive default that nobody revisits.

Match the tier to the task

Some rough patterns, which you should verify on your own data.

Small models tend to suit high-volume, well-defined work: classification, routing, tagging, extraction from clean text, short rewrites, moderation, and simple question answering over provided context. These tasks have clear right answers and little ambiguity, so extra reasoning power buys little.

Mid-tier models tend to suit most product features: drafting, summarization of long documents, coding assistance, tool-using agents with moderate complexity, data analysis, and anything with several constraints to juggle.

Top-tier models tend to earn their price on the hard tail: long multi-step agent tasks that need to stay coherent over many actions, ambiguous or underspecified problems, difficult code changes across a large codebase, careful analysis where mistakes are costly, and planning steps that steer a cheaper worker.

Those are tendencies, not rules. A small model with a good prompt and a few examples can beat a large model with a lazy one, and a well-defined hard task may still need the big model.

Test it: a simple method

You can decide with a few dozen examples and an afternoon. This is a lighter version of the process in how to evaluate AI agents.

  1. Collect 30 to 100 real inputs for the feature, including some awkward ones. Real logs beat invented examples.
  2. Decide what "good" means for each, either with a reference answer, a rule-based check, or a rubric.
  3. Run every model you're considering on the same inputs, with the same prompt.
  4. Grade the outputs. Use code checks where you can, an LLM judge with a specific rubric where you can't, and read a sample by hand no matter what.
  5. Record quality, average tokens, latency, and cost per call for each model.
  6. Compare on cost per acceptable result, not cost per call.

A small script does the running:

import time
import anthropic

client = anthropic.Anthropic()
CANDIDATES = ["small-model-id", "mid-model-id", "large-model-id"]  # real IDs from the docs


def run_all(cases, system_prompt, grade):
    for model in CANDIDATES:
        passed, total_time = 0, 0.0
        for case in cases:
            started = time.time()
            reply = client.messages.create(
                model=model,
                max_tokens=600,
                system=system_prompt,
                messages=[{"role": "user", "content": case["input"]}],
            )
            total_time += time.time() - started
            passed += grade(case, reply.content[0].text)
        print(f"{model}: {passed}/{len(cases)} passed, "
              f"{total_time / len(cases):.1f}s average")

Run each case more than once if the task has any variability, because a single sample can mislead you either way.

The outcome is often striking. It's common to find the smallest model passes 95 percent of a simple task at a tenth of the price, and that the remaining 5 percent are identifiable, which leads to the next idea.

Route instead of choosing once

You don't have to pick one model for everything. Several patterns let you use each tier where it fits.

Static routing. Different features use different models. Tagging uses the small one, the chat assistant uses the mid one, and the weekly report generator uses the large one. It's simple and predictable, and it's where most teams should start.

Cascade with fallback. Try the cheap model first. If its answer fails a check, such as failing schema validation, low self-reported confidence, or a rule-based test, escalate to a stronger model. You pay the big-model price only on hard cases. The checks matter: a cascade is only as good as your ability to tell when the cheap answer is bad. Combining this with validated structured output makes the check mechanical.

Planner and workers. Use a strong model to plan or decide, and cheaper models to execute the many small steps. In agent systems, this can cut costs a lot when the workers' tasks are narrow. It also fits the orchestration patterns in multi-agent systems.

Classifier router. A small model reads the request and decides which model should handle it. This adds a call, so it pays off only when the cost gap is large and the classification is reliable.

Keep routing logic boring. Elaborate schemes tend to break in ways that are hard to debug, and every extra path is one more thing to evaluate.

Latency and user experience

Cost isn't the only reason to go small. If a feature blocks the user, like autocomplete, inline suggestions, or a search-as-you-type assistant, speed may matter more than the last few points of quality. Streaming helps perceived speed on any model, since users see the first words quickly, and a smaller model helps everywhere.

For background work, the reverse holds. No one is waiting, so use the model that gives the best result per dollar and let it take its time.

Extended thinking or reasoning modes, where a model spends extra tokens working through a problem before answering, give more accuracy on hard tasks at the cost of speed and money. Turn them on where they measurably help, and leave them off for easy calls. They're billed as output tokens, so they add up.

Don't forget the prompt

Model choice interacts with prompting. A smaller model usually needs clearer, more explicit instructions, a few well-chosen examples, and a narrower task. A larger model tolerates ambiguity better. Before you conclude a small model can't do something, spend an hour improving the prompt using the guidance in prompt engineering that still works. The result is often that it can.

Prompt caching also changes the math. If a long shared prefix is cached, a larger model's input cost drops considerably, which narrows the gap. Recalculate after enabling it. The mechanics are in prompt caching.

Model versions and change

Models get retired, updated, and replaced by newer ones. A few habits make that painless.

Pin the specific model identifier in configuration, and don't scatter it through code. Use dated snapshot names when the provider offers them if you need stable behavior, and understand that aliases may point to newer models over time. Keep your test set and rerun it whenever you change a model, because a new version that's better on average can be worse on your particular task. Read release notes and deprecation schedules, and calendar the dates.

Also plan for the opposite problem: newer, cheaper models often match the old top tier. The choice you made a year ago may not hold. Review it periodically with the same test set.

An example decision, worked through

Imagine a helpdesk tool with three AI features: tagging incoming tickets, drafting replies for agents, and a weekly analysis of trends across all tickets. It's tempting to run all three on the same model. Here's how the model-choice loop would treat each.

Tagging is high volume and has a closed set of labels. You test the small and mid tiers on 200 labeled tickets. The small one agrees with your labels 94 percent of the time and the mid one 96 percent. Two points isn't worth paying several times as much across every ticket, so tagging goes on the small model, with a schema so bad outputs are caught.

Drafting replies is where quality is visible to people. Agents read every draft and edit it. You test the mid and large tiers, and have a few agents rate blind pairs. If the raters can't reliably tell which is better, the mid tier wins on cost and speed. If they consistently prefer the large one on complex tickets, you can send only tickets flagged as complex to the large model.

The weekly trend analysis runs once, with no one waiting, over a big pile of text. Here the cost of being wrong is a bad decision by a manager, and the volume is one call a week. The large model is a sensible choice, because the total spend is tiny and the quality matters.

The result is three features on three different tiers, each justified by numbers you collected. Contrast that with one model everywhere, where either the tagging is absurdly expensive or the analysis is weaker than it should be.

When the benchmark isn't your task

Public leaderboards are a fine way to shortlist models and a poor way to choose one. They measure general skills on standard problems, and your task has its own vocabulary, formats, and failure costs. A model that tops a coding benchmark may still stumble on your framework's conventions, and a modest one may do better on your support tone. Use published numbers to decide what to test, and use your own cases to decide what to ship.

Mistakes to avoid

Picking the top model everywhere "to be safe" and never revisiting. Picking the smallest model everywhere to save money, then losing the savings to retries and complaints. Judging by a handful of demos. Comparing models with different prompts, which tests the prompts and not the models. Ignoring latency until users complain. Hardcoding a model name in ten places.

A practical starting recipe

If you want something to do this week, here it is. Log usage per feature so you know where spend is. For the most expensive feature, build a set of fifty real cases and a way to grade them. Run it on one tier down. If quality holds within a margin you accept, switch and pocket the savings. If it doesn't, look at which cases failed, and consider a cascade that sends only those to the bigger model. Repeat for the next feature.

That loop takes a few hours per feature and often cuts a bill by half or more, with no change your users would notice. That's a better return than most optimizations, and it depends on measuring instead of assuming.