LLM tokens and costs, explained with real arithmetic
The first surprise most people get from an LLM API is the bill. A prototype that cost pennies in testing costs real money once users show up, and it's hard to say why, because the pricing page talks about "tokens" and you were thinking in requests. Once you can do the arithmetic in your head, budgeting stops being a guess.
This post explains what tokens are, how pricing is structured, why agents cost more than chatbots, and which levers actually move the number. Prices change often, so I'll use made-up round numbers for the math and leave the real figures to the provider's pricing page.
What a token is
A model doesn't read letters or words. It reads tokens, which are chunks of text produced by a tokenizer, a fixed vocabulary of pieces learned before training. Common words are often one token. Rarer words split into several. Punctuation, spaces, and numbers each take their share.
For ordinary English, a rough rule is that a token is about four characters, or three quarters of a word. So a thousand tokens is around seven hundred fifty words. That's only a rule of thumb. Code tends to use more tokens per line than prose, because of symbols and indentation. Languages other than English often use more tokens for the same meaning, which means the same content can cost noticeably more. Long numbers and unusual strings, such as UUIDs or base64, are expensive because they break into many small pieces.
Don't estimate when you can count. Most providers expose a token-counting endpoint or return usage numbers in every response. Anthropic's API includes input and output token counts in each response, and there are tokenizer tools for offline checks. Different model families use different tokenizers, so the same text can have different counts on different models.
How pricing is structured
Almost every API charges separately for input tokens and output tokens, quoted per million. Output costs more, often several times as much per token. There's a reason for that: generating text is sequential, each new token needs a pass through the model, while the input is processed in bulk and in parallel.
A worked example with invented prices. Say a model charges 3 dollars per million input tokens and 15 dollars per million output tokens. A request with 2,000 tokens of input and 500 tokens of output costs:
- Input: 2,000 ÷ 1,000,000 × 3 = $0.006
- Output: 500 ÷ 1,000,000 × 15 = $0.0075
- Total: $0.0135, about a cent and a third
A cent sounds trivial. Multiply by traffic. At 100,000 requests a day, that's $1,350 a day, or roughly $40,000 a month. A small difference in the prompt suddenly matters. Cutting the input in half saves $300 a day, and it took one afternoon of editing.
Notice which side dominated. Here output was more than half the cost despite being a quarter of the tokens. If your feature produces long answers, output is your biggest lever, and if it reads huge documents, input is.
Other line items exist. Many providers discount cached input, offer a cheaper batch mode for non-urgent work, and charge extra for special features such as web search or code execution. Reasoning or "thinking" tokens, produced by models that work through a problem before answering, are billed as output tokens even if you never see them. Check your provider's current pricing page for the specifics.
Why agents cost more than you expect
A chatbot call is one request. An agent is a loop, and the cost of the loop has a nasty property: it's quadratic-ish in the number of steps.
Recall how the agent loop works. At each step, the whole conversation so far is sent again: the system prompt, the tool definitions, every earlier message, every tool result. Step one sends 5,000 tokens. Step two sends those 5,000 plus what happened in step one, say 7,000. Step ten might send 30,000. The model reads the earlier text again each time, and you pay for it again each time.
Suppose an agent runs twenty steps and its context grows by 2,000 tokens per step, starting from 4,000. The input for each step is 4,000, then 6,000, then 8,000, up to 42,000. Add them up and you send about 460,000 input tokens for a single task, even though the final context is only 42,000. At 3 dollars per million, that's roughly $1.40 for the input alone, on a task that a human might call small. Run it a thousand times a day and the arithmetic becomes a budget line.
This is the reason long-running agents need discipline about context size. The habits in context engineering for AI agents aren't only about quality. Every token you keep in the window is a token you pay for on every subsequent step.
Extra cost drivers to watch in agents:
Big tool results. A tool that returns an entire web page or a whole file is stuffing thousands of tokens into every later step. Have tools return summaries, limit rows, and support pagination. The design advice in designing tools for LLM agents applies directly.
Many tool definitions. Every tool's name, description, and schema is sent with every request. Fifty tools can add thousands of tokens of fixed overhead to each step.
Retries and loops. An agent stuck repeating a failing call burns money quietly. Always set a step limit and a token budget.
Sub-agents. Multi-agent systems multiply everything, since each agent has its own loop and context. That's sometimes worth it, as discussed in multi-agent systems, and you should know you're paying for it.
The levers, in order of usefulness
Here's the order I'd try things, from cheapest in effort to most involved.
Measure first
Log input tokens, output tokens, model, and feature name for every call. Group by feature. You'll almost always find that one or two endpoints account for most of the spend, and they're rarely the ones you guessed. Without this, optimization is superstition.
Trim the prompt
Read your prompts as a stranger would. Look for repeated instructions, long examples that could be shorter, boilerplate legal text pasted in, and history that's no longer relevant. Cut whatever doesn't change the model's behavior. Test after each cut, since a few lines that look redundant occasionally hold the whole thing together.
Cache the stable prefix
If most of your prompt is identical from one request to the next, such as a long system prompt, tool definitions, or a reference document, prompt caching lets the provider reuse the processing for that prefix at a large discount. It's usually the biggest single saving for agents and for question-answering over a fixed document. The details, including how to structure prompts so the cache actually hits, are in prompt caching to cut cost and latency.
Limit output
Ask for what you need. If a downstream program wants a label, request the label and not an explanation. Set max_tokens to a sensible ceiling. Tell the model the desired length: "two sentences" is respected far more often than people expect. Remember output tokens are the expensive ones.
Use the right size of model
Not every task needs the biggest model. Classification, extraction, routing, and simple rewriting often work fine on a smaller, cheaper tier at a fraction of the price. Reserve the expensive model for the steps that need it. Choosing the right Claude model covers how to decide, and how to test the decision.
Batch what isn't urgent
If the work doesn't need an immediate answer, such as nightly classification, bulk summarization, or backfills, a batch interface can cost meaningfully less than real-time calls. The tradeoff is latency, since results may take hours.
Retrieve, don't stuff
Sending an entire knowledge base with every request is expensive and often worse for quality. Retrieving only the relevant pieces cuts input dramatically. The tradeoffs are laid out in long context versus RAG.
Cap and alert
Set per-user and per-day budgets, step limits for agents, and alerts when spend jumps. A bug that loops all night is cheaper to find at 2x normal spend than at 200x.
Common estimation mistakes
A few errors show up repeatedly when people forecast costs.
Forgetting that output and input have different prices, and averaging them. Estimate each side separately.
Ignoring the system prompt and tool definitions. If they total 6,000 tokens and the user's message is 100, the fixed part is the cost.
Testing with tiny inputs. Demo questions are short. Real users paste documents, long threads, and logs.
Counting requests and not steps. One user action may trigger fifteen model calls in an agent.
Assuming prices stay flat. Prices per token have generally fallen over time, and models come and go. Build so that switching models is a config change, and revisit your numbers every quarter or so.
A budget for one feature, start to finish
Suppose you're adding an "explain this error" button to a developer tool. Users paste a stack trace, and the assistant explains it. Work the numbers before writing code.
The system prompt is about 600 tokens. A typical stack trace with surrounding code is 1,500 tokens. The explanation runs 350 tokens. That's roughly 2,100 input and 350 output. With the invented rates from earlier, input costs about $0.0063 and output about $0.0053, so around a cent and a bit per click.
Now ask how often it'll be used. If ten thousand users click it twice a week, that's about 80,000 calls a month and roughly $920. Is that fine? Compare it to what the feature is worth. If it cuts support tickets or keeps users on the product, it probably is. If it's a novelty, you'd want a smaller model, a shorter prompt, or a daily cap per user.
The exercise takes ten minutes and shapes real decisions: whether to cache the system prompt, whether the explanation needs to be 350 tokens or 150, and whether to offer the feature to free users at all. Teams that skip it tend to discover the answer on an invoice.
A small cost tracker
Here's a minimal Python wrapper that logs usage from Anthropic's SDK and estimates dollars. Fill in your own rates.
import time
import anthropic
client = anthropic.Anthropic()
# Dollars per million tokens. Replace with current prices from the pricing page.
RATES = {"claude-sonnet-5": {"in": 3.0, "out": 15.0}}
def tracked_call(feature: str, **kwargs):
started = time.time()
response = client.messages.create(**kwargs)
usage = response.usage
rate = RATES[kwargs["model"]]
cost = (usage.input_tokens * rate["in"] + usage.output_tokens * rate["out"]) / 1_000_000
print(f"{feature}: in={usage.input_tokens} out={usage.output_tokens} "
f"cost=${cost:.4f} time={time.time() - started:.1f}s")
return response
In production you'd write those numbers to a database or metrics system instead of printing, and add fields for cached tokens if you use caching. But even this much, run for a week, tells you where the money goes.
What good looks like
A healthy setup isn't the cheapest one, it's the one where you know the cost per unit of value. Cost per resolved ticket, per summarized document, per successful agent task. If a task costs 40 cents and saves a person ten minutes, you're fine. If it costs 40 cents and the answer is wrong a third of the time, the number that matters is cost per correct result, which is why evaluation belongs in the cost conversation and not only the quality one.
So count tokens, log every call, put the expensive work where it earns its keep, and cache what repeats. The pricing page is simple enough. The bill surprises people because of the multiplication, and you can plan for that once you've done the multiplication yourself.