Skip to content

Getting reliable JSON out of an LLM

By SunnyKumar Jonwal 10 min read

The first version of most LLM features ends with a line like "respond with JSON only." It works in the demo. Then in production, one response in a few hundred arrives wrapped in a friendly sentence, another has a trailing comma, a third invents a field name, and your parser throws.

Reliable structured output is a solved problem if you use the right technique for your situation. This post covers the options, how they differ, and the validation habits that turn a mostly-works feature into a dependable one.

Why prompting alone isn't enough

A model generates text one token at a time, and nothing in a plain text prompt forces valid JSON. The model usually complies because it has seen a great deal of JSON, and "usually" is the problem. Typical failures include prose before or after the object, Markdown code fences around it, single quotes or unquoted keys, truncated output when a length limit hits, missing required fields, wrong types ("3" instead of 3), and values outside the set you intended.

You can reduce these with a clear prompt and an example, and you should. But if a program consumes the result, you need a technique that gives guarantees or at least detects and recovers from failures.

Option 1: ask nicely and parse defensively

The simplest approach: describe the format in the prompt, give an example, and parse the response with error handling. Strip code fences, find the first { and last }, and try json.loads. If it fails, retry.

This is fine for prototypes and low-stakes features. It costs an extra call whenever parsing fails, and it offers no protection against valid JSON with wrong content. It's a floor, not a plan.

Option 2: JSON mode

Some providers offer a mode that guarantees the output parses as JSON. That eliminates syntax failures, which is a real gain. It doesn't enforce your particular schema, so fields can still be missing or mistyped. Think of it as "valid JSON, unknown shape."

Option 3: force a tool call

With Anthropic's API, and similar features elsewhere, a reliable pattern is to define a tool whose input schema is the structure you want, and require the model to call it. You never execute anything. The tool call's arguments are your structured data, generated to match the schema.

Here's what that looks like in Python, with Pydantic defining the schema and validating the result:

from typing import Literal

import anthropic
from pydantic import BaseModel, Field, ValidationError

client = anthropic.Anthropic()
MODEL = "claude-sonnet-5"  # use a current model name from the docs


class Ticket(BaseModel):
    category: Literal["billing", "bug", "feature_request", "other"]
    priority: Literal["low", "medium", "high"]
    summary: str = Field(description="One sentence, under 200 characters.", max_length=200)
    needs_human: bool = Field(description="True if the message is angry, legal, or unclear.")


record_ticket = {
    "name": "record_ticket",
    "description": "Record the classification of a support message.",
    "input_schema": Ticket.model_json_schema(),
}


def classify(message: str) -> Ticket:
    response = client.messages.create(
        model=MODEL,
        max_tokens=512,
        tools=[record_ticket],
        tool_choice={"type": "tool", "name": "record_ticket"},
        messages=[{"role": "user", "content": f"Classify this support message:\n\n{message}"}],
    )
    data = next(block.input for block in response.content if block.type == "tool_use")
    return Ticket.model_validate(data)

tool_choice forces the call, so the model can't respond with prose. The schema, generated from the Pydantic class, tells it exactly which fields exist and what values are allowed. And model_validate checks the result on your side, because you should always verify what comes back.

This pattern needs no special feature beyond ordinary tool use, which is why it's so widely used. The background on how tools work is in tool use explained.

Option 4: native schema-constrained output

A growing number of providers offer structured output as a first-class feature: you supply a JSON Schema, and the system constrains generation so the output always conforms, using techniques that restrict which tokens are allowed at each step. Where it's available, it gives the strongest guarantee of syntactic and structural validity. Anthropic, OpenAI, and Google each offer versions of this on some models, with differences in supported schema features and limits, so check the current documentation for your provider and model.

Two things to remember even with constrained decoding. First, it guarantees the shape, not the truth. A perfectly formed object can still contain a wrong classification or an invented value. Second, schema support has limits: some keywords, deep recursion, or very large schemas may not be allowed. Design your schemas to fit.

Validate on your side, always

Whatever technique you pick, validate the result in your own code before using it. Pydantic in Python, Zod in TypeScript, Laravel's validator in PHP:

use Illuminate\Support\Facades\Validator;

$validator = Validator::make($data, [
    'category' => ['required', 'in:billing,bug,feature_request,other'],
    'priority' => ['required', 'in:low,medium,high'],
    'summary' => ['required', 'string', 'max:200'],
    'needs_human' => ['required', 'boolean'],
]);

if ($validator->fails()) {
    // retry with the errors, or route to a human
}

Validation catches what the model got wrong, and it protects you from unexpected changes if a provider updates a model. It also lets you check things a schema can't express: that an ID exists in your database, that a date is in the future, that two fields are consistent with each other.

Retry with feedback

When validation fails, don't just call again blindly. Tell the model what went wrong. Send the error message back and ask for a corrected response:

def classify_with_retry(message: str, attempts: int = 3) -> Ticket:
    messages = [{"role": "user", "content": f"Classify this support message:\n\n{message}"}]
    for _ in range(attempts):
        response = client.messages.create(
            model=MODEL, max_tokens=512, tools=[record_ticket],
            tool_choice={"type": "tool", "name": "record_ticket"}, messages=messages,
        )
        block = next(b for b in response.content if b.type == "tool_use")
        try:
            return Ticket.model_validate(block.input)
        except ValidationError as err:
            messages.append({"role": "assistant", "content": response.content})
            messages.append({"role": "user", "content": [{
                "type": "tool_result", "tool_use_id": block.id, "is_error": True,
                "content": f"Validation failed: {err}. Call record_ticket again with valid values.",
            }]})
    raise RuntimeError("Could not get a valid ticket after several attempts.")

A cap on attempts matters. After two or three tries, fall back to a human, a default, or a failure you can log and inspect. Track how often retries happen. A rising retry rate is an early warning that a prompt, schema, or model change went wrong.

Design schemas that models fill well

The schema is part of the prompt, and good design improves accuracy.

Describe every field. A description on each property tells the model what belongs there, with format hints and examples where useful. "priority" alone is ambiguous. "priority: low for questions, medium for degraded service, high for outages or data loss" is a spec.

Use enums for fixed choices. They remove a whole class of inventiveness and make downstream code simple.

Keep it flat and small. Deeply nested structures and dozens of fields raise the odds of mistakes and make each call slower and more expensive. If you need a lot of data, consider several smaller calls or a top-level structure with modest nesting.

Make optional things explicitly nullable. If a field may not exist in the source, allow null and say so, or the model will invent a value to satisfy "required." This matters a great deal for extraction tasks, and it links directly to the fabrication problem in why LLMs hallucinate.

Put a reasoning field first when it helps. For classification tasks, having the model write a short reasoning string before the label sometimes improves accuracy, since it gets to work through the problem before committing. Discard the field afterward if you don't need it. The order of fields in generated output usually follows the order in the schema.

Prefer simple types. Strings, numbers, booleans, and arrays of those are easy. Formats such as dates are best specified explicitly ("ISO 8601 date, YYYY-MM-DD"), and validated afterward.

Handle the awkward cases

A few situations deserve explicit thought.

Truncation. If the output hits your token limit mid-object, you'll get invalid JSON or a cut-off tool call. Check the stop reason. If it's a length limit, increase the cap or shrink the task. Treat truncated output as a failure, and don't try to parse it.

Refusals and off-topic responses. Input that the model declines to process, or that has nothing to extract, still needs a defined result. Build it into the schema: an is_relevant flag, or an error field, or nullable fields, so "nothing found" is a valid, expected answer and not an exception.

Large lists. Extracting a hundred items in one response risks truncation and drift. Chunk the input, extract from each chunk, and merge and deduplicate afterward.

Streaming. If you stream a structured response to show progress, you'll receive partial JSON. Use a parser that tolerates incomplete input for display, and validate only the final result.

Versioning. Schemas change. When you add or rename a field, decide how old stored data and old prompts are handled, and record the schema version with each result.

Structured output for agents and pipelines

The same ideas apply inside larger systems. When one step's output feeds the next, whether in a pipeline or between agents in a multi-agent design, a validated schema is the contract between the steps. Typed handoffs prevent a great many silent failures, since a malformed message fails loudly at the boundary and doesn't corrupt everything downstream.

And for extraction from untrusted text such as emails or web pages, remember that the content can carry instructions. Constraining output to a strict schema limits what a hijacked response can do, which is a small but real defense. It doesn't replace the broader measures in prompt injection and the lethal trifecta.

Common mistakes

Most structured-output bugs come from a short list of habits.

Marking every field required. If the source text might not contain a phone number, a required phone field gives the model no honest option. It will produce something that looks like a phone number. Make the field nullable, and tell the model to use null when the information isn't there.

Trusting the shape and skipping the content check. A response that parses cleanly and passes the schema feels finished. But "priority": "low" on a message about a server outage is valid and wrong. Sample real outputs by hand every week for the first month, and keep the ones that surprised you as permanent test cases.

Overloading one call. Asking for a classification, a summary, a sentiment score, an entity list, and a suggested reply in one object makes each part slightly worse. If the parts are independent, split them, and you can run them in parallel.

Ignoring token limits. A schema with a long free-text field and a small max_tokens will truncate at the worst moment. Size the limit for the longest realistic output, then add headroom.

Changing the schema casually. Renaming a field breaks stored data, dashboards, and any consumer downstream. Add new fields as optional first, migrate, and only then tighten.

Forgetting the model changes. A provider update can shift behavior on borderline inputs. Your validators and your regression set are what tell you it happened, which is one more reason to keep both around.

Testing

Include structured-output cases in your evaluation set: messy inputs, missing information, contradictory details, adversarial text, very long inputs. Measure the schema-valid rate, the correctness of the fields against ground truth, and the retry rate. Check both, because valid but wrong is the failure mode that structure hides.

A quick decision guide

For a prototype, prompt clearly and parse defensively. For most production features, force a tool call or use native structured output, then validate in your own code. If you need the strongest structural guarantee and your provider supports it, use schema-constrained decoding, and still validate semantics. In every case, cap retries, log failures, and plan for the "nothing to extract" case.

Structured output is one of those areas where a small amount of engineering removes a large amount of pain. Ten lines of validation and one retry loop are usually the difference between a feature you have to babysit and one you can leave running.