Skip to content

Add Claude to a Laravel app with the HTTP client

By SunnyKumar Jonwal 11 min read

Laravel developers tend to reach for a package first. For a model API, though, the HTTP client that ships with the framework is enough, and the resulting code is short, readable, and yours. There's no wrapper to upgrade and no hidden behavior. This post builds a small but production-minded integration with Claude's Messages API.

It assumes Laravel 11 or 12 and PHP 8.2 or newer. The API details, including model names and headers, come from Anthropic's documentation and do change, so check the current reference before you ship. Model names below are placeholders.

Configuration first

Put the key and the model in environment variables, then expose them through a config file so nothing calls env() outside config.

ANTHROPIC_API_KEY=your-key-here
ANTHROPIC_MODEL=your-model-id

Never commit the key. Add it to your .env on the server, and keep .env out of version control. If a key ever lands in a repository or a log, rotate it right away.

// config/services.php
'anthropic' => [
    'key' => env('ANTHROPIC_API_KEY'),
    'model' => env('ANTHROPIC_MODEL'),
    'version' => '2023-06-01',
    'timeout' => 60,
],

The version value is the API version header Anthropic requires. Confirm the current value in their docs.

A small service class

Wrap the calls in one class. Controllers, jobs, and commands then talk to your class and never to the API directly, which means you can change models, add logging, or swap providers in one place.

<?php

namespace App\Services;

use Illuminate\Http\Client\PendingRequest;
use Illuminate\Support\Facades\Http;
use RuntimeException;

class Claude
{
    private function client(): PendingRequest
    {
        return Http::baseUrl('https://api.anthropic.com/v1')
            ->withHeaders([
                'x-api-key' => config('services.anthropic.key'),
                'anthropic-version' => config('services.anthropic.version'),
            ])
            ->acceptJson()
            ->asJson()
            ->timeout(config('services.anthropic.timeout'))
            ->retry(2, 500, fn ($e) => $this->isRetryable($e), throw: false);
    }

    private function isRetryable($exception): bool
    {
        $status = method_exists($exception, 'response') && $exception->response
            ? $exception->response->status()
            : null;

        return in_array($status, [429, 500, 502, 503, 529], true);
    }

    public function message(string $prompt, ?string $system = null, int $maxTokens = 1024): string
    {
        $payload = [
            'model' => config('services.anthropic.model'),
            'max_tokens' => $maxTokens,
            'messages' => [['role' => 'user', 'content' => $prompt]],
        ];

        if ($system) {
            $payload['system'] = $system;
        }

        $response = $this->client()->post('/messages', $payload);

        if ($response->failed()) {
            throw new RuntimeException(
                'Claude request failed: '.$response->status().' '.$response->json('error.message', 'unknown error')
            );
        }

        return collect($response->json('content', []))
            ->where('type', 'text')
            ->pluck('text')
            ->implode('');
    }
}

A few things are worth pointing out in that code. The client sets a timeout, because a model call can take many seconds and you don't want a hung request tying up a worker forever. It retries only on statuses that mean "try again later," such as rate limiting and temporary overload, and it does not retry on a 400, which means your request is wrong and a retry won't fix it. And the response content is an array of blocks, so the method collects the text blocks and joins them, since a response can contain more than one.

Register it as a singleton if you like, or let the container build it on demand. Either works:

// app/Http/Controllers/SummaryController.php
public function __invoke(Request $request, Claude $claude)
{
    $data = $request->validate(['text' => ['required', 'string', 'max:20000']]);

    $summary = $claude->message(
        prompt: "Summarize this in three sentences:\n\n".$data['text'],
        system: 'You write plain, accurate summaries. Do not add facts that are not in the text.',
        maxTokens: 300,
    );

    return response()->json(['summary' => $summary]);
}

Notice the input validation with a length cap. Without it, one user could send a huge document and run up your bill, a cost problem explained in LLM tokens and costs.

Don't make users wait: queue it

Model calls are slow by web standards. For anything that isn't strictly interactive, push the work to a queued job and show the result when it's ready.

<?php

namespace App\Jobs;

use App\Models\Article;
use App\Services\Claude;
use Illuminate\Bus\Queueable;
use Illuminate\Contracts\Queue\ShouldQueue;
use Illuminate\Foundation\Bus\Dispatchable;
use Illuminate\Queue\InteractsWithQueue;
use Illuminate\Queue\SerializesModels;

class SummarizeArticle implements ShouldQueue
{
    use Dispatchable, InteractsWithQueue, Queueable, SerializesModels;

    public int $tries = 3;
    public array $backoff = [10, 60, 300];
    public int $timeout = 120;

    public function __construct(public Article $article) {}

    public function handle(Claude $claude): void
    {
        $summary = $claude->message(
            prompt: "Summarize in two sentences:\n\n".$this->article->body,
            maxTokens: 200,
        );

        $this->article->update(['summary' => $summary]);
    }
}

The job's own tries and backoff handle transient failures on top of the HTTP-level retry. Set the job timeout above the HTTP timeout so the queue doesn't kill it mid-request. And make sure the job is safe to run twice, since retries can happen: here it just overwrites a column, so a repeat is harmless.

Tool use from PHP

Everything so far is single-shot. To build something agent-like, you handle tool calls. The concepts are in tool use explained, and the loop looks like this in PHP:

public function run(string $prompt, array $tools, callable $execute, int $maxSteps = 6): string
{
    $messages = [['role' => 'user', 'content' => $prompt]];

    for ($step = 0; $step < $maxSteps; $step++) {
        $response = $this->client()->post('/messages', [
            'model' => config('services.anthropic.model'),
            'max_tokens' => 1024,
            'tools' => $tools,
            'messages' => $messages,
        ])->throw()->json();

        $messages[] = ['role' => 'assistant', 'content' => $response['content']];

        if ($response['stop_reason'] !== 'tool_use') {
            return collect($response['content'])->where('type', 'text')->pluck('text')->implode('');
        }

        $results = [];
        foreach ($response['content'] as $block) {
            if ($block['type'] === 'tool_use') {
                $results[] = [
                    'type' => 'tool_result',
                    'tool_use_id' => $block['id'],
                    'content' => json_encode($execute($block['name'], $block['input'])),
                ];
            }
        }

        $messages[] = ['role' => 'user', 'content' => $results];
    }

    throw new RuntimeException('Agent did not finish within the step limit.');
}

The $execute callback is where your application logic lives. It receives the tool name and the model's arguments, and it must validate them like any other untrusted input. If a tool queries your database, scope it to the current user, so a manipulated request can't read someone else's data. Cap the steps, as the loop does, so a confused model can't run forever.

For tool definitions, the JSON schema describes the arguments:

$tools = [[
    'name' => 'find_order',
    'description' => 'Look up one order by its public reference, such as ORD-1042.',
    'input_schema' => [
        'type' => 'object',
        'properties' => [
            'reference' => ['type' => 'string', 'description' => 'The order reference.'],
        ],
        'required' => ['reference'],
    ],
]];

Good descriptions matter more than they seem. The advice in designing tools for LLM agents is worth reading before you write more than a couple.

Get structured data back

When you need data and not prose, don't parse free text. Force a tool call whose schema is the shape you want, then validate with Laravel's validator. The full approach, including retries, is in getting reliable JSON out of an LLM; the pattern is the same in PHP, with tool_choice set to your tool and Validator::make on the arguments.

Testing without spending money

Http::fake makes this integration easy to test, and you should never call the real API in your test suite.

use App\Services\Claude;
use Illuminate\Support\Facades\Http;

it('returns the text from the response', function () {
    Http::fake([
        'api.anthropic.com/*' => Http::response([
            'content' => [['type' => 'text', 'text' => 'A short summary.']],
            'stop_reason' => 'end_turn',
        ]),
    ]);

    $result = app(Claude::class)->message('Summarize this.');

    expect($result)->toBe('A short summary.');

    Http::assertSent(fn ($request) =>
        $request->hasHeader('x-api-key') && $request['max_tokens'] === 1024
    );
});

it('throws when the API returns an error', function () {
    Http::fake([
        'api.anthropic.com/*' => Http::response(['error' => ['message' => 'bad request']], 400),
    ]);

    app(Claude::class)->message('Hi');
})->throws(RuntimeException::class);

Those tests check your code's handling of responses, which is deterministic. Testing the quality of model output is a different job and belongs in an evaluation set, as described in how to evaluate AI agents.

Streaming and conversations

Two features come up almost immediately after the first summary button: showing text as it's generated, and remembering a conversation.

Streaming makes a slow answer feel fast, because people see words appear within a second. The API can stream its response as server-sent events, and Laravel can pass that along to the browser with a streamed response. The catch is that the HTTP client's convenience methods buffer the whole body, so streaming needs the lower-level stream option and a small loop that reads events as they arrive. It's a good place to lean on the official documentation for the event names, since they're specific and versioned. Until you need it, skip it. Plenty of features work fine with a spinner and a queued job.

Conversations are simpler than they sound, and there's one fact to keep in mind: the API is stateless. It doesn't remember anything between calls. Each request has to carry the whole conversation so far. So you store messages in a table, with a conversation ID, a role, and the content, then load them in order and send them with every new user message. That also means a long chat gets more expensive with each turn, for the reasons explained in how the agent loop works. Decide early how you'll trim: keep the last N messages, or summarize older ones into a short note at the top.

Store the token counts from each response alongside the message. Six weeks from now, when someone asks why the bill doubled, you'll be glad the numbers are in your database and not only in a dashboard on someone else's server.

Security notes specific to Laravel

A few Laravel details deserve attention, because the framework makes the wrong thing easy in some places.

Authorize before you call. If a user asks about "my order 1042," check in your own code that they own order 1042 before any of its data reaches the prompt. Models can't enforce your authorization rules, and anything you put in the context can be repeated back to the user who asked.

Don't put secrets or internal data in prompts by accident. It's easy to interpolate a whole Eloquent model into a string and send every column, including flags and notes that were never meant for a customer. Select only the fields the task needs.

Escape on output. Blade escapes by default with double braces. The risk appears when someone uses the unescaped syntax on model output, or renders Markdown to HTML without sanitizing it. Model text can contain HTML, and if it read attacker-controlled content, it can contain attacker-chosen HTML.

Keep tool permissions narrow. A tool that runs "any SQL" or "any shell command" is a vulnerability with a friendly name. Expose specific, parameterized operations instead, and scope each to the authenticated user.

Watch your logs. Laravel's exception pages and log files can capture request payloads. If prompts contain personal data, make sure error reporting and log retention are set up with that in mind.

Production habits

A few practices keep the integration healthy once real users arrive.

Rate limit your own endpoints. Laravel's rate limiter on routes that trigger model calls stops one user from draining your budget or your quota. Limit per user, and apply a daily cap where it makes sense.

Log usage, not content. Record the model, token counts from the response's usage field, latency, and the feature name. Avoid logging full prompts and responses unless you've decided how to protect that data, since they may contain personal information.

Cache when the same input repeats. If many users request the same summary, store it and serve it from the database or cache instead of paying again.

Handle failure with grace. When the API is down or rate limited, your page shouldn't crash. Show a friendly message, queue a retry, or fall back to a non-AI path.

Keep prompts in code you can review. Put them in a dedicated class or a Blade-like template file, under version control, not scattered inside controllers.

Treat model output as untrusted. Escape it when rendering in HTML, which Blade's {{ }} does by default, and never pass it to eval, shell commands, or raw SQL. If the model reads user-supplied text, remember the risks laid out in prompt injection and the lethal trifecta.

Watch the model version. Keep the model ID in config, so upgrading is a one-line change, and rerun your checks when you change it.

When a package is worth it

There are community Laravel packages for Anthropic and other providers, and they can be a reasonable choice, especially if you want a common interface across providers or built-in streaming helpers. If you use one, read its source briefly, check that it's maintained, and see how it handles errors and retries. The trade is convenience against an extra dependency to keep updated. For a handful of calls, the service class above is often less code than the package's setup.

Where to go from here

You now have a config-driven service, retries, queued processing, a tool loop, and tests that cost nothing. Natural next steps are streaming responses to the browser, storing conversations so users can continue a chat, adding retrieval over your own content as in RAG that works, and turning repeated prompts into cached prefixes as described in prompt caching.

Start with one small feature that has a clear value, such as summarizing long text, drafting a reply, or classifying incoming messages. Measure what it costs and how often it's right. Then grow from there, with each addition small enough to test.