Skip to content

The state of AI agents: what works, what doesn't, what's next

By SunnyKumar Jonwal 9 min read

Talk about AI agents tends to swing between two moods. In one, agents are about to do most knowledge work, and anyone not building them is falling behind. In the other, they're demos that break on contact with reality. Both moods sell well, and neither helps you decide what to build on a Tuesday afternoon.

This post is an attempt at a middle view. It's a snapshot, written with the caveat that this field moves faster than almost any other, and that my knowledge has a cutoff. Where I'm describing capabilities, read them as the general shape of things and check current sources for the specifics. I'll try to say which claims are well supported and which are opinion.

First, a definition that helps

As laid out in what is an AI agent, an agent is a model that decides its own next step, uses tools to act on the world, and continues until it finishes or gives up. The useful contrast is with a workflow, where a developer fixes the sequence of steps and the model fills in pieces.

That distinction matters for this whole discussion, because most of what's succeeding commercially sits at the workflow end of the range. Anthropic's widely cited essay on building effective agents makes the same point: simple, composable patterns beat elaborate frameworks, and you should add autonomy only when the task needs it. The systems that get the most attention are the fully autonomous ones. The systems that get the most use are usually much more constrained.

Where agents are working today

Some categories have real adoption, and the evidence is in how many people use them daily.

Coding agents. This is the clearest success. Tools like Claude Code, Cursor, GitHub Copilot's agent modes, and others read a codebase, edit files, run tests, and iterate. Code is a friendly domain: there's a fast, objective check (does it compile, do the tests pass), the environment is text, and mistakes can be reverted with version control. Developers report large gains on well-scoped tasks, along with the need to review carefully, as covered in a checklist for reviewing AI-generated code. Published studies on productivity have been mixed and depend on the task and the developer's experience, so be cautious about any single number.

Customer support. Agents that answer questions from a knowledge base, look up order status, and hand off to a person when stuck are widely deployed. The successful versions are narrow, grounded in company documents, and honest about escalation. The failures usually involve an agent that promised something the policy didn't allow.

Research and analysis. Agents that search, read many sources, and produce a cited report save real time on literature scans, competitive overviews, and due diligence. Output still needs verification, because citations can be wrong and sources can be weak.

Data and operations work. Classifying documents, extracting fields from invoices, triaging tickets, and enriching records. Much of this runs as a workflow with model steps and works well because each step is small and checkable.

Personal and internal assistants. Agents connected to calendars, documents, and chat through connectors like MCP can answer questions across a person's tools. Useful, though this is also where security risk concentrates.

The pattern across all of these: bounded scope, good tools, clear success checks, and a person nearby.

Where they still struggle

The gaps are consistent, and they're the same ones the engineering literature keeps naming.

Reliability over long tasks. Each step has some chance of going wrong, and errors compound. An agent that's 95 percent reliable per step is right about 60 percent of the time over ten steps and worse over thirty. That arithmetic is why demos of short tasks impress and long unsupervised runs disappoint. Progress is real, since newer models hold together over longer sequences than older ones, yet the problem hasn't gone away. See how to evaluate an AI agent for how to measure it.

Judgment about when to stop. Agents sometimes press on after they should have asked a question, or give up too early. Knowing what they don't know is still a weak point, related to the causes described in why LLMs hallucinate.

Long-term memory. Models don't remember between sessions unless you build it. Memory systems help, and they bring their own problems, like stale or wrong memories carried forward. The patterns are covered in agent memory patterns.

Cost and latency. Loops resend context at every step, so expensive tasks get very expensive. Token costs constrain which use cases make economic sense, particularly at consumer scale.

Interfaces built for humans. Agents that operate screens are improving, and they're still slower and less dependable than API-based ones, as discussed in computer use and browser agents.

Evaluation. Many teams ship agents without meaningful tests, then can't tell whether a change helped. It's an unglamorous problem and one of the biggest.

The security problem is unresolved

If there's one topic where sober language matters, it's this one. Agents combine three things: access to private data, exposure to untrusted content, and the ability to act. That combination, which Simon Willison called the lethal trifecta, is dangerous because models can't reliably distinguish instructions from data. A hostile web page or email can steer an agent that reads it. Prompt injection remains an open problem: there are mitigations, and no one has a complete fix.

Researchers and companies have demonstrated attacks on coding assistants, browser agents, and connector-based assistants. That doesn't mean agents are unusable. It means their design has to assume that the model can be tricked, and to limit the damage when it is, with least-privilege permissions, sandboxes, approvals for consequential actions, and careful choices about what data an agent can see at once. The risks around AI coding tools are a concrete case.

Anyone telling you the security issue is solved is selling something. Anyone telling you it makes agents pointless is overreacting.

What's changed recently, and what hasn't

Things that have visibly moved: models are better at using tools and following long instructions. Standards like MCP have made connecting tools easier and more uniform. Coding agents went from autocomplete to doing multi-file tasks. Context windows grew, and prompt caching made long contexts cheaper. Tooling for tracing, evaluation, and permissions has matured.

Things that haven't: models still make confident mistakes. Real-world tasks still have messy, underspecified inputs. Humans still need to review anything that matters. Integration work, meaning permissions, data quality, and edge cases, still consumes most of a project's time, and it's rarely the exciting part.

Why demos mislead

A demo shows the best run of a carefully chosen task, on a clean environment, with the presenter ready to restart if it goes sideways. Production is the opposite: unknown inputs, messy data, thousands of runs, and nobody restarting anything. The gap between the two is the reason so many agent projects feel great in week one and frustrating in month three.

You can protect yourself with a few questions whenever you see an impressive demo. How many times was it run, and was this the best one? What did the inputs look like compared with yours? What did it cost per run? What happens when it's wrong, and who notices? Was a person quietly correcting it? A vendor with good answers is worth talking to. One who deflects is telling you something.

The same skepticism applies to dismissals. A story about an agent failing spectacularly is usually true and usually about a system that was given too much access and too little supervision. It tells you about that design, not about the whole field.

How to place bets

Here's practical advice for someone deciding what to build or learn, and I'll flag it as opinion.

Start with a workflow, then add autonomy where it pays. Most products don't need a fully autonomous agent. A fixed sequence with model calls at the hard steps is easier to test, cheaper, and more predictable. Move to an agent loop when the path genuinely can't be known in advance.

Choose tasks with fast, cheap verification. If you can check the result automatically, as with code and tests, structured data with validation, or a calculation, agents are far safer bets. If verifying costs as much as doing the task, be careful.

Keep humans on the consequential steps. Sending money, deleting data, emailing customers, and merging code are all good approval points. The cost of a click is small compared with a wrong irreversible action.

Invest in evaluation early. A test set of a few dozen real cases will teach you more than any leaderboard, and it lets you swap in better models as they arrive.

Design for model change. Models improve and get cheaper quickly. Keep the model ID in config, keep prompts in version control, and avoid building so much scaffolding around a weakness that you'll be stuck when the weakness disappears.

Budget for cost. Do the token arithmetic before launch. Set limits per user and per task.

Learn the fundamentals. Tool design, context management, retrieval, evaluation, and security are the durable skills. Specific frameworks come and go, and the concepts carry over from one to the next.

What this means for your career

If you build software, the practical question is what to learn. My suggestion is to get hands-on with a coding agent and use it on real work for a month, paying attention to where you trust it and where you don't. Then build one small tool-using agent of your own, so the loop stops being a diagram. Then learn to measure it. Those three steps, done in order, give you more than a year of reading headlines.

Roles are shifting, though not as fast as the loudest predictions claim. Work that's routine, well specified, and easy to check is the most exposed, and work that depends on judgment, context, and accountability is changing more slowly. People who can define problems clearly, review output critically, and design safe systems are in demand, because every agent needs someone who does those things.

What I'd watch next

A few developments are worth following, with the caution that predictions in this field age badly.

Longer, more dependable autonomy. Measurements of how long a task an agent can complete reliably have been rising. Whether that continues, and how it translates from benchmarks to messy real work, is an open question.

Better memory and context handling. Approaches that let agents manage their own context, keep notes, and pick up where they left off are improving, and they're central to long-running work.

Agent-to-agent and cross-vendor protocols. Standards for how agents discover and call each other's tools are forming. Which ones stick isn't clear yet.

Security research. New defenses, and new attacks, arrive regularly. Expect this to shape which deployments are considered acceptable, and by whom.

Regulation and liability. Rules about automated decisions, disclosure, and responsibility are emerging in different places. If you build for regulated industries, follow them closely.

Economics. Falling prices per token expand what's viable, while agents also use far more tokens per task. Which effect wins will differ by use case.

A sensible stance

Take agents seriously as a tool without treating them as magic. They're genuinely useful in bounded settings with good tools and honest checks, and they're unreliable in open-ended settings without either. The people getting value from them tend to be building modest systems carefully: narrow scope, real tests, limited permissions, a human on the important decisions, and a habit of measuring cost.

If you're new to this, you don't need to predict where it all ends up. Build one small agent, as in building your first agent with the Claude API, watch where it fails, and learn the failure modes firsthand. That knowledge transfers to whatever tools show up next year, and it will make you a better judge of the claims, hopeful and skeptical, that you'll hear along the way.