Skip to content

Agent memory: short-term, long-term, and files

By SunnyKumar Jonwal 10 min read

A language model has no memory. Each request starts from a blank slate, and whatever the model appears to remember is text your program put back into the prompt. That fact is the starting point for every design decision about agent memory: it isn't a feature the model has, it's a system you build around it.

The good news is that the building blocks are simple. Most useful memory setups are a few files, a small database, or a search index, plus rules about what to save and when to bring it back. This post walks through the options and how to choose between them.

Two kinds of memory to keep apart

People use the word "memory" for several different things, and mixing them up leads to muddled designs.

Short-term memory is what the agent knows during the current task or conversation: the messages, tool results, and intermediate work in the context window. It vanishes when the session ends, or when the window overflows and you trim it.

Long-term memory is what survives across sessions: a user's preferences, decisions made last week, facts learned about a codebase, procedures that worked. It lives outside the window, in storage, and has to be deliberately loaded back in.

Within long-term memory, it helps to borrow three labels from cognitive science:

  • Semantic memory: facts. "The user's timezone is IST." "Our API uses cursor pagination."
  • Episodic memory: past events. "Last Tuesday we tried approach A and it failed because of rate limits."
  • Procedural memory: how to do things. "Release notes follow this template." "Run the linter before committing."

You don't need to implement each separately, but asking which kind you're storing tells you how to store and retrieve it.

Pattern 1: keep the recent conversation

The baseline is the message history itself. It costs nothing to build, since you're already sending it. The problems appear as it grows: cost climbs, the model gets distracted by old material, and eventually you hit the limit.

The simplest fix is a sliding window: keep the last N turns and drop the rest. It's cheap and predictable, and it forgets things that mattered ("the user said their deadline is Friday" twenty turns ago). A rolling window works for chatty, low-stakes interactions and fails for anything that needs continuity.

Pattern 2: summarize as you go

Instead of dropping old turns, compress them. When the history passes a threshold, ask the model to write a summary of what happened, and replace the raw turns with that summary. Keep the most recent few turns verbatim so the conversation stays fluid.

A good summary keeps decisions, open questions, key facts, and the user's stated goals. It drops the pleasantries and the raw tool output. Summaries drift, so a long-running agent that summarizes its own summaries can slowly distort the story. Guard against that by keeping structured facts (see below) separate from narrative summaries, and by occasionally checking summaries against the source.

The trade-offs of when to compress, and what it does to caching, are discussed in context engineering and prompt caching.

Pattern 3: notes in files

This one deserves more attention than it gets. Give the agent a place to write notes, usually a plain text or Markdown file, and let it read them back when needed.

A coding agent might maintain NOTES.md with the plan, decisions, and what's left. A research agent might keep a findings.md and append to it as it goes. When the context resets, the agent reads the file and picks up where it left off.

Why it works so well:

  • It's inspectable. You can open the file and see exactly what the agent believes. Debugging memory is much easier when it's readable text.
  • It's editable. You can correct a wrong note or add one, and the agent will honor it.
  • It survives anything. Crashes, restarts, model upgrades, and window overflows don't touch it.
  • It needs no infrastructure. A folder is enough.

Project-level instruction files, such as the CLAUDE.md that Claude Code reads at the start of a session, are the same idea applied to procedural memory: standing knowledge about how to work in this codebase, loaded on every run. Writing a good CLAUDE.md goes into what belongs there.

A few conventions make notes files reliable. Ask the agent to keep them short and current, not an ever-growing log. Give them a structure, such as Goal, Decisions, Open questions, and Next steps. Tell it to update the file at meaningful checkpoints, not after every action.

Pattern 4: structured facts

Some memories are really key-value pairs: a preferred language, a default project, an approved vendor list. Storing them as free text makes retrieval fuzzy. Storing them in a small structured record, such as a JSON document or a database row per user, makes them exact.

Load the whole record into the system prompt when it's small. It's cheap, deterministic, and easy to audit. Add fields deliberately, so the record stays tidy, instead of letting the model dump arbitrary strings into it.

Pattern 5: searchable memory

When there's too much to load, such as thousands of past conversations, documents, or events, you need retrieval. This is RAG applied to the agent's own history: store memories as chunks with embeddings, and at each step retrieve the few most relevant to the current situation.

The mechanics match what's covered in RAG that works, with a few twists specific to memory:

  • Write good memories. Storing raw transcripts makes retrieval noisy. It's usually better to store distilled statements ("User prefers tabs over spaces in Go projects") along with when and where they were learned.
  • Add metadata: timestamp, source, confidence, and a scope (this user, this project, everyone).
  • Rank by more than similarity. Recent memories, frequently used ones, and higher-confidence ones often deserve a boost.
  • Deduplicate. The same preference learned five times shouldn't take up five slots.

Pattern 6: let the agent manage its own memory

Instead of your code deciding what to save, give the model memory tools: save_memory, search_memory, update_memory, delete_memory. The agent decides when something is worth remembering and when to look something up.

This is flexible and fits agentic designs. Some model APIs now offer a memory tool that lets the model read and write files in a directory you control, which is a variation on the notes-file idea with the model in charge of the bookkeeping. The cost is that the model's judgment about what to save is imperfect. It may hoard trivia, miss important facts, or save something wrong.

A minimal version is only a few lines:

from pathlib import Path

MEMORY_DIR = Path("agent_memory")
MEMORY_DIR.mkdir(exist_ok=True)

def save_memory(topic: str, content: str) -> str:
    """Save or overwrite a memory under a short topic name."""
    safe = "".join(c for c in topic.lower() if c.isalnum() or c in "-_")[:60]
    if not safe:
        raise ValueError("Topic must contain letters or numbers.")
    (MEMORY_DIR / f"{safe}.md").write_text(content.strip())
    return f"Saved memory '{safe}'."

def recall_memory(topic: str = "") -> str:
    """List memory topics, or return one memory's contents."""
    if not topic:
        names = sorted(p.stem for p in MEMORY_DIR.glob("*.md"))
        return "\n".join(names) or "(no memories yet)"
    path = MEMORY_DIR / f"{topic}.md"
    return path.read_text() if path.exists() else f"No memory named '{topic}'."

Expose those as tools, and tell the agent in the system prompt when to use them: "At the start of a task, list your memories and read any that look relevant. When you learn something durable about the user or project, save it."

What to store, and what to skip

The biggest determinant of memory quality is discipline about what goes in. A memory full of noise is worse than none, because it gets retrieved and misleads.

Worth storing:

  • Stable preferences and constraints
  • Decisions and the reasons behind them
  • Corrections the user made ("don't suggest library X, we can't use it")
  • Non-obvious facts about the environment
  • Procedures that worked, and pitfalls to avoid

Not worth storing:

  • Anything derivable from the current code or documents, which can go stale and conflict with the source
  • One-off details that won't matter next time
  • Raw tool output
  • Speculation the agent hasn't verified
  • Secrets, credentials, and sensitive personal data

A practical test: would a new teammate benefit from being told this on their first day? If yes, save it. If it's trivia or a duplicate of something in the repo, skip it.

Staleness, conflicts, and forgetting

Memories rot. A preference changes, a project gets renamed, a decision gets reversed. If you never remove or update anything, the agent will confidently act on outdated beliefs.

Some habits that help:

  • Timestamp every memory and show the date when it's retrieved, so the model can weigh recency.
  • Prefer updating over appending. When a new fact contradicts an old one, replace it, and keep a short note of the change if history matters.
  • Verify before acting. A memory saying "the config lives in app/settings.py" is a hint, not a fact. Have the agent confirm against the real files before relying on it.
  • Expire or review. Periodically prune memories nobody has used, or ask the user to confirm the important ones.
  • Make deletion easy, for you and for the user.

Privacy and safety

Persistent memory turns a stateless system into one that accumulates information about people, which brings responsibilities.

Store as little personal data as you can, and be clear with users about what's remembered. Give them a way to view and delete it. Depending on where you operate and who your users are, data protection rules may apply to what you keep, how long you keep it, and how it can be erased, so check the requirements rather than assuming.

Scope memory correctly. A memory learned in one user's session must never surface in another's. Namespace by user and tenant, and filter at retrieval time in code, not by asking the model to be discreet.

Beware memory poisoning. If an agent saves whatever it reads, then a malicious web page or email can plant an instruction in memory that gets replayed in every future session. That turns a one-time prompt injection into a persistent compromise. Defenses include saving only from trusted sources, having a human or a separate check approve writes to important memories, and tagging memories with their provenance. This is part of the broader problem discussed in prompt injection and the lethal trifecta.

Keep secrets out. Never let an agent write API keys or passwords to memory, and scan writes for patterns that look like secrets.

Two example designs

A coding agent. Procedural memory in an instruction file (CLAUDE.md) for conventions and commands. A NOTES.md per task for plan and progress. No vector store, because the code itself is the source of truth and the agent can search it on demand. This is deliberately light, and it works because most of what the agent needs lives in the repository.

A support assistant. A structured record per customer (plan, language, open issues) loaded into every conversation. Summaries of past tickets, stored as short distilled statements and retrieved by similarity when relevant. Strict per-customer scoping enforced in the retrieval layer, and a retention policy that deletes data after a set period.

Notice how different the two are. Match memory to the job, and resist the urge to build the fanciest version.

How to test memory

Memory bugs are subtle, so test them on purpose.

  • Recall test. Teach the agent a fact in one session, and check that it uses it correctly in the next.
  • Update test. Contradict an earlier fact and confirm the new one wins.
  • Irrelevance test. Confirm unrelated memories don't leak into answers.
  • Isolation test. Confirm one user's memory never appears for another.
  • Injection test. Feed it a document containing a fake "remember this" instruction and confirm it doesn't save it blindly.

Add these to the evaluation set described in how to evaluate an AI agent.

A sensible starting point

If you're building this from scratch, start small. Keep the message history, add a summary when it grows, and give the agent a notes file. That's often enough. Add structured facts when you find yourself repeating the same user details, and add searchable memory only when the volume forces you to.

Every layer you add is another thing to keep correct, private, and current. The best memory systems tend to be the ones with the fewest moving parts that still do the job.