Skip to content

A practical Claude Code workflow for real projects

By SunnyKumar Jonwal 9 min read

Claude Code is Anthropic's coding agent. It runs in your terminal (and in IDE extensions, a desktop app, and on the web), reads your project, edits files, runs commands, and iterates on the results. Under the hood it's the loop from how the agent loop works with a strong set of coding tools attached.

Tools like this reward a good workflow. Two people using the same agent on the same codebase can get very different results, and the difference is rarely a secret prompt. It's habits: how they scope work, how they verify it, and how they keep the agent's context useful. This post lays out habits that tend to work.

Start with the setup, not the prompt

Open a terminal in your project root and start a session with claude. The first thing worth doing in a new repository is running /init, which has Claude look around and draft a CLAUDE.md: a file of standing notes about the project that gets loaded at the start of every session. Edit what it produces. The best version contains the commands to build, test, and lint, the conventions that aren't obvious from the code, and any traps a newcomer would hit. There's a whole post on writing a CLAUDE.md that helps.

Then check your permission settings. Decide which commands can run without asking (reads, test runs, linters), which should prompt (edits, installs, network calls), and which are off limits. The setup you want on a personal laptop with real credentials is stricter than the one you'd use inside a disposable container. Least privilege for AI agents covers the reasoning.

Give it a tight feedback loop

Coding agents are at their best when they can check their own work. Tests, type checkers, linters, and a way to run the app all give the agent something objective to react to, the same way they do for you.

Before starting a task, make sure the loop exists. Can the agent run the relevant tests with one command? Does that command run quickly? If the suite takes twenty minutes, point it at a fast subset and tell it how. If there are no tests around the area you're changing, ask it to write a few first, or have it describe the current behavior so you can compare.

A phrase that pays off: "run the tests after each change and fix failures before moving on." Without a check like that, the agent has only its own confidence to go on, and confidence is cheap.

Plan before you build

For anything bigger than a small edit, separate thinking from doing. Claude Code has a plan mode, toggled from the terminal with Shift+Tab, in which it can read and explore but can't change files. Use it to get a plan first.

A good exchange goes like this. You describe the goal and the acceptance criteria: "Add rate limiting to the login endpoint. Five attempts per minute per IP, return 429 with a Retry-After header, cover it with tests, don't touch the session code." The agent explores, reads the relevant files, and proposes a plan. You read the plan. This is the cheap moment to catch a wrong assumption: maybe it planned to add a dependency you don't want, or missed the existing middleware that already does part of the job.

Fixing a plan costs seconds. Fixing a hundred-line change built on a bad plan costs an hour. Approve when the plan matches what you'd do yourself, then let it run.

Work in small, verifiable steps

Big vague requests ("refactor the billing module") produce big vague diffs. Break work into steps you can check in one sitting. Ask for one change, look at the diff, run the tests, and only then move on.

Three habits make this easier:

  • Commit after each working step. Small commits give you checkpoints you can return to, and they make review painless. If the next step goes sideways, you reset to the last good state and lose nothing.
  • State what's out of scope. Agents like to be thorough. "Fix the failing test only; don't refactor surrounding code" prevents helpful but unwanted changes.
  • Point at files. "Look at app/Services/Invoice.php and the tests in tests/Feature/InvoiceTest.php" saves exploration and gets a more focused result than "look at how invoices work."

Be specific about what done means

The difference between a useful result and an almost-right one is often the acceptance criteria. Vague: "make the search faster." Specific: "the search endpoint takes over two seconds on the products table; find why using the query log, fix it, and show before and after timings."

Include what to avoid, what to preserve, and how you'll judge it. If you have a preference between approaches, say so and say why. As in any prompt, the reasoning helps: "avoid a new dependency, because this service is deployed to an environment where we can't install packages" lets the agent make sensible calls about cases you didn't list. The general habits in prompt engineering that still works apply here.

Manage context like a budget

Long sessions accumulate clutter: old file contents, dead ends, superseded plans. That costs money and, more importantly, dilutes the agent's attention. Some habits keep it sharp.

Start fresh for unrelated tasks. /clear resets the conversation, and it's the single most effective habit for staying focused. When a long session is still on the same task but getting heavy, /compact summarizes the history so you keep the thread without the bulk. You can pass guidance about what to preserve. For work that spans days, keep a short notes file (a plan, decisions, what's left) so a new session can pick up where the old one stopped, an application of the ideas in context engineering and agent memory.

And watch for over-reading. If the agent is opening dozens of files to answer a question, give it a narrower starting point, or hand the exploration to a subagent.

Use subagents and worktrees for parallel or noisy work

Exploration is noisy. "Find every place we call the payment API" can consume a great deal of context reading files that turn out to be irrelevant. Claude Code can delegate such tasks to subagents, helpers that run in their own context window and report back a short summary, which keeps your main session clean. There are built-in helpers for exploration and planning, and you can define your own. Skills, hooks, and subagents explains how.

When you want to run several tasks at once, use git worktrees: separate working directories, each on its own branch, sharing one repository. Start a session in each. One agent fixes a bug while another builds a feature, and their edits can't collide, because they're in different folders. When each finishes, review and merge its branch like any other pull request.

Review the output like a pull request

Claude Code will produce plausible code that looks right and isn't. Treat every change as you'd treat a colleague's pull request, and sometimes with more care.

Read the diff. Look for changes you didn't ask for, weakened or deleted tests, hardcoded values, swallowed errors, and new dependencies. Run the full test suite, not just the tests the agent chose to run. Look at anything touching authentication, permissions, input handling, or money with extra suspicion, since bugs there are expensive. There's a full list in a checklist for reviewing AI-generated code.

It's also worth asking the agent to review its own work in a fresh context, or to explain why it made a choice. Explanations aren't proof, but they often surface a shaky assumption.

Automate the repeatable parts

Once a pattern works, capture it. Custom slash commands turn a frequently used prompt into a one-word command. Skills package instructions and scripts for tasks like "write release notes in our format." Hooks run shell commands at fixed points, for instance formatting a file after every edit or blocking edits to certain paths, and unlike instructions, they run every time.

For scripting and CI, there's a non-interactive mode: claude -p "your prompt" runs a single task and prints the result, which makes it usable in scripts and pipelines. Common uses include summarizing a diff, triaging test failures, or drafting a changelog. Keep permissions tight in automated runs, since nobody's watching to say no.

Adapting the routine to the kind of task

The default sequence stretches to fit different work, with a few adjustments.

For a bug fix, start by reproducing the problem. Ask the agent to write a failing test that captures the bug before it touches the code. That test defines done, keeps the fix honest, and stays behind as a regression guard. Then ask for the smallest change that makes it pass, and watch for fixes that paper over the symptom, like catching an exception instead of preventing it.

For a new feature, spend more of your effort in plan mode. Ask the agent to look at how similar features are built in the codebase and to follow those conventions, because the most common failure is code that works and looks foreign. Build in slices that each leave the app in a working state: the data model first, then the endpoint, then the interface, with tests at each layer.

For a refactor, lean hard on the existing tests, and add characterization tests where coverage is thin, so you can tell whether behavior changed. Keep each commit purely mechanical, with no behavior changes mixed in, so review is easy and a revert is clean. Refactors are also where a large context helps least and a focused, file-by-file approach helps most.

For unfamiliar code, use the agent as a guide before you use it as an author. Ask it to explain the request flow, draw out the module boundaries, and list where a given concept is implemented, then verify a few claims yourself. You'll come away with a better mental model, and your later instructions will be sharper for it.

Common ways it goes wrong

A few patterns account for many bad sessions.

The giant task. Asking for an entire feature in one go, then receiving a sprawling change nobody can review. Split it.

No feedback loop. Letting the agent edit without any way to run the code. It will report success anyway.

Accepting without reading. Approval fatigue is real, and so is the moment you merge something you didn't understand. If you can't explain the change, don't ship it.

Context sprawl. Hours into one session covering five topics, with the agent working from stale assumptions. Clear and restart.

Fighting the tool. When the agent keeps making the same mistake, more emphasis rarely helps. Look at what it's missing: a convention in CLAUDE.md, a failing check it can't run, a file it doesn't know exists.

Trusting it with secrets. Keep credentials out of the repository and out of reach, and remember that files the agent reads become part of what's sent to the model.

A default routine

If you want a starting point, this sequence works on most tasks:

  1. Fresh session, clear goal, acceptance criteria, and out-of-scope notes.
  2. Plan mode first. Read the plan and adjust it.
  3. Let it implement in small steps, running the tests after each.
  4. Review each diff and commit each working step.
  5. Run the full suite and lint, and read the final diff once more top to bottom.
  6. Note any lessons (a missing convention, a gotcha) in CLAUDE.md so the next session starts smarter.

It's not glamorous, and it's also how people get consistently good results from agentic coding tools. The tool keeps improving, and these habits carry over.