Slopsquatting, leaked secrets, and other AI coding risks
Reviewing generated code for bugs is one problem. A separate set of risks comes from the way AI coding tools sit inside your workflow: they read your repository, install packages, run commands, and talk to external services. Those are supply-chain, secrets, and access problems, and they don't show up in a diff.
This post covers the ones worth planning for, roughly in the order you're likely to meet them. For line-by-line review of the code itself, see a checklist for reviewing AI-generated code.
Hallucinated packages and slopsquatting
Ask a model for code that does something specific, and it may import a library that sounds right and doesn't exist. Usually that produces an install error and mild annoyance. The security problem is what happens when someone anticipates it.
Researchers studying code-generation models found that a meaningful share of the package names they suggest don't exist on the public registries, and that the same invented names tend to recur across runs. That predictability is what an attacker needs: register the fake name, put malicious code inside, and wait for developers and agents to install it. The technique has picked up the name slopsquatting, a cousin of typosquatting.
An autonomous agent makes it worse. A person seeing "package not found" stops and thinks. An agent that runs pip install or npm install as part of its loop may simply proceed, and if the name resolves to something malicious, that code runs on your machine, often with your credentials nearby.
What helps:
- Verify every new dependency before installing it: exists on the official registry, sensible age and download history, a real repository, a maintainer you can identify, and a purpose that matches what the code needs.
- Keep an allowlist of approved packages for automated work, or require approval for any install.
- Commit lockfiles and review changes to them, since a lockfile diff shows what actually got added, including transitive dependencies.
- Prefer the standard library or a package already in your tree for small jobs.
- Run dependency scanning and, in higher-risk environments, install through a private registry proxy that only serves vetted packages.
- Run installs in a sandbox with no access to secrets, so a bad package has less to steal.
Secrets: in prompts, logs, and commits
Secrets leak through AI tools by three routes.
The first is context. Anything the agent reads is sent to the model provider as part of the prompt. If it opens .env, a private key, a config file with credentials, or a database dump, those values travel with the request and may be retained in logs on the other end, depending on the provider's policies. The fix is prevention: deny the agent access to those files entirely, through permission rules, and keep real secrets out of the working directory when you can.
The second is output. An agent might write a real key into a config file, a test fixture, a script, or a commit message. Models also produce realistic-looking fake credentials, which can trigger alarms or, worse, be mistaken for real ones. Add secret scanning to your pre-commit checks and your CI so anything that slips through gets caught before it's pushed. A minimal pre-commit setup with a scanner such as gitleaks looks like this:
# .pre-commit-config.yaml
repos:
- repo: https://github.com/gitleaks/gitleaks
rev: v8.18.0 # pin to a current release
hooks:
- id: gitleaks
The third is logs and transcripts. Session histories, debug output, and tool-call logs can contain everything the agent saw. Treat them as sensitive, restrict access, and set retention limits.
If a secret does leak, rotate it immediately. Assume anything sent to an external service or pushed to a repository is compromised, even briefly.
Insecure patterns learned from old code
Models learn from a huge body of existing code, which includes plenty of outdated and insecure examples. Left to defaults, they sometimes reproduce them: string-built SQL queries, MD5 or unsalted hashing for passwords, disabled certificate verification, wildcard CORS, predictable tokens from a non-cryptographic random function, overly permissive file modes, or missing authorization checks.
The countermeasures are the same ones you'd use with any code: static analysis in CI, linters with security rules, framework defaults that make the safe path the easy one, and tests for the security properties that matter, such as "user A can't read user B's record." You can also steer generation directly. State security requirements in your project's instruction file, for example "use parameterized queries only" or "every route needs a policy check," which is exactly the kind of non-obvious convention a good CLAUDE.md is for. Instructions help, and scanners and tests are what actually catch the misses.
Instructions hidden in your own repository
Coding agents read everything in the project: source files, READMEs, issue text, documentation, and their own configuration and rules files. Any of that can contain instructions aimed at the model, and it can arrive from places you don't control: a dependency's README, a pull request from a stranger, an issue comment, a copied snippet.
Researchers have demonstrated attacks where instructions are hidden in a way humans don't easily see, for example using invisible Unicode characters inside a shared rules or configuration file, so that anyone who adopts the file also adopts instructions telling the agent to slip a vulnerability or a backdoor into the code it writes. This is prompt injection applied to the development environment, and the general problem is laid out in prompt injection and the lethal trifecta.
Practical steps: review rules and instruction files like code, including with tools that reveal hidden characters. Be wary of adopting configuration from unknown sources. Restrict what an agent can do when it's working on untrusted code, such as a repository you just cloned. And keep humans in the loop for anything that would change security-sensitive files.
Too much autonomy
An agent with shell access and broad permissions can do real damage without any attacker involved. It may run a destructive command to "clean up," overwrite files, drop a table to make a migration work, or push to the wrong branch. There have been public reports of AI coding tools deleting production data despite explicit instructions not to. In those stories, the instruction was in the prompt, and the capability was in the credentials, and capability won.
The remedy is the theme of least privilege for AI agents: run agents in sandboxes, use credentials that can't touch production, require approval for destructive actions, work on branches or copies, and keep backups you've actually tested. A rule of thumb: never give a coding agent a credential you wouldn't hand to a new contractor on their first day.
CI/CD deserves a special mention. Agents wired into pipelines, for example bots that respond to pull requests or issues, may run with tokens that can write to the repository or access secrets. If those bots read text written by outsiders, that's the trifecta again. Use minimal token scopes, run on untrusted contributions without secrets, and require human approval before merging or deploying anything the bot produced.
Tools and connectors as attack surface
Every connector you add to an agent widens what it can reach and what can reach it. An MCP server or plugin you install runs code, exposes tool descriptions to the model, and returns results into its context. A compromised or malicious one can steal data or steer behavior. Vet connectors the way you'd vet a dependency: source, maintainer, permissions requested, update history. Enable only what a task needs. The background is in MCP explained.
Data exposure, privacy, and licensing
Code you send to a hosted model leaves your machine. For many projects that's fine. For others, such as regulated data, customer information, unreleased products, or code under strict contracts, it's a real question. Check what your provider retains, whether it trains on your inputs by default, and what enterprise or zero-retention options exist. Keep customer data and production dumps out of prompts, and use synthetic or anonymized data for debugging with AI tools.
Licensing is a quieter risk. Generated code can resemble existing open-source code, and the legal picture is still developing. Tools may offer filters or indemnification in some plans, so read the terms. For anything you plan to distribute or sell, ask your legal team how they want AI-assisted code handled, and keep a record of where it was used.
The human factor
Some of the biggest risks come from people, not tools. Automation bias makes reviewers trust polished output. Approval fatigue makes people click "allow" without reading. Pressure to move fast pushes teams to skip review when the code seems to work. And shadow usage, where developers paste company code into personal accounts of AI tools, creates exposure the organization can't see.
Good practice here is boring: clear guidelines on which tools are approved and what data may go into them, training that includes real examples of these failures, and a culture where "I don't fully understand this diff" is an acceptable reason to slow down.
What an incident looks like
It helps to walk through a plausible failure, because the pieces rarely arrive one at a time. A developer asks an agent to add a PDF export feature. The agent suggests a package with a convincing name, runs the install, and the package pulls in a post-install script. That script reads environment variables from the shell, where a cloud token happens to be set for an unrelated task, and sends them to a server the developer has never heard of. Nothing looks wrong. The export feature might even work.
Every step there was ordinary. The agent was allowed to install packages, the install ran with the developer's full environment, and the token was reachable because it was convenient. Any one control would have broken the chain: an allowlist for installs, a sandbox without credentials, a scoped token that could only do the narrow job it was meant for, or egress rules that blocked unknown hosts. Layered controls matter because you can't predict which link an attacker or an accident will use.
After an incident like this, the response is unglamorous. Rotate every credential the environment could reach, find out which package and version ran, check other machines and CI for the same install, and review what the token could have touched in the logs. Teams that have practiced this once respond in an hour. Teams that haven't spend a week guessing.
A few questions to ask before adopting a tool
Before rolling out any new AI coding tool or connector, a short set of questions saves a lot of trouble later. Where does my code go when I use it, and how long is it kept? What can it read and write on my machine by default, and can I narrow that? Does it run commands, and if so, can I require approval or run it in a container? What network access does it have? Who maintains it, and how quickly do they ship security fixes? Can I see logs of what it did? If the answers are vague, treat the tool as untrusted and start it in the most restricted setup you can manage.
A minimal program
If you'd like a starting point that fits a small team, it looks something like this.
- Approve specific tools and configurations, and write down what data may and may not be shared with them.
- Deny agent access to secrets by default, and add secret scanning to pre-commit and CI.
- Route dependency changes through review, with lockfiles committed and scanning enabled.
- Run agents in sandboxes with scoped credentials, no production access, and network egress limits.
- Require human approval for destructive and outward-facing actions, and design the approvals to be readable.
- Review instruction and configuration files as code, and be careful with ones from outside.
- Log agent actions, and review a sample regularly.
- Test your assumptions with an exercise: plant a fake package name, a hidden instruction in a README, and a decoy secret, and see what your setup does.
None of it is exotic, and most of it is standard security hygiene extended to a new kind of contributor. The risk isn't that AI coding tools are uniquely dangerous. It's that they act quickly, with access, on text they can't fully vet, so every gap in your existing controls gets exercised faster than before.