Skip to content

Google says AI is speeding up vulnerability discovery. Here's what that means

By SunnyKumar Jonwal 9 min read

On September 30, Google Threat Intelligence Group published a report arguing that AI is changing both the pace and the profile of vulnerability discovery. The coverage, from SecurityWeek, Help Net Security, Infosecurity Magazine, and others, centers on a handful of numbers that deserve a careful read, because they're the kind of figures that get quoted without their caveats.

This post walks through what the report found, what I think it supports and what it doesn't, and what a developer or a security team can do about it. I'm working from secondary coverage as of October 8, 2026, not the full report, so if a number matters to a decision you're making, read Google's original.

The numbers

The report says the number of vulnerabilities disclosed each month roughly doubled over the year. Monthly disclosures rose from 5,045 in January 2026 to 10,477 in July, and peaked at 10,740 in August.

Exploitation went up too. Google counts the average number of vulnerabilities exploited in the wild per month, and says it rose from 10.5 in 2025 to 18 in 2026 so far. That's about a 71 percent increase, which is the figure you'll see in headlines. Exploitation of high-risk vulnerabilities more than doubled, from 28 in 2025 to 75 in the first eight months of 2026. The report says most of the growth in exploitation came from n-days, meaning flaws that are already public and patched (or patchable), and not from zero-days.

Then comes the AI angle. The researchers identified a set of vulnerabilities as likely AI-discovered, and compared them with the rest. Half of the likely AI-discovered flaws, 50 percent, resulted in remote code execution, against 26 percent of other CVEs. In other words, the flaws that AI-assisted discovery turns up are more likely to be the severe kind.

A separate item surfaced in the same week's news. The UK AI Security Institute reported on tests of a model it calls GPT-6 Astra in a simulated environment, where the model completed simulated supply-chain attacks in 29.2 percent of runs, including creating fake identities and submitting malicious code. That's a controlled evaluation of capability, and a different kind of evidence from the Google data, and I'd keep the two apart.

What the report supports

Take the claims one at a time.

Disclosures are up sharply. That part is a count, and it's hard to argue with. Defenders have more advisories to process than they did at the start of the year, and the first half of this post's message is simply that the workload went up.

Exploitation is up, mostly through flaws that already have fixes. That one has a practical implication I'd take seriously. When most of the growth comes from n-days, the weakness being exploited isn't an unknowable bug. It's the gap between a patch being available and a patch being installed. Attackers are getting faster at turning public information into working exploits, and many organizations aren't getting faster at patching.

The likely-AI-discovered group skews toward remote code execution. If that holds up, it matches an intuition: automated tools are good at finding certain classes of bugs, such as memory-safety errors and injection flaws, and those classes tend to be the dangerous ones.

What it doesn't prove

This is where I'd slow down, because the report's headline invites conclusions that the data may not carry.

"Likely AI-discovered" is an inference. As I read the coverage, researchers can't see how every researcher or company found every bug, so they estimate which ones probably came from AI-assisted work. Any classification like that has an error rate, and a different method could yield different numbers. It's a finding to take seriously, not a measurement to treat as exact.

More disclosures don't automatically mean more vulnerabilities exist. The count of published CVEs depends on how many people are looking, how many organizations are assigning identifiers, and how much effort goes into reporting. A big increase in the number of bug hunters, human and automated, would raise disclosures even if the underlying number of flaws stayed constant. It might also be that old bugs are finally being found. That's good news in a sense, because a bug found by a defender is better than one found by an attacker.

Correlation isn't cause. AI tools got much better during the same period that disclosures rose. That's consistent with AI driving some of the increase, and other explanations, such as new bug bounty programs and broader coverage by vendors, could contribute. The report may address these, and I haven't seen all of its analysis.

Exploited-in-the-wild counts also depend on what gets detected and reported. If more security firms are watching, the same activity can look like a rise.

None of that means the trend is imaginary. It means the right posture is "probably real, size uncertain."

The part I find most useful: the n-day gap

If I had to pick one takeaway for an engineering team, it would be the n-day finding. When a patch is released, attackers can compare the old and new versions and work out what changed, a technique called patch diffing. That used to take skilled people days or weeks. Tools that can read code and explain differences shorten that time. If the report's reading is right, the window between "patch published" and "exploit exists" is shrinking.

For defenders, this puts a premium on patch speed for internet-facing systems. A monthly patch cycle that was reasonable five years ago may be too slow for the services attackers can reach. The NetScaler case in the Citrix post is an extreme example, since exploitation preceded the patch. In the n-day cases, it's the slowness of the installation that does the damage.

A patch-window example with made-up numbers

To make the n-day point concrete, here's a hypothetical. These numbers are invented for illustration, so don't read them as data.

Say a vendor ships a fix for a serious flaw in a web-facing product on a Tuesday. In the old world, working out how to exploit it from the patch might take a skilled researcher a week, and attackers began scanning for vulnerable servers the following Monday. A team patching on a monthly cycle, with the maintenance window two weeks out, was sometimes exposed for a short stretch and mostly fine.

Now suppose tools shrink that analysis step from a week to a day. Scanning starts Wednesday. The same team, still waiting for its window two weeks out, is exposed for thirteen days instead of seven. Nothing about the team's behavior got worse. The goalposts moved. The only lever that helps is the one under the team's control, the delay between "fix available" and "fix installed."

That's why I'd put a number on it. Pick a target for how fast internet-facing systems get critical fixes, track the actual time, and look at the worst cases each month. If your slowest critical patch took six weeks, find out why. It's usually a process problem, such as a missing approver, an untested rollback, or an application that nobody wants to restart, and process problems can be fixed.

Questions to ask your own team

A short list that I think gives a quick read on where you stand.

How many internet-facing services do we run, and can someone list them without a day of digging? When a critical advisory lands for something we use, who's responsible for deciding, and how long until they know whether we're affected? What's our slowest critical patch from the last quarter, and what slowed it down? Do we have any services nobody owns? Which of our dependencies are unmaintained?

If the answers come quickly and with specifics, you're in good shape. If they trigger a long silence, you've found where to start.

What this means for bug hunters and researchers

If you do security research, the same trend gives you a mixed picture. Automated tools make some discovery cheaper, which raises the bar for finding bugs that others haven't already found. Depth, context, and judgment count for more: understanding how a system is used, chaining small issues into a real impact, and writing a clear report that a maintainer can act on.

A caution applies for everyone using AI tools to look for flaws. Verify before you report. Programs and maintainers have complained about a flood of automated reports that describe problems that don't exist or can't be reproduced, and those waste the time of the people trying to fix real bugs. A report with a working proof of concept and a clear impact statement gets attention, and a pasted model summary doesn't.

Stay inside the rules, too. Test only systems you own or are authorized to test, follow each program's scope, and use responsible disclosure. Faster tools make it easier to cross a line by accident.

What to do about it

A few practical moves, in rough order of payoff.

Shorten the time to patch for anything exposed to the internet. Set a target, such as days for critical flaws in exposed systems, and measure how you do. Automate what you can: unattended updates for operating system packages, automated pull requests for dependency bumps, and base images rebuilt on a schedule.

Know what you run. You can't patch what you can't list. A current inventory of services, versions, and dependencies, including those inside container images, makes each advisory a lookup instead of an investigation.

Reduce your attack surface. Every service you don't expose is a service whose next vulnerability doesn't matter to you. Put management interfaces behind a VPN or an identity-aware proxy, turn off features you don't use, and remove old test systems.

Use the same tools on your own code. If AI-assisted analysis finds flaws that attackers would want, run it on your code first. Static analysis, fuzzing, and model-assisted code review can all find bugs before release. Be realistic about the limits: these tools produce false positives and miss things, and they work best alongside tests and human review. The practices in a checklist for reviewing AI-generated code apply in both directions.

Be strict about AI-written code. If more bugs are being found by machines, it would be odd to ship machine-written code without review. The risks in slopsquatting, leaked secrets, and other AI coding risks don't shrink because your review queue is longer.

Prefer memory-safe languages and well-tested libraries for new components where you can. A class of bugs that doesn't exist can't be found by anyone.

Maintainers and small teams

The report's picture is harder on volunteers than on companies. Open-source maintainers receive more bug reports, some of them low-quality or automated, and have less time to triage. If you depend on open-source libraries, which means all of us, the useful response is to support the projects you rely on: report bugs with clear reproduction steps, send patches, and fund maintainers where you can. A flood of machine-generated reports that don't verify the problem wastes the scarce attention that fixes real flaws, so if you use automated tools to find bugs, confirm each one before filing it.

A realistic outlook

I don't think this report means that attackers have won, or that software is about to collapse under a wave of exploits. It suggests that the economics are shifting: finding and exploiting bugs gets cheaper, and defenders need to match that with faster patching and smaller attack surfaces. The same tools help defenders too. A team that uses them well can find its own bugs earlier, and in that race, the side that moves first benefits.

The sensible response doesn't need to wait for better data. Patch faster, inventory what you run, shrink what you expose, and review your own code with the tools attackers are going to use anyway. If the next report shows the trend flattening, you'll have lost nothing by doing it.