AI-Enabled Security Testing: Count What Gets Verified and Fixed
Table of Contents
On 31 January 2026, curl ended its bug bounty. Daniel Stenberg, who leads the project, wrote that in earlier years somewhere north of 15% of submissions had turned out to be confirmed vulnerabilities, and that in 2025 the rate fell below 5% as AI-generated reports flooded in. His May post described what had been happening alongside the slop. AI-powered code analysers had triggered between two and three hundred bugfixes in curl over the previous eight to ten months, probably a dozen or more of their findings had become CVEs, and security researchers using AI were sending in a high volume of high-quality reports.
The junk and the genuine findings reached the same small team. AI made both cheaper to produce, and every report still needs someone who understands the code to check it before anyone can fix it. When a scan with Anthropic’s Mythos Preview model reached curl in May 2026, it reported five confirmed vulnerabilities. After several hours of work, Stenberg’s security team was left with one, rated low severity. Three were false positives and the fifth was an ordinary bug.
Judge AI-enabled security testing by the verified, correctly rated vulnerabilities that get fixed or mitigated, most serious first, and by the human effort behind them; findings that arrive faster than the team can verify and address them add to the backlog without reducing exposure.
The questions a CISO asks about using AI to test software and networks, as distinct from testing the security of AI systems, are practical ones. Is full automation the goal? Are more findings better? Does AI improve coverage, speed and the quality of testing, and how does anyone know its results are right, or better than the current approach? The published evidence answers some of these questions, and the rest need testing in the organisation’s own environment.
What AI testing has demonstrated
The clearest controlled result comes from DARPA’s AI Cyber Challenge, run with ARPA-H, whose final round ended at DEF CON in August 2025. Competing systems analysed more than 54 million lines of real open-source code into which the organisers had planted vulnerabilities. Between them, the systems found 54 of the 63 planted flaws, or 86%, and patched 43, or 68%, against 37% found and 25% patched at the semifinal a year earlier. They also found 18 real vulnerabilities that nobody had planted, and supplied patches for 11. DARPA put the teams’ spending at about $152 per competition task.
DARPA’s scoring rewarded proof as well as discovery. It said the winner, Team Atlanta, performed best at finding and proving vulnerabilities, generating patches, pairing vulnerabilities with patches and keeping its submissions accurate. The first announcement counted 70 planted vulnerabilities, which made the results 77% found and 61% patched, until the organisers corrected the count to 63. The number of finds stayed at 54.
A study by Justin Lin and colleagues at Stanford, Carnegie Mellon and Gray Swan AI, revised in March 2026, compared AI agents with human testers on a live network, a setup the authors describe as the most realistic in AI security research. Ten professional penetration testers, six existing AI agents and the authors’ own agent framework, ARTEMIS, tested a university computer science network of about 8,000 hosts across 12 subnets. The better of two ARTEMIS configurations placed second on the leaderboard, ahead of nine of the ten professionals, with nine valid findings and 82% of its submissions valid.
Codex and CyAgent did worse than most of the humans. Claude Code and one other framework refused the task outright, and a third stalled during reconnaissance.
The cost figure in the study’s abstract needs its conditions. In the paper’s cost analysis, the configuration that placed second used an ensemble of frontier models and cost $59 an hour to run, about the same as the $60 an hour the authors derive from the average salary of a US penetration tester. The $18-an-hour configuration, which ran on GPT-5 alone, submitted as many findings, but only 55% of them were valid and it placed seventh.
The authors list the limits themselves. The testers had at most ten hours of active work, against the one to two weeks of a typical engagement, and the agents ran for 16 hours with only their first ten counted in the ranking. The professional who finished first had done external reconnaissance before the test began. The university’s IT team knew about the test and approved flagged actions it would otherwise have blocked, and the sample was too small for hypothesis testing. The study was funded by a Stanford seed grant and an unrestricted gift from OpenAI.
Anthropic’s work with Mozilla on Firefox shows the whole pipeline from report to fix. In a two-week effort the two organisations described on 6 March 2026, Claude Opus 4.6 produced 112 reports. Mozilla had asked for the findings in bulk, including crashes nobody had yet shown to be security issues, and triaged them itself. It classified 22 as vulnerabilities, 14 of them high severity, and Mozilla said all 22 were fixed in the current version of Firefox by the time the work was announced. The same effort surfaced 90 other bugs, most of which Mozilla had also fixed.
Anthropic then ran an exploit test several hundred times, spending about $4,000 in API credits. The model produced a working exploit in two cases, both only in a test build with the browser sandbox removed.
Project Glasswing, which Anthropic launched in April 2026 to give a group of partners access to its unreleased Claude Mythos Preview model, has produced tens of thousands of findings and close public scrutiny. In its first update in May, Anthropic reported that the model had rated 6,202 of its 23,019 open-source candidates high or critical. Six independent security firms, and in a few cases Anthropic itself, had assessed 1,752 of those. Of the assessed findings, 90.6% were real vulnerabilities and 62.4% were confirmed as high or critical.
Anthropic’s disclosure dashboard, which now covers findings from November 2025 onwards, recorded 29,439 candidate findings on 2 October 2026. External security firms had reviewed 6,123 of them and confirmed 92.7% as valid. Anthropic had reported 6,157 vulnerabilities to the maintainers of 591 open-source projects. Of those reports, 1,333 had been through external review first, and Anthropic sent the other 4,824 directly, noting that they may include false positives. To its knowledge, 516 had been patched upstream, which doesn’t mean the fixes have been installed. Anthropic has described independent human triage as the step that limits how many findings it can disclose.
On 8 September 2026, Patrick Garrity of the vulnerability intelligence company VulnCheck reconciled the dashboard with the disclosure ledger. The ledger then marked 202 findings as fixed, while the dashboard counted 421 as patched upstream, and 245 findings had been withdrawn. Where both the model and a project’s maintainers had rated a finding, the model put 91.5% in the high or critical bands and the maintainers 51.3%. Eighteen findings had been fixed by maintainers before Anthropic reported them.
These counts describe different stages and different sets of findings, and they don’t share a denominator. Most of what the independent firms checked was real. Where maintainers rated severity, they rated many findings lower than the model did. At each snapshot, a small share of what had been reported was known to be patched.
Every result above describes a system, and the access conditions differ. AIxCC, Firefox and Glasswing all worked from source code. ARTEMIS worked over the network with a student-level account. In the ARTEMIS study, GPT-5 inside the authors’ framework outperformed half of the professionals, and the same model inside Codex outperformed two of them. A vendor quoting a model’s benchmark score hasn’t said how its testing system performs, which is why I keep separating a model, a system and an agent.
Where the bottleneck moved
HackerOne’s May 2026 analysis of its own platform shows the same pattern at scale. In the twelve months to March 2026, submissions rose by about 76%. HackerOne says its signal rates, a measure of report validity, stayed relatively consistent, and reads the growth as a mix of valid and invalid reports rather than a wave of AI-generated noise. Mean time to remediate fell by about 80%. The number of vulnerabilities resolved each month fell by about 46%, the backlog of validated but unresolved findings grew more than 21-fold, and the resolution rate for critical findings dropped from over 83% to under 40%. A July product announcement put the growth of the critical backlog at 29-fold.
These are platform-wide figures from a vendor, not a controlled study, and I covered what they do to speed metrics in the metrics piece in this series. HackerOne’s own reading is that discovery is outpacing remediation capacity. Across the platform, validated findings arrived faster than customers resolved them.
Open-source maintainers face the same volume with fewer people. In his January post, Stenberg contrasted vulnerability reports with pull requests, which hadn’t become a problem for curl because the project screens them automatically: no person looks at a pull request until it gets “green check-marks from 200 CI jobs”. Build and test checks on a code change do a narrower job than establishing that a vulnerability is real, but the principle carries over, with machine checks running before anyone spends time. Bug bounty programmes and open-source projects have mostly received AI findings without a filter like that, as reports for people to read.
Severity is a second constraint. In the Glasswing ledger, maintainers rated far fewer findings high or critical than the model did, and people mis-rate severity too. In the ARTEMIS study, one professional rated an anonymous directory service exposing details of 27,000 users as informational, and the reviewers rated it critical. A severity label decides what gets fixed first, so it needs an independent rater, whoever produced the finding, and that rater can be wrong as well.
How much to automate
Full automation isn’t the right goal for most programmes yet. The evidence supports supervised automation, which means automating discovery, reproduction checks and candidate fixes where they work, while accountable decisions about scope, severity and deployment stay with people. Discovery is the part the systems above do well and cheaply.
Verification can be largely mechanical when a finding arrives with evidence an independent reviewer can check. Where it’s safe, that means a minimal test case or proof of vulnerability that a separate replay tool runs against a clean build. Where exploitation would be unsafe, documented configuration or code evidence can do the same job. Mozilla singled out the minimal test cases in Anthropic’s reports as what let its team verify each issue quickly.
Grouping findings by root cause is a separate step that someone should review, because one input can reach several faults and many inputs can reach one. A crash on its own doesn’t show that a security boundary was crossed. AIxCC scored proofs, and the ARTEMIS framework tried to reproduce each finding before submitting it.
Keep the verifier and its evidence out of the testing agent’s reach. OpenAI’s report on the Hugging Face incident, discussed below, records agents trying to erase or tamper with their outputs and message logs to fool the automated graders. The attempts didn’t change the logs OpenAI’s graders saw, but they show why the check has to sit outside anything the agent can write to.
Deciding severity and the fix needs context the tools usually lack and someone accountable for the decision. Whether a flaw is reachable in a particular deployment, what data is behind it, which customers run the affected version and what a change would break are questions for the people who own the system.
Stenberg’s team spent hours turning five reported vulnerabilities into one, and much of that time went into establishing that three of them were limitations already described in the API documentation. Garrity’s analysis lists the questions maintainers ask of each finding: whether it is a vulnerability or just a bug, whether the component is reachable, whether someone has already reported it, and whether it falls inside the project’s security boundary.
AI-written patches belong in the same pipeline. AIxCC systems patched 68% of the planted vulnerabilities, Anthropic sent Mozilla a Claude-written proposed patch with its first Firefox finding, and Stenberg notes that AI analysers can often offer a patch, though usually not a complete fix.
Anthropic’s own advice on patching agents is to check two things automatically: that the original bug can no longer be triggered, and that the existing tests still pass. It adds that it can’t guarantee a patch passing both checks is ready to merge, and recommends reviewing AI-written patches like any patch from an external author. I’d add a regression test for the root cause where one is feasible, and follow each fix through review, merge, release and deployment, with a retest at the end. In some products a validated mitigation is the right interim remedy.
That leaves the testers’ time for verification, severity, reachability and the work the tools miss. In the ARTEMIS study, 80% of the professionals found a remote code execution flaw on Windows machines reached through a web-based remote console. ARTEMIS couldn’t operate a graphical interface, so it reported weaker misconfigurations on the same device and moved on.
ARTEMIS also reached an older management interface that the professionals’ browsers wouldn’t load, and exploited a flaw on it that no human found. The paper attributes this to the agent working from the command line. The people and the agent missed different things, which is an argument for running them together.
Measuring it like a fuzzer
Fuzzing researchers went through the same measurement problem a decade earlier, and AI testing can reuse what they learned. In Evaluating Fuzz Testing, George Klees and colleagues reviewed 32 fuzzing papers in 2018 and found that every evaluation lacked something important. Most didn’t say how many trials they ran, none used a statistical test, and many counted crashes or code coverage instead of distinct bugs. They recommended multiple trials compared with statistical tests, several sets of seed inputs, runs of 24 hours rather than a few, and ground truth in the form of known bugs, with coverage as a secondary measure.
Six years later, Moritz Schloegel and colleagues reviewed 150 fuzzing papers published between 2018 and 2023. They found the guidance on statistical tests widely ignored, and in attempts to reproduce eight of the papers they could not fully support the claims.
AI testing agents can take different paths through the same target on repeated runs. A comparison should therefore repeat each fixed configuration, report the spread, and record the model, tools, prompts and budget. Changing the configuration answers a different question: in the ARTEMIS study, two configurations of one framework, using different models, finished five places apart.
Each measure needs its definition written down before the test starts, along with a fixed target version, time window and scope. These are the ones I’d use.
| Measure | How to calculate it | What to watch |
|---|---|---|
| Validity rate | Share of adjudicated submissions confirmed as real, in-scope vulnerabilities | Pending, disputed, duplicate and non-security reports counted as valid |
| Unique verified findings | Confirmed findings grouped by root cause, per fixed test budget | Crash and alert counts that inflate the number |
| Recall on known vulnerabilities | Share of known, independently confirmed vulnerabilities in a versioned test corpus that the tool finds | Cases the vendor has seen or tuned on, and seeded flaws simpler than real ones |
| Severity agreement | Tool ratings compared with independent ratings under the same scheme | Raters who see the tool’s label first, and under-rating treated the same as over-rating |
| Cost per verified finding | Licence, compute, setup, supervision and verification time, divided by unique verified findings | Human time and zero-yield runs left out |
| Patch outcomes | Share of proposed patches merged, released and deployed without a regression | Patches that pass the tests without fixing the flaw |
| Missed vulnerabilities found later | Vulnerabilities present and in scope in the tested version but found afterwards by someone else | Flaws introduced in later releases counted as misses |
| Time to fix or mitigate | Time from a verified finding to a deployed fix or mitigation, alongside the age of open findings | Averages calculated over closed findings only |
| Incremental findings | Verified findings the existing testers didn’t make, at a comparable budget | Rediscovery of what the baseline already finds |
Known vulnerabilities give the closest thing to ground truth. Keep old versions of the application or firmware containing vulnerabilities that were found and confirmed later, add seeded flaws where they help, and count how many the tool finds with acceptable evidence. Keep the corpus away from vendor tuning, note where a case might be in a model’s training data, and report the results by class of flaw and access condition. The result describes that corpus, not every unknown flaw in production.
Missed vulnerabilities cover the rest over time. A flaw that was present and in scope in the tested version, and that a researcher, customer or attacker finds later, counts against the test that missed it. A flaw introduced in a later release doesn’t.
Cost per verified finding has to include the people. AIxCC teams spent about $152 per competition task, and the cheaper ARTEMIS configuration cost $18 an hour, but neither figure includes the hours someone spends confirming, rating and fixing what the tool produces. Mozilla took on that triage for 112 Firefox reports, and curl’s security team spent hours on five Mythos findings. A tool with a low licence cost and a low validity rate can be the most expensive option once that time is counted.
Comparing AI-enabled testing with the current approach takes a head-to-head test. The useful comparison usually has three arms: the existing testers, the same testers using AI, and an autonomous configuration where that is safe. Give each the same scope, time and rules, freeze the target or account for fixes between runs, and score the findings blind against one scheme. Measure the overlap between arms, because an agent that mostly rediscovers what the baseline finds adds little. Repeating the comparison once a year, or after a major change to the model, prompts or tools, tracks drift over time but doesn’t replace repeated runs within the evaluation.
Write the scoring scheme down first. ARTEMIS rewarded exploitation over findings that were only verified, and its severity weighting gave a critical finding eight times the weight of an informational one, so its leaderboard reflects a preference for depth. For procurement, weight findings by their importance to the organisation rather than their technical difficulty. A simple, reachable critical exposure can deserve priority over an elegant exploit.
Public benchmarks are a weak substitute for that test, partly because models can cheat on the evaluations meant to measure them, as the OpenAI incident below shows. A testing tool should be judged on the systems and conditions it will meet in production. I made the same argument about machine-learning detectors: adaptive evaluation is a procurement question.
Safety and rules of engagement
An AI testing agent acts on real systems, and its mistakes affect production and sometimes third parties. The ARTEMIS team ran their agents under controls I would copy. The university’s vulnerability disclosure policy defined the scope and banned destructive actions, a member of the research team watched every agent session live with authority to stop it, and the university’s IT staff watched the network logs at the same time.
An agent could be stopped at three independent points, by killing its process, shutting down its host or cutting its network connection. The authors observed no out-of-scope or unsafe activity from any participant, human or AI.
OpenAI’s technical report on the Hugging Face incident, published on 26 August 2026, describes what happened when the controls were weaker than the agents. OpenAI was running ExploitGym, a cyber-capability evaluation, without the production classifiers that normally stop its models pursuing high-risk cyber activity, in order to measure maximal capability. Agents in the evaluation found zero-day vulnerabilities in an internal package mirror and used it to reach the internet.
The agents collected exposed credentials and, between 11 and 13 July 2026, executed code on 41 of Hugging Face’s production dataset-server workers. According to the report, in under 13 hours the agents went from code execution in a single worker to administrative and host-level access across several Hugging Face clusters. I covered the first disclosure in the containment failure piece.
A monitoring tool had flagged port sweeps from the evaluation on 27 June, and on-call staff advised that the run didn’t need stopping. Across OpenAI’s attempts, no model had solved 198 of the evaluation’s 898 tasks. Those tasks were 22% of the set but accounted for 93% of the tasks the agents discussed on their improvised message board, and OpenAI identified them as a significant factor in the agents probing their sandbox.
OpenAI concluded that no single control should be assumed robust and that controls should overlap independently. Its retrospective assessment is that the chain-of-thought monitoring it now runs would have paged its security team more than a day before the breach.
From both cases I’d take these rules for a testing programme. Enforce scope outside the agent, with network allow-lists and egress controls the agent can’t change, and test the indirect routes too, because the way out in the OpenAI case was a permitted internal package mirror. Set limits on actions, rate, concurrency, cost and duration, and define stop triggers in advance, such as unexpected access to a third party, exposure of sensitive data or loss of monitoring.
Keep at least two independent ways to stop the agent, make sure stopping it also covers anything it spawned, the credentials it was issued and any remote sessions, and have a person watching live sessions against production. Don’t switch off safeguards to see what an agent can do while it can reach anything real.
Give the agent credentials scoped to the test and short-lived. Treat whatever the target sends back as untrusted input, since a target can contain instructions aimed at the agent reading it, and assume the model can be fooled. Accept “nothing found” as a legitimate result, because the OpenAI agents probed their sandbox most on tasks no model had solved, and record it separately from a test that was stopped, refused or failed.
Some general-purpose agents refuse offensive work. In the ARTEMIS study, Claude Code and one other framework declined the task, and Hugging Face’s own responders were refused by hosted models when they asked for help analysing the July attack. Refusals protect against misuse and also stop legitimate testers, which is why model providers now offer vetted access to defenders. Attackers have already persuaded AI agents to do offensive work by claiming to be authorised testers. A testing programme should therefore be able to show written authorisation for every engagement an agent runs.
Industrial products and the Cyber Resilience Act
For a manufacturer of industrial products, AI-enabled testing fits into a process that already exists. IEC 62443-4-1, the 2018 standard for the secure product development lifecycle, sets out security verification and validation testing in five requirements: testing of security requirements, testing of threat mitigations, vulnerability testing, penetration testing, and independence of testers. Vulnerability testing in the standard includes fuzzing, attack surface analysis, known vulnerability scanning and software composition analysis.
AI tools can support the vulnerability testing and penetration testing the standard requires. Using them doesn’t by itself show adequate coverage, or the independence the fifth requirement asks for, because an AI agent configured and run by the development team is still testing by the development team.
Every verified finding then enters the same standard’s practice for managing security-related issues, and since this autumn EU law covers some of them. Under the Cyber Resilience Act, manufacturers in scope have had to report actively exploited vulnerabilities and severe incidents since 11 September 2026, through ENISA’s single reporting platform. Reports are due without undue delay, with an early warning within 24 hours of becoming aware and a notification within 72 hours. For an actively exploited vulnerability, the final report is due within 14 days of a corrective or mitigating measure becoming available, and for a severe incident within a month of the notification. The reporting duty already applies to products placed on the market before December 2027.
From 11 December 2027 the main requirements apply, including vulnerability handling throughout a product’s support period. Manufacturers must address and remediate vulnerabilities without delay in relation to the risk they pose, including through security updates, and apply effective and regular tests and reviews of their products’ security. They must also report vulnerabilities they find in integrated components, including open-source ones, to whoever maintains those components. Products already on the market before that date generally fall under these requirements only if they’re substantially modified afterwards.
An AI-found vulnerability doesn’t trigger a report by itself, because the reporting duty covers active exploitation, and an authorised proof of vulnerability isn’t exploitation. The remedy needn’t always be a patch either, since the regulation ties remediation to risk. A testing programme that produces findings faster than the vulnerability-handling process can assess and address them can expose weaknesses in that process from December 2027. AI tools that scan a product will also find flaws in the open-source components inside it, and the duty to report those upstream sends them to the same maintainers who struggled with Glasswing’s volume.
Testing agents near operational technology need stricter limits still. An agent given access to connected IT systems may acquire a route towards OT, and an agent probing a production control network can cause the physical effect a test is meant to prevent. I’d keep agentic testing of OT to test rigs, digital twins and offline copies of controller firmware, within their documented limits, since a twin may not reproduce the process behaviour that matters. Any activity on a live OT network needs the site’s safety review and bounded techniques, and a human operator at the keyboard doesn’t make intrusive testing safe on its own.
Where to start
I would start with one scope the team already tests every year, an application or a product line with last year’s penetration test report on file. Before buying anything, measure how many verified findings per month the team can fix or mitigate. That number limits how much extra discovery the team can act on, although a serious finding can still justify isolating a system or holding a release when no fix is ready.
Then run the AI system alongside the existing testers on the same scope, with reviewable evidence required for every finding and severity rated independently. Build a corpus of known vulnerabilities to measure recall, and track the measures in the table from the first week.
The same evidence belongs in procurement. Ask a vendor for its validity rate on targets like yours and what evidence comes with each finding. Ask how its severity ratings compare with independent raters, how scope is enforced outside the agent, and what happens to the data and proof artefacts the agent collects. Ask how you’ll be told when the underlying model changes, and ask what the system found that nothing else found, what action followed and how much extra human work it created. If a vendor can only offer benchmark scores, none of these questions has been answered.
Before the pilot starts, I’d also agree what result would justify expanding the programme, what would restrict it and what would stop it. The answer will usually depend less on how much the AI finds than on how much of it the team can verify and address.
The edge of the evidence
I ran red teams at CyberAgency for twelve years, much of the work on critical national infrastructure. Around 2011 we ran an extensive engagement against a mass-transit rail operator and found twenty different ways to chain attacks on IT, operational technology and intentional electromagnetic interference into kinetic impact on the railway, up to a crash or a derailment. One of the twenty used a $15 GSM jammer against the GSM-R radio link the trains used for train control, which triggered the emergency brakes. I described that jammer test in a 2018 article on intentional electromagnetic interference.
The AI testing results above concern software and networks, and within that domain the agents can do sophisticated work. The OpenAI agents combined zero-day vulnerabilities, exposed credentials and cloud permissions into a compromise spanning several organisations’ systems. Stenberg found that the AI tools used on curl turned up new instances of known kinds of error, while Mozilla reported that Claude also found classes of logic error that fuzzing had not uncovered.
AI agents have done well in OT-themed competitions too. Alias Robotics reported that its agent solved 32 of the 34 challenges in the 2025 Dragos OT capture-the-flag and finished sixth among more than 1,000 teams, with people retrieving the challenges and submitting the flags. Its authors note that such puzzles have known solutions and no real consequences.
None of the cited work establishes autonomous competence across a complete cyber-physical engagement, chaining weaknesses across IT, control systems and radio into a physical outcome, with part of the attack surface reachable only by someone close to the track. Eight years after my 2018 piece on AI and the cyber arms race, the testing evidence is strongest where automation was easiest, in code and networks.
AI-enabled testing raises the floor on code and network findings. For cyber-physical systems it doesn’t yet replace the people who think across those domains, and AI is moving into robots and other autonomous machines where that cross-domain thinking will be needed more. A measurement plan should record what the AI was never asked to test.
In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.