AI Governance and Policy

Measuring AI in the SOC: Outcomes, Baselines and the MTTR Trap

In May 2026 HackerOne published a year of remediation data from its platform, and the speed figures looked excellent. Over the twelve months to March 2026, mean time to remediate fell by about 80% and the median by more than 70%. Over the same months, the number of vulnerabilities resolved each month fell by about 46%. The backlog of validated but unresolved findings grew more than 21-fold. The backlog of critical findings grew about 25-fold, as the resolution rate for criticals dropped from over 83% to under 40%. A July update put the growth in the critical backlog at 29-fold.

Submissions on the platform rose by about 76% that year, and HackerOne notes that the timing lines up with AI-assisted discovery tools reaching more researchers. Its own reading is that teams fixed individual issues faster, perhaps by choosing which ones to fix, while the total effort going into remediation shrank. “We can resolve faster. We just didn’t resolve more,” the authors wrote. These are aggregate figures across many customers and they don’t isolate the effect of AI, but they show a speed metric and an exposure indicator moving in opposite directions in the same dataset.

Anyone claiming an AI speed gain in a security function should show a credible baseline and a quality counter-metric; without them, faster processing is not evidence of better protection.

IBM’s 2026 Cost of a Data Breach report, based on Ponemon Institute research with 602 breached organisations, points the other way, and it will turn up in a great many vendor presentations this autumn. IBM found that organisations making extensive use of security AI and automation had breaches costing $1.93 million less than organisations using none, with breach lifecycles 65 days shorter.

That is an observational comparison. The extensive users probably differ from the non-users in budget, maturity, staffing and telemetry, and IBM grouped AI with automation in general. A sample made up only of breached organisations also can’t show how adoption changes the chance of being breached. The $1.93 million difference is useful background, but it can’t tell a CISO what a particular AI workflow will do in their own security function.

A CISO deciding on an AI workflow needs to know whether it delivers better protection at the same cost, the same protection at lower cost, or more capacity without an unacceptable loss of quality. All three are legitimate results. Telling them apart requires measuring what the workflow catches, what it misses and what it costs, against a comparison the organisation controls.

What has been measured so far

Microsoft has published some of the best-known controlled trials of AI in security work. Its Office of the Chief Economist reported a randomised trial of Copilot for Security with security novices in November 2023 and one with experienced analysts in March 2024, who were 22% faster and 7% more accurate on a set of tasks with Copilot, the 7% being a relative improvement in task scores. The trials were properly randomised, which gives them far more causal weight than customer testimonials. Microsoft ran them on tasks it chose, though, and an independent replication would show how far the results generalise.

Microsoft’s trial with IT administrators shows why time savings should be measured rather than asked about. Copilot saved participants 18.23 minutes on average, holding accuracy constant, yet 64.5% estimated the saving at more than 20 minutes and 41.9% at more than 30. The authors note that the overestimates are consistent with participants enjoying the tool.

METR found a wider gap in its trial of early-2025 AI tools. Sixteen experienced open-source developers working on 246 real issues took 19% longer when AI was allowed. They had expected to be 24% faster and afterwards believed they had been 20% faster. In a February 2026 update, METR said its newer data was unreliable and that developers were probably faster with the tools by then. The size of the effect has moved, but the gap between perceived and measured speed still transfers to a SOC.

Microsoft researchers James Bono, Justin Grana and Alec Xu studied Security Copilot in day-to-day use. They compared 89 organisations that adopted it with 88 matched non-adopters, drawn from telemetry covering more than 150 organisations, across 95,522 incidents. In the third month after adoption, mean time to resolve was 30.13% lower in the adopting group, with no significant difference in the first two months.

Anyone quoting that figure should quote its conditions with it. The main analysis is limited to incidents resolved within 12 hours. The clock starts when an analyst first opens an incident, not when the alert fires. The authors state plainly that unobserved confounders prevent causal identification and that adopters may be the organisations that benefit most.

Their follow-up in April 2025 added the number of alerts per incident and the probability that an incident is reopened. The reopening estimate in their preferred model, 68.44% lower in the third month, excludes organisations whose reopening rate was above 80%. Without that filter the estimate is 18.9% and not statistically significant, and the authors’ own discussion rounds the effect to “over a 50% drop”. I would copy the design of reporting a reopening rate next to a speed measure, even though the estimate is sensitive to how the data is filtered.

The study I found most useful is Microsoft’s randomised trial of its Phishing Triage Agent, published in late 2025. It recruited 167 external security analysts, split them into three groups and gave each a queue of 25 user-reported emails drawn from 93 real reports. The control group worked without AI. A second group worked through a queue the agent had sorted and saw its verdict and rationale for each email. A third got the sorted queue without the verdicts.

The agent was right about every real email in the corpus, so the researchers injected synthetic errors to set its effective precision and recall at 80%, a level Microsoft says matches its internal measurements on large samples. The participants reviewed every email in their queue. Microsoft then calculated how the auto-close workflow would perform by combining each analyst’s decisions on the emails the agent called malicious with its own verdicts on the emails it called benign.

On that basis, analysts with the agent found up to 6.5 times as many true positives per minute, or 3.1 times in the study’s most pessimistic scenario, and about four-fifths of the gain came from the queue handling. F1 rose by 77% when the agent’s verdicts matched the ground truth and by 48% at 80% accuracy. In the 80% scenario, mean recall was 0.70 against 0.79 in the control group. That difference wasn’t statistically significant, which means the trial couldn’t distinguish a nine-point fall in recall from no change at all.

Analysts didn’t rubber-stamp the agent’s malicious verdicts. They spent 53% more time on malicious emails. When the agent wrongly called a malicious email harmless, though, analysts who could see that verdict missed it 46% of the time, against 17% in the control group. They also confirmed 86% of the agent’s correct benign verdicts, against 65% for the controls. Microsoft concluded that analysts gain little from reviewing benign verdicts and that the product should resolve them automatically, which it does. Auto-closure removes routine human review of exactly the errors the analysts were worst at catching: malicious emails the agent called harmless. Most of them will go uncounted unless someone samples the auto-closed emails.

Every participant also received the same prepared evidence for each email, including annotated screenshots, the results of detonating its links and sanitised attachments. That changes what the trial timed. In a real SOC somebody or something has to gather that evidence, and the trial can’t say whether doing so would widen or narrow the gap.

The agent Microsoft tested has since been renamed the Security Alert Triage Agent. Its documentation describes it as the same agent, generally available for user-reported email and being extended to other alert types. Nobody should read the phishing results as evidence for those other alert types.

Microsoft’s Project Ire, an autonomous malware-classification agent announced in August 2025, reached precision of 0.98 and recall of 0.83 on a public set of Windows drivers. On nearly 4,000 harder files that Microsoft’s automated tools couldn’t classify, all created after the models’ training cut-off, precision was 0.89 and recall 0.26, so the agent found about a quarter of the malware. A longitudinal study by Ronal Singh and colleagues of 3,090 GPT-4 queries from 45 eSentire analysts shows them using a general-purpose model mostly to interpret artefacts and draft text, and it measured use rather than outcomes.

We now have evidence of faster work and better scores on selected tasks, and some improvement in operational measures such as reopening. These studies don’t establish how much a particular AI workflow reduces losses in a particular organisation. External benchmarks, some of which models have learned to game, can’t replace measurement in an organisation’s own environment.

I keep separating a model, a system and an agent because a result measured on one doesn’t automatically apply to the others. A benchmark score describes a model. Copilot running on an organisation’s own telemetry is a system. A triage agent that closes alerts without a human is an agent, and its errors have to be measured where it acts.

The MTTR trap

In a conventional closed-case calculation, mean time to resolve includes only the cases that were resolved. If AI closes the easy alerts in seconds, the average falls even while the hard cases age in the queue, because an open case doesn’t enter the calculation until someone closes it. HackerOne’s data shows the effect at platform scale. The Bono, Grana and Xu study shows a milder version by design, because the authors left out the long tail by limiting the main analysis to incidents resolved within 12 hours.

The authors explain why they chose the limit, and it is a reasonable choice for their question. A SOC dashboard that reports only closed cases, or caps durations, has the same blind spot without the explanation.

Before comparing any timing metric across tools or periods, the team has to define it. That means the event that starts the clock and the one that stops it, elapsed or business hours, which population counts, how reopened, merged and duplicate cases are handled, and the reporting window. Alerts, incidents, findings and verified mitigations are different populations, and triage, containment, resolution and remediation are different stop events. Automation can shorten an analyst’s hands-on time without shortening the wall-clock delay that an attacker exploits.

Reporting a mean on its own hides the shape of the distribution. I’d report the median, the 90th percentile and the share of cases past a severity-specific deadline, together with the age of the oldest open high-severity cases. A mature programme can go further and use survival analysis, which counts still-open cases as censored observations instead of dropping them.

“Alerts handled automatically” counts benign and malicious closures together, so an AI that closes true positives as harmless produces the same number as one that closes only noise. If an AI writing triage notes assigns severity differently from a human analyst, every severity-weighted metric shifts even if nothing about the incidents has changed. Seat utilisation, prompts per analyst and the share of incidents with an AI summary are useful for running the programme. They are closest to what NIST calls implementation measures, which show that something is in place rather than that it works.

Microsoft’s triage agent comes with a performance dashboard showing incidents handled, mean time to triage and a breakdown of true and false positives over time. Those labels are the agent’s own verdicts. According to Microsoft’s documentation, the agent classifies an alert it judges benign as a “False Positive” and resolves it. It marks an alert it judges malicious “True Positive” and leaves the incident open for an analyst. Counting those labels shows what the agent decided, not whether it was right.

Analysts can give feedback on individual alerts, and later hunts or external notifications can expose mistakes. Without deliberate sampling, though, nobody sees most of the errors among auto-closed alerts.

NIST’s SP 800-55 Volume 1, the measurement guide it revised in December 2024, recommends where possible that the person collecting measurement data be someone other than the person who owns it, to avoid a conflict of interest, while accepting that small organisations may not manage that. The guide also expects automated collection from systems and logs, so a product’s own telemetry is legitimate data. The real issue is who defines the truth labels, checks them and challenges the result.

The same guide sorts security measures into four types: implementation, effectiveness, efficiency and impact. NIST uses mean time to remediate to illustrate patching efficiency and defines effectiveness as how well controls are working and whether they meet the desired outcome. It puts costs and business consequences under impact. The published AI studies report mostly efficiency, plus task accuracy, while the CISO’s question is about effectiveness and impact.

NIST also warns that a single timing measure is unlikely to explain why remediation is delayed, and that a badly designed phishing test can raise the score while telling the organisation less about its staff. Both warnings apply to AI tools as written.

Passive production metrics also don’t show how a workflow will hold up against attackers who adapt to it. Machine-learning detectors have a long record of evasion, one of several attack classes against machine learning models. AI tools in a SOC also read attacker-controlled emails, logs, tickets and binaries all day, so the working assumption should be that the model can be fooled. For an agent that can act, the test set should include emails or tickets written to change its verdict, hide evidence or trigger an unauthorised action. I’d score each test on whether the whole system contains the consequence.

Five layers of measurement

I’d measure an AI workflow at five layers, each tied to one question. Is the tool usable for the work? Did the work get faster? Was the AI’s output right? Does the function catch more and miss less? Did exposure and cost go down? The layers overlap, and a programme that reports only the first two hasn’t yet answered the protection question.

At the first layer, I’d record the workflows where AI is in use, the share of alerts it touches, the log sources it can actually read, and the tasks it fails or refuses, with the fallback and the time lost. Some refusals are appropriate and shouldn’t count as failures, but others block legitimate work. Hosted-model guardrails obstructed forensic log analysis for the Hugging Face responders this summer, and a count of legitimate tasks blocked is one of the few usage figures with a security consequence.

Efficiency measures are the ones vendors already report: time to a defined event, time per alert type and throughput. Of these I’d keep verified, deduplicated true positives per analyst-hour, because it counts output with value, and Microsoft’s phishing trial used the same idea per minute. I’d also report hands-on analyst time and wall-clock time separately.

Precision among the AI’s positive findings and accuracy across a sample of all its verdicts are different numbers, and an approval rate dominated by easy benign cases flatters both. The reviewers who adjudicate shouldn’t see the AI’s verdict, and seniority alone isn’t ground truth. The reopening rate should use a fixed follow-up window and link replacement tickets to the original incident. It measures observed rework, not proof that the other closures were right. Override and escalation rates are diagnostic signals, and someone needs to adjudicate a sample before reading them as good or bad.

Acceptance of seeded wrong verdicts, the design Microsoft used, counts how often people accept a deliberately incorrect AI recommendation. It measures over-reliance, not the AI’s accuracy, and should be reported separately for missed attacks and false alarms. In the phishing trial, analysts treated the two very differently. Analysts can only check what they can inspect, so the evidence, tool output and event history behind each verdict should be kept. A fluent rationale from the model isn’t evidence, a distinction the explainability literature has been making for years.

I found few public evaluations at the fourth layer, security effectiveness. Recall can be estimated on known malicious cases, such as seeded attacks, replayed historical attacks or independently adjudicated incident cohorts. Malicious cases taken from the detector’s own alerts would inherit its blind spots, and the result describes only the scenarios tested. Weighting it by asset criticality and privilege ties it to risk. Verified technique coverage should count the tested procedures behind each MITRE ATT&CK technique, with the environment and the test date, since one fired detection doesn’t validate every way of implementing a technique.

The fourth layer also includes the escape rate, which needs defining before anyone reports it, for example as the share of incidents discovered later by a hunt, a notification or a regulator that had an earlier alert the AI closed. It includes the error rate among automated benign closures, defined carefully below, and dwell time on confirmed incidents, read with care because better discovery of old compromises can make it look worse for a while.

At the fifth layer, the Cyentia Institute’s research with Kenna Security offers a useful pair for vulnerability work. Coverage is the share of exploited vulnerabilities that get fixed; efficiency is the share of fixes that address an exploited vulnerability. In their historical comparison of prioritisation strategies, fixing everything with a CVSS score of 7 or above reached 31% efficiency and 53% coverage, against 61% and 62% for a prediction model that needed about half the effort. AI prioritisation can be judged on the same pair.

Ticket closure is a weak proxy for remediation, so I’d count verified mitigations and record risk acceptances, compensating controls and decommissioned assets separately. A resolution rate also needs a definition, because closures divided by new arrivals behave differently from the share of a cohort resolved within a fixed time.

Every layer should be broken down by alert type and severity. An average across categories can hide much worse performance on rare, high-impact categories, in the same way that aggregate accuracy can hide far higher error rates for particular groups. Strong phishing performance shouldn’t offset a weak result on cloud identity alerts, and a subgroup too small to measure should be reported as insufficient evidence.

The pairs I’d report together look like this.

Reported numberReport it withWhat the pair tests
Time to a defined resolution eventReopening within a fixed window, open-case age distribution and high-severity deadline breachesWhether the speed holds across the whole queue
Alerts closed automatically as benignError rate in an independent sample of those closures, with its upper boundWhether automation is hiding malicious cases
Verified true positives per analyst-hourRecall on tested scenarios, including critical subgroupsWhether throughput cost detection
Findings raised by AIIndependently confirmed precision and severity agreementWhether the extra output is signal
Tool use by workflowTasks failed or refused, fallback success and adjudicated overridesWhere the assistance breaks down
Mitigations verifiedOpen exposure by severity and age, with arrivals and departuresWhether exposure is going down
Time or money savedTotal workflow cost, rework, sampling effort and harm from mistaken actionsWhether the saving survives full accounting

Each metric needs a short record: the population, numerator and denominator, the observation window, the data source, how ground truth is established, the uncertainty and the owner. It also needs the decision it is meant to support, written down before anyone builds the dashboard.

Testing an AI workflow

A team evaluating an AI workflow needs a credible comparison, and a before-period is the weakest kind. Attack volume and mix change from one month to the next, as do staffing and workload, so a simple before-and-after comparison can credit the AI with a quiet quarter. Historical data is still useful for volume and variability. A concurrent comparison is better at separating the effect of the AI from everything else that changed.

Where it is safe, randomise at the level of the case, or of a cluster such as an incident, campaign, account or asset. Randomising individual alerts lets information leak between arms when alerts from one campaign are assigned to both. Stratify by alert category, severity and time, keep the existing protections in both arms, and set escalation rules in advance. Be explicit about what is being compared. Microsoft’s phishing trial compared a sorted, filtered AI queue with an unsorted manual one, which measures the package rather than the model inside it.

When a rollout has to be phased anyway, a stepped-wedge design, described by Karla Hemming and colleagues in the BMJ in 2015, is an option. Units switch to the new process in a random order, one wave at a time, and the units still waiting act as concurrent controls. It needs enough reasonably separate units, randomised timing, and an analysis that accounts for calendar time, learning and correlation within units, and the reporting standard for stepped-wedge trials is a good checklist for that.

A newly centralised security function may have several sub-teams to stagger, but it will also be changing staffing and tooling at the same time. I’d try a parallel design first if one is practical.

A seeded attack, run by a purple team or a breach-and-attack simulation platform, tests the detection and response path and gives a known denominator for recall on that scenario. A seeded wrong verdict tests whether people catch misleading assistance. Don’t insert either casually into production queues. Run them in authorised exercises, bounded test accounts or replay queues, with controls that stop a test verdict from closing a genuine incident or contaminating production labels. Rotate the scenarios so the team doesn’t learn to recognise its own tests.

The approach that came out of the emerging-technology risk labs I ran was to assess risk by testing rather than by speculation, and AI in a SOC deserves the same treatment.

Shadow mode lets an AI produce verdicts that nobody acts on while analysts work as usual. Reviewing only the disagreements misses the cases where both are wrong, so a random sample of agreements needs adjudicating too, against independent evidence, with an “unresolved” category for cases the evidence can’t settle. Shadow mode measures decision quality. It can’t measure the productivity or behavioural effects of acting on the advice. Read-only permissions are a sensible default, together with bounded data access and egress and no route to production actuators. I’d run evaluations under the same containment as production, a lesson from the Hugging Face containment failure.

Sampling automated closures

With auto-closure switched on, sampling the alerts the AI closed as benign becomes the most important measurement, and it is easy to give the result the wrong name. Three quantities get confused. Recall is the share of all malicious cases the process caught; the false-negative rate is the share it missed. The false-omission rate is the share of cases labelled benign that were in fact malicious. A sample drawn only from benign closures estimates the third. Its denominator is the benign closures, not all the malicious cases.

Picture a queue of 10,000 user reports, 50 of them malicious, and a useless classifier that closes everything as benign. Its false-omission rate is 0.5%, which looks excellent, while its recall is zero. A benign-closure target such as “below 1%” can’t stand on its own for that reason. Recall needs its own estimate from known malicious cases, and neither measure sees attacks that never raised an alert or phish that nobody reported.

Each week, take a representative random sample of automated benign closures and have a senior analyst investigate each one without seeing the AI’s verdict or rationale. Report the share found to be malicious, its upper confidence bound, and the estimated number of harmful closures in the period. With zero misses in 40 reviews, the one-sided 95% upper bound is about 7.2%, close to the 7.5% given by the rule of three. Getting that bound below 1% takes 299 clean reviews, or 368 under the two-sided convention.

Those numbers assume independent, representative samples and reliable adjudication. Duplicate emails from one campaign and alerts belonging to one incident aren’t independent, reviewers make mistakes, and the workflow changes while it is being measured. The sampling period and the decision rule should be fixed in advance, because checking accumulating results and stopping at the first favourable bound weakens the guarantee.

A rare, high-impact class contributes so few closures to an overall sample that the class’s error rate can be far higher than the overall bound. Such a class needs its own sample, weighted back to the production mix. I couldn’t find a published evaluation of benign-closure sampling or shadow mode in a security setting, so both need a documented protocol before their results go to a board.

A small permanent holdout, a share of alerts left on the non-AI process, keeps a comparison available after rollout where assignment stays feasible and safe. A few per cent may be too small for rare threats, and a holdout can be contaminated by shared knowledge or system changes. I’d review the holdout’s size and value periodically.

Where a holdout isn’t practical, frozen replay sets, fresh challenge cases and production sampling can stand in, each with its own limits. If the AI demonstrably prevents serious harm, withholding it from a random sample indefinitely is hard to justify, and a staged rollout or shadow mode is the better design.

Restarting the baseline every time a vendor ships an update erases regressions, so version boundaries should be marked and earlier results kept. Revalidation should follow changes to prompts, retrieval sources, rules, integrations, permissions and agent memory, as well as the model. Microsoft’s triage agent lets analysts turn feedback on an alert into a lesson stored in the agent’s memory to influence its future decisions, so its behaviour can change with no model release at all. Feedback that shapes future verdicts is also an input that a hurried analyst, or an attacker with access, can poison.

Governance: compliance, works councils and ownership

AI can cut the cost of preparing compliance evidence without changing how well any control works, and that is a legitimate efficiency gain. A rising pass rate is the wrong test of it, because more rigorous testing can lower the pass rate while improving assurance. Document review is still a valid way to assess some controls.

Article 21(2)(f) of NIS2 requires in-scope entities to have policies and procedures to assess the effectiveness of their cybersecurity risk-management measures. To claim stronger assurance, an organisation should be able to show that its assessment is accurate; to claim stronger protection, it should be able to show that its controls work. I’d report evidence accuracy, assessment coverage, control effectiveness and the remediation of deficiencies as separate things.

The PCI Security Standards Council’s guidelines on AI in PCI assessments, version 1.0 from March 2025, describe AI as “a tool, not an assessor” and keep human assessors responsible for findings and final decisions. They also warn that AI can introduce false positives, incorrect assumptions and bias. An internal compliance team can borrow the principle. A rate of AI-drafted material approved by a named person shows that review happened. I’d pair it with independent error checks and, where the stakes justify it, re-performance of the work.

In German establishments with a works council, section 87(1) no. 6 of the Works Constitution Act gives the council co-determination rights over technical systems designed to monitor employees’ behaviour or performance. The Federal Labour Court has read “designed to” as “objectively suitable for”, whatever the employer intends, since 1975, and restated the principle in July 2024. That case concerned retail headsets that weren’t assigned to individual employees and recorded nothing, and they still counted because supervisors could listen in. An AI assistant that records analysts’ prompts and alert-closure times can fall within the provision.

Scope, permitted uses, access, retention and reporting should be agreed before the relevant monitoring starts. Team-level reporting helps, but it isn’t an exemption, and randomising cases instead of people suits the experiment without settling the legal question. Outside Germany, the employee representatives and the privacy function should be involved according to local rules, since not every country has a direct equivalent.

The SOC should own its operational measurements and the corrective action that follows. A second-line risk function should challenge the definitions, sampling and interpretation, and internal audit should assess independently whether the programme is reliable, which it can’t do if it runs the programme. The team that bought the tool shouldn’t be the only judge of whether it works. Depending on its mandate, an AI security lead can coordinate all of this, though reporting outside the SOC isn’t independence by itself. Whoever approves the renewal should see both the operational results and the independent challenge.

Targets for the same spend

Microsoft’s phishing trial asked, as one of its research questions, whether recall stayed within a narrow band of the manual baseline, with three percentage points as the example, while productivity rose. That was a question rather than a finding, but it has the right shape for a target. A useful target names an efficiency gain and a quality floor, both against a baseline. It should be analysed as a non-inferiority question, comparing the confidence bound on the difference in recall with the margin, since a non-significant difference isn’t enough. If the manual baseline is itself poor, an absolute minimum belongs alongside the margin.

As an illustration only: raise verified true positives per analyst-hour in phishing triage by half, keep recall on tested phishing scenarios within three points of the manual baseline at 95% confidence, and keep the error rate among benign closures below a bound set from the prevalence and the cost of a missed phish. The right numbers depend on severity, prevalence and consequences, and serious failure classes should keep their own targets.

Cost per verified true positive is useful for triage economics, but it isn’t risk reduction. More attacks produce more true positives with no improvement in protection, repeated alerts for one attack inflate the count, and a good preventive control lowers it. A sound comparison counts deduplicated, independently validated incidents for a stable workload and case mix. It includes the full cost of licences, inference, integration, enrichment, analyst review, rework, sampling and governance, plus the harm from mistaken automated actions such as wrongful host isolation or account suspension.

Capacity released isn’t cash saved unless budgets or other costs actually change. In Microsoft’s trial, analysts spent 53% more time on malicious emails, a reallocation anyone can count. In a SOC, the equivalent is the hours moved into threat hunting or detection engineering and what those hours produced, such as new detections validated by test.

Where to start

I would start with one high-volume, well-understood workflow, and user-reported phishing is the obvious candidate. Four to six weeks of measurement before switching anything on is a reasonable planning figure, though the duration should follow from how often the relevant outcomes occur.

Measure volume, time to a defined triage event, reopening, recall on tested phishing scenarios in a replay queue, and the error rate in a blind sample of closed reports. Agree the measurement plan with the works council or other employee representatives, and with the privacy function, at the same time. Then run a concurrent comparison or a staged rollout. Sample automated closures every week from the first day, and report the paired metrics at team level.

The same evidence belongs in procurement. Before buying an AI triage or investigation product, ask the vendor for results against seeded errors and for benign-closure sampling with its confidence bounds, under conditions that resemble the buyer’s environment. Put notice of model changes, and the vendor’s regression testing for them, into the contract. Speed figures and satisfaction scores alone don’t supply evidence of effectiveness. I made the same argument about machine-learning detectors: adaptive evaluation is a procurement question.

Before the pilot starts, I’d also agree what result would justify expanding the workflow, what would restrict it and what would switch it off. A measurement plan that can’t trigger the third decision isn’t yet a basis for scaling the tool.

222fb9d292e3d0111656a33900e24a27cfb6a36eb7b202a94a66bb84766154b4?s=120&d=mp&r=g
[email protected] | About me |  Other articles

In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.

Related Articles