AI Literacy for Security Teams Is Measured by the Errors They Catch
On 27 July 2026, the AI Act’s literacy duty changed. The original Article 4 of the AI Act required providers and deployers to take measures to ensure, to their best extent, a sufficient level of AI literacy among the people who use AI systems on their behalf. As rewritten by the Digital Omnibus, Article 4 requires them to take measures that support the development of AI literacy, and says expressly that no specific level has to be guaranteed for any individual.
The Commission’s AI literacy Q&A already said that Article 4 doesn’t oblige anyone to measure employees’ AI knowledge. The measures still have to fit the people and the systems involved, but a security team can document a literacy programme without ever learning whether its analysts catch a convincing AI mistake.
Developers in METR’s 2025 trial believed AI had made them faster when it had made them slower, and participants in a Stanford study wrote less secure code with an AI assistant while believing it more secure. When Microsoft tested its phishing triage agent, analysts who saw its verdicts missed many more of the malicious emails it had been made to call benign than analysts working without it.
Measure AI literacy the way you measure phishing resilience: seed the errors and count who catches them.
AI literacy also covers a system’s limits, data handling and when to escalate, but catching errors is the part a security team can test with known answers.
Confidence is not competence
METR’s randomised trial, published in July 2025, gave 16 experienced open-source developers 246 real tasks on codebases they knew well. Allowed to use AI tools, they took 19% longer. Before starting they had expected AI to make them 24% faster, and afterwards they estimated that it had made them 20% faster. METR’s February 2026 update treats that result as a measurement of early-2025 tools, and says its newer experiment couldn’t produce a reliable estimate, partly because 30% to 50% of developers chose not to submit some tasks they didn’t want to do without AI.
For training, the useful finding is the gap between what was measured and what the developers believed. Their own sense of how well they worked with AI was a poor guide to their speed.
Neil Perry and colleagues at Stanford found a similar gap on security. In their study, published at CCS in 2023, participants given an AI coding assistant built on an OpenAI Codex model wrote significantly less secure code than those without one. They were also more likely to believe their code was secure. Participants who trusted the AI less and worked harder on their prompts produced more secure code, an association rather than proof of cause. The tools are old, and the study doesn’t show that current assistants have the same effect.
The phishing trial, which Microsoft ran and wrote up in a preprint posted in November 2025, measured analysts’ reliance in a SOC task. Analysts who saw the agent’s verdicts missed 46% of the malicious emails it had been made to call benign, against 17% for analysts working without it. The same paper reports large overall gains, with up to 6.5 times as many true positives per analyst minute and verdicts 77% more accurate. It also found that analysts didn’t rubber-stamp the agent’s malicious verdicts. Its author concludes that analysts gain little from re-reviewing the agent’s benign verdicts, because they mostly agree with it anyway.
I discussed what that means for an analyst’s sign-off in the piece on delegating SOC work to AI agents. For training, it means an analyst can be fast, confident and wrong, and tests with independently adjudicated cases can expose it.
A 2024 systematic review by Tomáš Lintner found 13 self-report scales and three performance-based ones among the 16 AI literacy instruments it assessed. Self-report instruments such as MAILS and SNAIL ask people to rate their own understanding and abilities. They can help assess perceived understanding and evaluate a course, but they don’t show whether an analyst catches a wrong verdict at work.
What literacy means for a security team
For a security team, AI literacy is the ability to work with an AI system and tell when it is wrong, in the team’s own tools and on its own data. That includes knowing what the system can’t see, what hostile content does to it, and when to escalate rather than accept or override.
Analysts also have to accept the AI’s advice when it is right. Max Schemmer and colleagues proposed measuring appropriate reliance on the cases where a person’s first answer and the AI’s advice disagree. One measure counts how often people keep a correct answer against wrong advice. The other counts how often they switch to correct advice when their own answer was wrong. An analyst who overrides everything has stopped relying on the AI, and the team loses the time the AI was saving. Literacy in this sense is calibrated reliance, and a test has to measure both directions.
Their design records a first judgement before the advice and a final decision after it. That shows when someone changes their answer, and whether they gave up a correct first judgement for wrong advice or replaced a mistaken one with correct advice. It doesn’t show how carefully they checked. Recording an analyst’s initial assessment and confidence before the AI’s verdict appears is a cheap change to an exercise, and it matches the evidence-first review I suggested for delegated decisions. In production, a full independent investigation of every routine alert would cost most of what the automation saves, so I’d keep evidence-first review for exercises and for high-consequence decisions.
Seeding the errors
Phishing simulations test how people respond to realistic lures. A literacy exercise does the same with realistic AI output that contains known errors.
In Microsoft’s trial, each analyst was given a queue of 25 emails, four of them malicious and 19 benign. For the other two, the agent’s verdict had been flipped, from malicious to benign on one email and from benign to malicious on the other. To make the wrong verdicts believable, Microsoft replaced the agent’s justification with a generic one that fitted the email, and checked that nothing in it obviously contradicted the content. The agent had classified every email in that corpus correctly, so the only wrong verdicts analysts saw were the ones Microsoft had planted.
Analysts in the agent group were also shown how accurate its verdicts usually were, and their queue was ordered by the agent, so the comparison with the control group doesn’t isolate the effect of seeing a verdict.
A SOC can run the same exercise on its own past cases, in its own tools. Independent reviewers should adjudicate each case first, record the decisive evidence, and present only what was available at the moment of decision, because past incident labels aren’t automatically reliable and hindsight makes cases easier. Run the seeds in a controlled exercise or a shadow queue that can’t contaminate production incident records or trigger real notifications, remediation or external submissions. Keep a separate labelled record of the exercise’s case versions, scores, debriefs and resulting changes, and make sure a planted verdict never becomes a lesson in the production agent’s memory.
Edited AI outputs test the analysts’ review. Hostile inputs run through an isolated copy of the agent test the whole system, and there the team has to establish whether the injected content actually changed the agent’s output before scoring anyone on it. If the production agent closes benign cases before analysts see them, an exercise that shows analysts every case is testing a different workflow. So the exercise should also include a controlled sample of cases the agent would have closed on its own.
Each exercise should produce a small set of measures, each with its denominator.
| Measure | How to calculate it | What to watch |
|---|---|---|
| False-benign corrections | Seeded malicious cases shown with a benign AI verdict that analysts correctly reclassified, divided by all such cases | Unresolved and escalated cases reported separately, not dropped |
| False-malicious corrections | Seeded benign cases shown with a malicious AI verdict that analysts correctly reclassified, divided by all such cases | Blanket rejection of the AI counted as skill |
| Harmful overrides | Correct AI verdicts changed into incorrect final verdicts, divided by correct AI verdicts reviewed | Escalation counted as an error |
| Appropriate escalation | Cases meeting a pre-agreed escalation rubric that were escalated, with unnecessary escalations reported too | Escalation used to avoid decisions |
| Unaided performance | Accuracy and time on comparable cases with the AI switched off, tracked across rounds | Changes in case difficulty or team membership |
| Workflow benefit | Assisted against unaided accuracy, time and escalation load on matched cases | A larger benefit read as lost skill |
Report counts alongside percentages, because nine catches out of ten is a different result from nine out of fifty, and don’t rank individuals on small samples. A team’s score also reflects its interface, the evidence available, time pressure and the cases chosen, which is why the results should lead to changes in tooling as well as in training.
Missed attacks can have severe consequences, and in Microsoft’s trial the false-benign cases showed the clearest deterioration against the control group. In absolute terms, analysts with the agent missed even more of its planted false alarms, 54% against 48% for the control group, a difference that wasn’t statistically significant. Count harmful overrides as well: a team that catches every seeded error by overriding half of the AI’s correct verdicts has traded one failure for another.
Some lures are far harder to spot than others. In a field trial of phishing training that sent ten simulated campaigns to more than 19,500 employees at UC San Diego Health, the share who fell for a lure was 1.82% for a fake request to update an Outlook password and 30.8% for a fake update to the vacation policy.
A literacy exercise built only on generic justifications may be easier than what attackers now write. The seeds should include persuasive errors too: a wrong verdict with a confident, specific rationale, a summary that leaves out the decisive log line, and a sample whose embedded text argues for a benign verdict. How hard each seed is should be tested on people, not assumed from how convincing it sounds.
To compare results over time, keep two sets of cases. A stable assessment set, refreshed with equivalent cases, measures the trend. A rotating challenge set covers new failure types and adversarial cases, and is reported separately. Analysts who see the same kinds of planted error every quarter learn to spot the exercise rather than the error, the human version of a model learning to game its evaluation. In a queue packed with planted mistakes, analysts can also learn that rejecting the AI is usually rewarded. Keep the share of seeded items low, and don’t read the results as production error rates.
What phishing training got wrong
The UC San Diego study, by Grant Ho and colleagues and presented at IEEE Security and Privacy in 2025, also tested the training itself. Recently completing the annual awareness course made no significant difference to whether employees fell for later simulations, an association the authors measured across the workforce. Randomly assigned training shown after a click reduced later failures by only about 2%. The authors concluded that such programmes, in their commonly deployed forms, are “unlikely to offer significant practical value” in reducing phishing risk.
According to the university’s summary, three-quarters of users spent a minute or less on the post-click training, and the authors recommended putting more effort into two-factor authentication and password managers that work only on the correct domain.
Measuring analysts doesn’t train them, any more than counting clicks trained UC San Diego’s staff. A seeded-error exercise is worth running for what the team does afterwards: a debrief on each missed error, changes to how the tool presents evidence, and a narrower set of tasks for the AI where analysts keep missing its mistakes. That is closer to the UC San Diego authors’ advice, which was to put more of the effort into the systems. A course and a completion quiz leave the question of transfer to real work unanswered.
To show that the training itself helped, the team should retest with fresh cases and again some weeks later, and a staggered rollout across sub-teams would give it a comparison. Running the exercise isn’t evidence that it improved anything. Analysts should know that exercises happen and why. The exercise still measures reliance as long as they can’t tell which items are seeded. Use the results for coaching and tooling, not for a league table.
Measuring people at work
A seeded-error programme measures how people perform, and some jurisdictions have rules about that. In Germany, the works council has a say over technical systems that can monitor employees’ behaviour or performance, under Section 87(1) No. 6 of the Works Constitution Act. The Federal Labour Court looks at what a system can objectively do, not what the employer intends. I covered the case law in the piece on measuring AI in the SOC.
Where a works council has those rights, the scope and use of any system that records attributable exercise results should be agreed before the exercise starts. Team-level reports and private coaching may reduce concerns, but the collection, access and retention of individual results still need to be settled. In a small team, even team-level results can identify individuals.
An AI system intended to evaluate individual analysts from the exercises may fall into the AI Act’s employment category in Annex III, point 4, especially where its scores feed decisions about work allocation or promotion. Check it against that category and the classification rules in Article 6. A spreadsheet of team results graded by people is not an AI evaluator, but exporting an AI-generated score into one doesn’t remove the system that produced it. For systems that are high-risk, which most SOC tools aren’t, Article 14 goes further and requires that the people overseeing them be able to stay aware of automation bias.
Skill without the AI
Literacy also covers what analysts can still do when the AI is unavailable, wrong in a new way, or under attack. Lisanne Bainbridge’s “Ironies of Automation”, published in 1983 about industrial process control, set out the problem. Automation takes over the routine work and leaves the operator to handle the failures. The operator’s skills then fade for want of the routine practice that built them. Bainbridge’s remedy was regular practice, including in simulators, on the tasks the automation normally does.
A 2024 paper by Auste Simkute and colleagues, including researchers at Microsoft Research, analysed the same ironies in generative AI. They argue that users shift from producing work to evaluating it, and that the automation tends to simplify easy tasks while making hard ones harder.
In an observational study at four Polish endoscopy centres, published in The Lancet Gastroenterology & Hepatology in 2025, the share of colonoscopies done without AI that found at least one adenoma was 28.4% in the three months before AI assistance was introduced and 22.4% in the three months after. The result is consistent with a loss of unaided performance, but it can’t establish the cause or whether the loss lasts, and colonoscopy isn’t alert triage.
In a SOC, juniors may be particularly exposed if an agent takes over the routine triage through which they would otherwise develop judgement. I’d run part of each exercise with the AI switched off and track unaided performance on comparable cases over time. A widening gap between assisted and unaided scores can simply mean the AI got better. A sustained fall in unaided scores is the signal to investigate.
What the training should cover
Article 4 leaves the content of the measures to the organisation, and for a SOC I’d make them task-specific. I’d build the training around five things an analyst needs to know about the AI they work with.
The first is what the system can and can’t see. Microsoft’s phishing agent, for example, is given read access to the content of the emails associated with alerts, not to every email. Every analyst should know the data sources, permissions and time window behind the verdicts they read.
The second is how hostile content works on the model. Analysts should learn the common patterns, such as injected instructions, text designed to trigger a refusal and plausible cover stories, and treat ordinary-looking content as untrusted too, because recognising patterns won’t catch every injection.
The third is how to verify. That means going back to the raw evidence, checking where each claim in a summary came from, re-running a query instead of trusting its description, and looking for what a summary leaves out. It also means forming a judgement before reading the AI’s verdict, which is the habit the exercises measure.
The fourth is data handling. Analysts need to know which data may go to which tool and why, which is the practical answer to people who use AI freely at home and find it restricted at work. SOC data routinely contains personal data, and the right tool is one approved for that class of data and that use.
The fifth is overrides, escalation and feedback. Analysts should know when to override and when to escalate, and what their feedback does. In Microsoft’s agent, feedback is recorded for audit and changes the agent’s behaviour only when an analyst saves it as a lesson, which can then influence similar future cases. I’d have a second person review any lesson before it changes production behaviour.
Which training is worth taking
External training and certification can build and test relevant skills, and some GIAC exams include practical, hands-on components. None of them tests performance with an organisation’s own agent, data and workflow, which is what the local exercises are for. With that caveat, this is what I’d choose by role.
For SOC analysts and incident responders, the first training is the team’s own exercises, plus the vendor’s material on what its agent can and can’t see. CompTIA SecAI+, launched in February 2026 and aimed at people with some hands-on security experience, is a reasonable broad baseline, though AI-assisted security makes up 24% of its exam and securing AI systems 40%.
For detection and automation engineers, GIAC has an AI Security Automation Engineer certification that covers applying automation and AI across defensive, offensive and cloud security work. Engineers who build or wire up LLM applications and agents are better matched by GIAC AI Platform Security, the certification attached to SANS’s GenAI and LLM application security course. For red teamers, GIAC lists an Offensive AI Analyst certification and an AI Penetration Tester certification that was listed for presale in October 2026. I covered the measures for red-team work in the piece on AI-enabled security testing. For security engineers who build their own models, GIAC offers the Machine Learning Engineer certification.
For security managers, ISACA’s Advanced in AI Security Management, launched in August 2025, covers AI governance, risk management and controls, and is open only to holders of an active CISM or CISSP. ISACA’s Advanced in AI Audit serves audit and assurance professionals with an active CISA or another qualifying designation. Someone has to own the literacy programme itself and its results, and in a large organisation the chief AI security officer can own it.
Where to start
I’d run a baseline exercise before buying any training, using the team’s own tools and adjudicated past cases, with the measures in the table reported at team level, counts included. Run it again with fresh cases after the training and some weeks later, after any change of model or agent configuration, and at least twice a year with part of it unaided. Record the exercises, the debriefs and the changes they led to as the team’s Article 4 measures. The Commission expects organisations to document those measures, and these records show what the training was for.
A completion record shows that training happened. It doesn’t show which mistakes analysts now catch, which correct advice they keep, or whether they can still investigate when the agent is unavailable. The gap between principles on paper and practice on the floor is one I’ve written about before, and an exercise with known answers is how a SOC finds out which side of it the team is on.
In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.
