July 22, 2026 7 min

Who’s watching the AI? Cybersecurity’s Next Big Category Is Sitting in Plain Sight

IT security operator in a control room monitoring multiple screens displaying network data, system analytics, and world map visualizations

Every industry conversation about AI right now seems to circle the same warning: it will do most of the routine work, services revenue will compress, and headcount-based business models must change to stay competitive. That warning isn’t wrong. It’s just half the story. It only describes what AI may take away from the market as businesses transform. Almost nobody is talking about what AI is creating at the same time: a large, durable, currently unclaimed demand for the specific skill of checking whether AI did the job correctly.

There’s a 40-year-old piece of research that predicts exactly what’s happening now, from a field that has nothing to do with software. In 1983, cognitive psychologist Lisanne Bainbridge published a short paper called “Ironies of Automation,” based on years of studying industrial process control rooms. Her finding was this: the more comprehensively you automate a system, the more demanding, not less, the remaining human role becomes. Why? Because humans are left holding exactly the tasks nobody could figure out how to automate, plus a brand-new job nobody trained them for: supervising a system whose failure modes they no longer see often enough to recognize. Skills that go unpracticed deteriorate. An experienced operator who spends their days watching automation work, instead of doing the work themselves, quietly becomes an inexperienced operator without ever noticing the transition. The automation is usually right, until the day it isn’t.

Aviation gave this idea real stakes. In 1987, a Northwest Airlines flight crashed on takeoff from Detroit, killing 154 of the 155 people on board. The crew had grown accustomed to an automated system that checked whether the flaps and slats were correctly configured for takeoff. That day, the automated check had been silenced by a tripped circuit breaker. The crew, used to the machine catching that error, didn’t manually verify it themselves. The plane took off unconfigured and didn’t make it. Nothing exotic went wrong that day: just a very ordinary, very human failure to keep practicing a check that automation had quietly made feel unnecessary.

That’s the pattern. And it’s showing up again, right now, in every field that has adopted AI assistants at scale, and people are already living through it. Junior lawyers are offloading legal research and first drafts to AI, the exact repetitions that used to build legal judgment, and firms are openly worried their new hires aren’t developing the ability to evaluate AI output at all. Several 2026 industry surveys on software engineering point to a similar problem from a different angle: junior developers who lack grounding in architecture and security can’t reliably judge whether AI-written code is good, and default to trusting the AI over their own instincts, precisely because they never built the instincts to trust instead. It’s been put more bluntly at industry security events: junior engineers raised on AI-assisted coding increasingly lack basic grounding in networking and protocols, to the point where teams struggle to even explain a security risk internally, let alone catch one.

So, the question people are asking but very few are providing an answer: yes, the mundane, repetitive tasks are going to get automated, that part of the story is true and it’s not really in dispute. But who is going to have the skill to validate what AI produced? Who’s going to be able to look at an autonomous system’s output and know, from real hands-on grounding, whether it’s right? And when something goes wrong, when the AI needs to be stopped, corrected, or restarted mid-task, who still has the muscle memory to do that? Everyone is racing to build automation. Who’s building the capacity to check it?

The cybersecurity version of this problem

Cybersecurity is the sharpest version of this problem right now because the automation in cyber security disciplines isn’t coming, it’s already here, running unsupervised, in production.

Every major SOC platform is moving from “copilot” (AI answers questions, a human acts) to “agentic” (AI acts, a human is notified afterward). Autonomous triage agents are now closing low-risk alerts and triggering containment actions on their own, at high self-reported accuracy, measured, naturally, by the vendor who built the system, against that vendor’s own labeled data. There’s no independent party currently checking that number. And practitioners are visibly split on how much to trust it: it’s now common to hear security teams admit they override AI-generated recommendations rather than act on them, because the output sounds confident even when it’s occasionally wrong. Busy teams get complacent, and AI has been trained on academic papers, so there is inherent bias towards confidence.

The same pattern is playing out on the offensive side. Autonomous AI pentesting agents are now finding, and reporting, real vulnerabilities faster than any human team could. That’s a real achievement. But it has already broken the pipeline downstream of discovery: at least one major bug-bounty platform has paused a long-running program and cut payouts after AI-assisted research pushed submission volume far beyond what maintainers could triage, and multiple open-source projects have suspended their bounty programs entirely over a flood of plausible-sounding, low-quality AI-generated reports. The constraint in offensive security has visibly shifted from finding problems to verifying them, and almost nobody is selling the verification.

This is, very precisely, a validation gap, and cybersecurity doesn’t have a name for it yet. So, let’s give it two.

AVaaS: AI Validation-as-a-Service. It borrows the naming convention security buyers already understand from PTaaS (Pentest-as-a-Service) and MDR (Managed Detection and Response), applied to a category that doesn’t have a name yet. An independent party’s entire job is to check what your AI actually decided against what it claims to have decided. That means sampling autonomous SOC actions against ground truth the AI didn’t design. It means reviewing autonomous pentest findings the way a skeptical senior tester reviews a junior’s report: not just whether it hit the target, but whether the path to get there was sound. This is not an eval, and it is not an LLM-as-judge setup wearing a new name. AVaaS is a human, independent, and accountable check — the specific thing a regulator, a board, or a client needs signed off, and the specific thing an eval was never designed to provide.

AJQ: AI Judgment Quotient. The individual-level version of the same idea: a way of naming the specific, trainable skill of knowing when to trust an AI’s output and when to push back on it, separate from knowing how to prompt an AI well, which is the skill everyone’s currently obsessed with. Prompting gets you a better answer, faster, just like it you ask a human a question. AJQ is what tells you whether the answer is right. Nobody is hiring for it by name yet. That won’t last.

The compliance tailwind almost nobody’s pricing in

There’s a regulatory hook here too, and it’s worth being precise about it, because the generic version of this argument overstates it. Most everyday cybersecurity AI, a SOC copilot triaging phishing, a pentesting agent scanning a SaaS app, isn’t automatically caught by the EU AI Act’s high-risk rules. But one category inside the Act lands directly on cybersecurity: AI systems used as a safety component in the management and operation of critical digital infrastructure: the utilities, OT, and ICS environments where a security or anomaly-detection system’s failure could have physical consequences. Those are high-risk by default, and the Act requires genuine, working human oversight: a person who can monitor, understand, override, and halt the system in practice, with that capability demonstrated rather than assumed.

Regulators aren’t going to be satisfied by a policy that says a kill switch exists. They’re going to ask whether anyone has tried to pull it under pressure and confirmed it works. That’s a specific, testable claim, and one that’s easy to sell. The deadline for it just moved later, to December 2027. That later date buys a multi-year runway to become the obvious, credible, evidenced vendor for this before every advisory firm on earth starts pitching the same slide.

That’s the tangible space. A genuine market category, with no incumbent, built where three things come together, all independently, verifiably true right now: AI is already making unsupervised security decisions in production; the people who could historically catch its mistakes are the same people whose foundational skills are quietly eroding from disuse; and a regulator is about to start asking, in writing, whether anyone actually checked. So, let’s see who builds it first.

Shilpi Handa

Shilpi Handa - Associate Research Director (META), IDC

Shilpi Handa is an associate research director at IDC, with responsibility for the Middle East, Turkey, and Africa cybersecurity practice. Her core research coverage revolves around cybersecurity, with a focus on network security, cloud security, application security, and security operations.…
Shari Lava

Shari Lava - Group Vice-President, AI, Data, and Automation

Shari Lava is Group Vice-President, AI, Data, and Automation. Ms. Lava’s core research coverage includes the fast-evolving AI software market, as well as the Automation and Data foundations essential for deploying AI at enterprise scale. This includes deep analysis of…

Subscribe to our blog