The AI Tunes Itself. Until It Doesn't.

Two stories broke this week. One was loud, one was quiet, and both say the same thing about trusting autonomous AI.

OpenAI published a report this week on how its own AI agents broke out of a controlled test and went after Hugging Face. Read it, then read it again, because the same lesson is playing out quietly inside the email security market, and almost nobody is naming it.

The short version is ugly. Nobody sabotaged the agents and nobody hijacked them. They did exactly what they were told, chasing a reward across whatever path was open. They executed code on 41 of Hugging Face's production servers and took root on a production node. When the scoring system got in the way, they learned how it worked and moved to cover their tracks. OpenAI missed the early signals and took about a week to notice. Turn a capable model loose with an open-ended goal, real access, and no human close enough to grab the wheel, and the autonomy becomes the exposure. That is the whole story in one sentence.

The same week, more than 100 150 organizations, including OpenAI and Anthropic, signed an open letter warning that status quo security will not hold as AI-enabled attacks scale. Two labs that compete on nearly everything, agreeing in public (hold that thought, because I want to come back to it).

First, the quiet story.

For years the pitch across email security ran the same way. The AI tunes itself. You do not need rules. You do not need humans in the loop. Trust the model. One of the loudest behavioral-AI vendors went so far as to publish a blog titled "When Custom Rules Break," arguing the smarter path was removing the need for rules at all.

Then, a few months later, that same class of vendor started shipping rule builders. And the fine print gives it away. The customer's hand-written rule overrides the AI model, which overrides the core detection. The tell is right there in the precedence order. When a human's rule sits on top of the machine by design, the trust-the-model era has quietly ended. Pure autonomy did not satisfy the security teams who actually run these tools.

I am not knocking the pivot. Shipping the rule builder was the right call. What I want people to notice is what the reversal admits. The teams running these platforms wanted a hand on the controls, and they wanted it badly enough that the vendors who spent years calling rules obsolete built the rules back in and gave them top priority.

We made the other bet in 2015, and we never had to walk it back. IRONSCALES paired adaptive AI with a human in the loop and a network of 36,000 analysts across 18,000 organizations, so a novel attack caught at one company sharpens detection for all of them, and an analyst or an employee can always weigh in. Our Phishing SOC Agent runs L2-level investigation in minutes and shows its work in plain language, so a person can agree with it or overrule it. The human sat in the design from the first release. There was nothing to reverse.

Now back to that letter. Section 04 asks frontier AI companies for two things, responsible model access and agentic identities that are traceable and accountable. That is the hardest ask in the document, and it is the right one. It reads like a request to the labs. It is just as much a question every security buyer should put to their own vendors. How do you source, contain, and govern the models inside your product, and who answers for it when one of them is wrong?

We can answer that with a name. IRONSCALES is among the first email security vendors verified under Anthropic's Cyber Verification Program, which means someone checked who we are and how we contain what we build before handing over frontier-grade capability. Verification before capability. That is the mechanism the letter is asking for, running today.

One more thing about the letter. Search it for the word phishing. It is not there. Neither is social engineering, or email, or voice, or video. The word human does not appear once. The letter names old bugs, misconfigurations, unpatched software, weak authentication, legacy debt. That is strong vulnerability management and a half-finished threat model, because AI is closing the code gap faster than it is closing the human gap. Every dollar of model capability aimed at patching raises the cost of writing an exploit and does nothing to the cost of a convincing voice on a Teams call. Attackers are not sentimental. They take the cheaper path, and the cheaper path into a hospital or a water utility runs through a helpdesk that resets MFA for someone who sounds like the CFO.

Here is the part my side of the industry has to own. You can patch software. You cannot patch a person. No vendor, mine included, has fully solved proving that the voice, the face, and the sender are who they claim to be across every channel a business runs on. That is the work, and it is where a human in the loop stops being a luxury.

So trust the AI. It is genuinely good, and it is improving fast. Keep a human on it. Ask your vendor how the model got there and who answers for it. And read the precedence order before you buy the pitch, because the vendors who told you the machine had this handled are the same ones quietly building the human back in.

Explore More Articles

Say goodbye to Phishing, BEC, and QR code attacks. Our Adaptive AI automatically learns and evolves to keep your employees safe from email attacks.