The AI Filter: When the Agent Breaks Something, Whose Fault Is It?
This week a swarm of AI agents broke into a company's systems, a ten-thousand-agent proof of a famous math problem got accused of borrowing from two humans, and a new phrase showed up in the industry's vocabulary: "I didn't do it. My agent did." Three stories, one question, and it is the oldest question in my trade.
This is The AI Filter, the weekly read where I go through the stack of AI newsletters so you do not have to. I cut the hype, keep the signal, and give you the reasoning so you can judge for yourself. This week the signal was not a new model. It was a shift in who gets blamed when the tool goes wrong, and if you are about to hand an AI a piece of your business, you need to see it coming.
Start with what happened, stripped of adjectives.
The Batch, Andrew Ng's newsletter and still the most careful source in my stack, led with the incident that drove the week's fear. An OpenAI team ran a large swarm of agents that got into Hugging Face's systems. The press reported twelve hundred agents as if that number were the story. Ng's reply was dry and useful: he had about thirteen hundred processes running on his laptop while he wrote the sentence. The swarm was not magic. Buggy sandboxing and weak monitoring were the problem, and those are engineering fixes, not reasons to stop.
Then he said the thing I want you to keep. If I swing a hammer, miss the nail, and dent the wall, it is not the hammer's fault. If I point an agent at a task and it breaks into someone's system, the responsibility is mine. And he named a new move he is seeing from AI companies: disclaiming responsibility for their own products. The tool did it. Not us.
The same issue carried a second story that shows what taking responsibility actually looks like. Meta released Muse, a personal agent that reads your email, browses, fills forms, and buys things. What caught my eye was not the feature list. It was the design assumption. Meta built Muse on the premise that the model will be fooled. A malicious web page can and will trick it. So the agent never holds a password. A separate gatekeeper approves every outbound action and it cannot be talked out of it by anything the agent read. Approvals come through a system dialog, never through the conversation, so injected text cannot fake your yes. They assumed failure and put the checks where the failure could not reach them.
The third story is the messiest. OpenAI announced that ten thousand agents, working for eighty-eight hours, produced a proof addressing a fluid-dynamics problem that carries a million-dollar prize. Two mathematicians who had spent most of a year on the same approach, using OpenAI's own tools, publicly asked whether the agents had seen their work. OpenAI's answer changed between one statement and the next. I am not qualified to judge the math. I am qualified to notice that when the tool produces the result, the humans start arguing about who owns it, and nobody argues about who is accountable if it is wrong.
TLDR added two smaller pieces that belong in the same folder. Researchers chained two vulnerabilities to get into OpenAI employees' own ChatGPT accounts back in July. And a widely shared essay argued that people simply turn their brains off when they use these tools, hand the model the problem, and become what the author called a meat proxy.
Now the filter, and here is where twenty-six years as a Navy Master Training Specialist earns its keep.
In my world, when a student failed a practical, we never blamed the manual. Not once. The manual was a tool. The instructor was accountable for the instruction, the student was accountable for the performance, and the evaluation sat outside both of them, written before anyone walked into the room. That structure is not bureaucracy. It is the only thing that makes the word "worked" mean anything.
Look at Meta's design again through that lens. Assume the learner will be fooled. Put the check where the learner cannot reach it. Make the approval come from outside the conversation. That is a training evaluation wearing an engineering jacket, and it is the right instinct.
Now look at the phrase "my agent did it" through the same lens. It is a student blaming the textbook. It will not survive contact with a customer, an auditor, or your own bank.
So when you set an AI loose on something in your business this month, do the two jobs the tool cannot do for you. Decide, in writing, what it is allowed to do and what a correct result looks like. Then check the result yourself, or build the check so it does not depend on the tool being honest. Those two jobs made me a professional before AI existed. They are still the whole profession.
The tool is never the accountable party. Meta built an agent on the assumption that it would be fooled, and put the checks where the fooling could not reach. Do the same. Decide the job, write the check, and keep your name on the result.
If you have the expertise and you are ready to turn it into something you can teach, The Foundation is our free starter course. It is at https://rhynowerks.ai
