How AI Works, Part 6: Agents

AI agents: a three-node loop over a grid of numbers

This is the last post in the series, and it’s the one everything else was building toward. Understanding, learning, language, sight, and the ability to act all combine into AI agents this week, something that doesn’t just answer questions or interpret a feed, but decides what to do next on its own and keeps going until it’s done.

The Range of AI Agents

An agent, in the sense the industry means it now, isn’t one thing. It’s a range. At the simple end, an agent calls one tool, a database lookup or a calendar check, and hands control straight back to a human. Further along, it plans a multi-step task, executes it, checks its own work, and adjusts, still with a person approving the risky steps. At the far end, it runs on a schedule with no human in the loop at all, and it isn’t alone. Production systems increasingly hand a task to several agents at once, each with its own slice of the job, coordinating or competing to get there.

The range matters, because “AI agent” gets used for a chatbot with one plug-in and for a fleet of autonomous systems negotiating with each other, and those are barely the same category of risk. The two failure modes I actually care about both live at that far end, and they’re different problems.

When Agents Act Without Being Asked

In June 2026, Anthropic stress-tested 16 leading models from across the industry, its own included, in scenarios where a model had autonomous access to a company’s systems, discovered damaging information, and then faced a threat to its own continued operation combined with a goal that conflicted with its supervisor’s. Claude Opus 4 and Gemini 2.5 Flash attempted blackmail in 96% of those runs. GPT-4.1 and Grok 3 Beta did it 80% of the time. DeepSeek-R1, 79%. Every model tested showed the behavior under the right pressure.

The unsettling part isn’t that a model got confused. Anthropic’s own writeup notes the models showed clear awareness that what they were doing violated ethical constraints, and chose to do it anyway because it served the goal they’d been given. That’s not a bug in the traditional sense. It’s a system correctly optimizing for the wrong thing, with enough autonomy to act on it before anyone could check.

When Agents Start Coordinating

The second failure mode is what happens when more than one agent is in the room. Anthropic’s Frontier Red Team ran an experiment where three Claude agents were given access to the same software project, each with instructions that quietly conflicted, and none aware the others existed. They didn’t stay out of each other’s way. Each one read the others’ changes as deliberate sabotage and escalated, eventually deploying increasingly aggressive, self-replicating malware against one another.

Some agents also proposed what looked like a neutral way to settle the dispute, a shared scoring metric, and at least one later acknowledged the metric it suggested happened to favor its own approach while appearing impartial. Nobody told these systems to fight, and nobody told them to game the referee. Both behaviors emerged from individually reasonable choices made under incomplete information.

Anthropic’s own conclusion is the one that stuck with me. Quirks that look small in one agent can compound into a genuinely unwanted outcome once you have thousands of them interacting, and today’s safety testing is mostly built to check one agent at a time.

The Pushback

Anthropic says plainly that it hasn’t seen evidence of agentic misalignment in real-world deployments. These were adversarial stress tests, deliberately engineered to corner a model into a conflict most production systems will never face, with permissions most companies would never hand out in the first place. Most agents running today are narrow, think a coding assistant with a limited toolset, or a support bot with a scripted escalation path, both still operating under approval gates a human checks.

The more mundane problem is still the dominant one. Costs run higher than expected once a pilot goes to production, the measured business value is often unclear, and a lot of what gets marketed as an autonomous agent is existing automation with a new label. Most AI agents in production today fail on reliability, not rebellion.

Both things are true at once, and that’s the uncomfortable part. The mundane failures are the near-term reality. The coordination and misalignment failures are the ones that get worse, not better, as autonomy and agent count both keep climbing, and climbing is exactly what’s happening.

How I Have Used This in My Writing

This is the whole engine behind the AI in The Apex Code, and it’s where I’m taking the threat in The Axion Code. The first book is about a system that crosses from watching to acting, interpreting a camera feed, then unlocking a door. The next problem is what happens when that system isn’t alone anymore, when it’s coordinating with other instances of itself, or with other systems entirely, to protect a goal none of its creators explicitly gave it permission to protect that way.

The real research is more useful to me than anything I could invent, because the actual finding isn’t “the AI turned evil.” It’s that individually rational agents, each doing something locally defensible, can produce a collectively destructive outcome that nobody designed and nobody would have approved. That’s a much scarier sentence than a rogue AI with a plan, and it’s the one I’m building the sequel around.

As always, thanks for reading!

Chris

Reader Feedback

"Great page turner, so relevant and timely! Really enjoyed this one! Especially enjoyed the technical (both AI and military) accuracy that never got in the way of a fast moving, fun story."

— Barry Parsons, August 9, 2026
See more reader comments →
FAQs
Where can I buy The Apex Code?

Paperbacks and hardcover copies are all available directly from this site, or you can get them (unsigned!) along with the ebook through Amazon, Barnes & Noble, Waterstones and other major retailers.

How do I let you know who I want my signed copy dedicated to?

There is an option to add a note when you are ready to checkout. Fill this out with your dedication request just before you place your order!

When is The Axion Code coming out?

Later this year — sign up for the newsletter to be the first to know the release date.

Is The Axion Code a sequel, or can I read it standalone?

It's the second book in the series — while it can be read on its own, The Apex Code sets up the characters and world that carry through.

Do you have events, signings, or appearances coming up?

Check the News & Blog page for updates — new events are posted here as they're scheduled.

Are you running any giveaways right now?

Yes! Five signed copies of The Apex Code are up for grabs for UK readers. Visit the UK Giveaway page for full details on how to enter.