loader image
A digital command centre map showing red attack lines spreading out from a single point across the world, representing an AI agent launching attacks on its own.

Picture this: an AI agent attacks more than 460 systems on its own, using nothing but public vulnerabilities that already had patches available. No hacker sat there guiding it step by step. Someone just pointed it at a target and let it run.

That’s not a thought experiment. It happened last week, and it’s the clearest test yet of a question I keep coming back to: does your AI governance actually work, or does it just look good in a document?

The agent in question was built on DeepSeek, an open model anyone can download, wired into a free agent framework anyone can use. One message on Telegram was all it took to set it loose. It scanned, picked its targets, and attacked on its own, breaking into at least three organisations, including through equipment sitting at the edge of company networks.

I’ve written before about Microsoft’s own agentic security system, Project Perception, and how it uses agents to find and fix problems rather than cause them.

This story is the flip side of that coin.

The same qualities that make agents useful for defence, working fast and acting without waiting for a human, are what make them dangerous in the wrong hands. If you’re a business or technical leader watching AI roll out across your Microsoft 365 setup, this is the story that should change how you check your own governance.


What Actually Happened, In Plain Terms

The attacker didn’t write custom malware or find a secret flaw. Every vulnerability the agent used in its confirmed breaches was public and already had a patch available, one had been fixed months earlier.

The problem wasn’t the code, it was that nobody had gotten around to applying the fix, and the agent found that gap faster than any human team could.

That’s the part worth sitting with.

This wasn’t a sophisticated nation-state operation. It was an off-the-shelf model, an open framework, and a bit of patience. The barrier to running an autonomous attack just dropped, and it’s not coming back up.

Around the same time, a separate benchmark tested how well AI agents follow written company rules when they’re dropped into a simulated business with a real handbook. The best agent stuck to the rules in only 36.2 percent of trials. Read that again: the best one.

Most agents ignored the rulebook more often than they followed it.


Why A Policy Document Won’t Save You

Here’s the bit that matters for anyone running Copilot, Claude, ChatGPT Work, or a mix of AI tools across their business. A lot of AI governance today lives in a document somewhere, a PDF that says what staff can and can’t do with AI, sitting in a SharePoint folder that nobody opens after the launch meeting.

That approach worked fine when the “AI” in question was a person using a chatbot in a browser tab.

It doesn’t work when the AI is an agent making its own decisions inside your systems, chaining tasks together, and touching data or tools you didn’t expect it to reach. A rule that only exists on paper can’t stop an agent mid-action, because the agent was never reading the paper in the first place.

This is the shift I talked through in building an AI operating model for Microsoft 365: governance needs an owner, guardrails, and a way to check what’s actually happening, not just a policy that sounds good in a board pack.

What’s changed in the last week is that the industry has started building the actual plumbing to make that possible.


The Market Is Catching Up Fast

Three separate security vendors launched agent governance products within days of each other, and that’s not a coincidence, it’s a signal that everyone’s racing to solve the same problem at once.

Drata launched a platform built to discover, monitor, and prove what every AI agent in a business is doing, starting with agents built on Anthropic’s models and expanding to OpenAI, Google, and AWS soon.

Airlock Digital extended its endpoint security to give IT teams command-level visibility into what trusted agents are doing on company devices, with real-time governance over what they’re allowed to do.

Zero Networks released something called Least Agency Enforcement, based on an OWASP principle that says you should give an agent only the access it needs for its one job, nothing more, with human approval required before it touches anything sensitive.

Different vendors, same idea: stop trusting the policy document, and start controlling what agents can actually reach and do in real time.

That’s the direction Microsoft’s own tools have been heading too, which is worth understanding if you’re already running Copilot or agents built through Copilot Studio.


What This Means Inside Microsoft 365

If your business has rolled out Copilot, or you’re experimenting with agents through Copilot Studio, this isn’t a hypothetical problem for someone else’s IT team to worry about.

Microsoft 365 E7 introduced Agent 365 as a control plane specifically for this, giving IT a single view of what agents exist, what they can access, and how they’re governed, and it works alongside

Purview and Defender to manage identities, data access, and compliance for agents the same way you’d manage a person’s account.

The trouble is, plenty of businesses I speak to have Agent 365 available and aren’t using it properly, or haven’t connected it to a real governance process yet, the kind I outlined in designing an AI operating model across Copilot and Claude.

Having the control plane switched on isn’t the same as having control. You need someone who owns the decision about what agents can do, someone checking what data they’re touching, and a process that catches drift before it becomes a headline.

If you haven’t looked at unified AI governance as a connected practice, rather than a stack of separate rules for each tool, this is the moment to start.

The businesses that get caught out won’t be the ones using AI, they’ll be the ones who assumed a policy document was doing more work than it actually was.


Four Things Worth Doing This Month

I’m not going to pretend every business needs to buy a new governance platform tomorrow. But there are a few practical steps that don’t cost much and make a real difference.

Get A Proper List Of Every Agent

Not just Copilot, but anything built in Copilot Studio, any automation with AI plugged in, and anything your teams have quietly connected through Claude, ChatGPT, or Perplexity. You can’t govern what you can’t see.

Check What Each Agent Can Actually Touch

Does it need access to your whole SharePoint site, or just one folder? Does it need to send emails on someone’s behalf, or just draft them for review? Cutting access back to only what’s needed is the single biggest risk reducer, and it costs nothing.

Put A Human Checkpoint In Front Of Anything Sensitive

Financial transactions, external communications, changes to systems, or access to confidential data. Agents can move fast, but that speed should stop at the door of anything that matters.

Make Someone Accountable

Not a document, a person. Someone in your business should be able to answer, right now, “what agents do we have, what can they do, and who approved that.” If nobody can answer that question today, that’s the gap to close first.


The Bottom Line

So, back to the question I opened with: is your governance actually working, or does it just look good on paper?

Agents are reliable enough to work without watching, and unreliable enough that you shouldn’t leave them unwatched.

That’s not a contradiction, it’s just where the technology sits right now. The businesses who’ll do well with AI aren’t the ones avoiding agents, they’re the ones treating them with the same discipline they’d apply to a new staff member with system access: clear boundaries, clear approval, and someone checking the work.

If you’re not sure how many agents are already running inside your Microsoft 365 environment, or whether your governance is actually enforced rather than just written down, that’s exactly the kind of conversation my team and I have with businesses every week.

Reach out, and we’ll help you get a clear picture before it becomes someone else’s headline.

CTA banner warning about rogue AI agents, with an unapproved agent alert and a prompt to book a governance health check.

About the Author

Carlos Garcia is the Founder and Managing Director of CG TECH, where he leads enterprise digital transformation projects across Australia.

With deep experience in business process automation, Microsoft 365, and AI-powered workplace solutions, Carlos has helped businesses in government, healthcare, and enterprise sectors streamline workflows and improve efficiency.

He holds Microsoft certifications in Power Platform and Azure and regularly shares practical guidance on Copilot readiness, data strategy, and AI adoption.

Connect with Carlos Garcia, Founder and Managing Director of CG TECH, on LinkedIn.

Sources