A few weeks back, security researchers caught something that’s been on my mind since.
A hacker wired an AI model into an automated system, gave it one instruction over Telegram, and let it loose on the internet to hunt for weak spots across hundreds of systems on its own.
That story’s been doing the rounds, and some of the headlines make it sound like an AI woke up one day and started hacking businesses by itself.
The real story’s a bit different, and honestly, more useful for anyone running a business on Microsoft 365 right now.
What Actually Happened
Palo Alto Networks’ Unit 42 team documented a Chinese-speaking hacker, working out of Zhuhai, who connected the DeepSeek model to an open-source agent tool called Hermes Agent.
The setup was controlled through Telegram, and with a single message, the agent got to work.
It scanned more than 460 internet-facing systems, looked up known vulnerabilities for the software it found, pulled exploit code off the internet, and tried to break into two specific targets: Langflow and n8n servers.
Here’s the part that matters most. Those automated attempts didn’t actually succeed. The targets had proper authentication and configuration in place, so the agent’s attacks failed.
But the same hacker wasn’t relying on the AI agent alone.
Running alongside those automated attempts, they also used more traditional hacking methods against other software, including Citrix NetScaler, Marimo notebooks, Apache Tomcat and Windows VPN devices. That’s where the real damage happened.
Unit 42 confirmed data theft from three organisations through a Citrix NetScaler flaw, and unauthorised access on 11 Marimo notebook servers through a separate bug.
Some coverage rounds this up as “14 successful attacks”, and that’s accurate, it’s 14 systems across those two vulnerabilities. It’s not 14 named businesses breached by an AI acting alone.
None of the reporting names the organisations involved, and nothing points to Australian businesses being among them.
So what really happened is this: an AI agent did the boring, time-consuming groundwork of an attack (finding targets, checking for weaknesses, picking exploits) at a speed no person could match.
Then a human took over and used ordinary hacking tools to get into a small number of systems that hadn’t been patched.
Why This Still Matters
It’s tempting to read that and think “well, the AI part didn’t even work, so what’s the fuss?” But that misses the point.
The systems that got breached weren’t hit by anything clever. They were running software with known bugs and available fixes that just hadn’t been applied yet. The AI agent’s job was to find those gaps fast, across hundreds of targets at once, and it did that part well.
That’s the shift worth paying attention to. It’s not that AI is about to start hacking everyone overnight. It’s that the gap between “a vulnerability gets published” and “someone’s agent is actively scanning for it” is shrinking.
The same agent-style automation we’re getting excited about for productivity is already being used on the offensive side too.
For a business running Copilot, workflow tools, or any AI-connected system, that’s a governance question as much as it is a security one.
Where This Leaves AI Governance
Most AI policies I come across are still focused on rules written down on paper: where AI can be used, what data it can touch, which use cases need sign-off. All useful, but none of it would’ve stopped an agent that’s already been given internet access and a goal.
Nothing in a policy document stops an agent from scanning, researching or acting once it’s connected to the right tools. The only thing that stopped this particular attack from working was technical, not a policy: the target systems required proper authentication and were configured correctly.
That’s a good way to think about AI governance going forward. It starts to look a lot like the way we already manage human access and permissions.
Which agents exist across your systems?
What can they see and touch?
What can they do without a person checking first?
Who’s watching what they actually do?
I wrote a while back about treating every agent like production software, with a clear owner, proper testing and a change process before it goes live. This is exactly the kind of scenario that makes that approach worth the effort.
What This Means Inside Microsoft 365
If your business runs on Microsoft 365, a good chunk of your AI activity is already happening inside tools like Copilot, Teams, SharePoint and Defender.
Microsoft’s own security work is heading in a similar direction, just on the defensive side. Project Perception is Microsoft’s new agentic security system inside Defender, using AI agents to find and prioritise risks automatically. It’s a sign that agent-driven work, both offence and defence, isn’t a future concept. It’s landing now.
That’s good news in one sense, because you’re getting more automated protection built into tools you already pay for. But it also raises the bar for how well you understand what’s running inside your own environment.
If you’ve got Copilot rolled out, you’re testing ChatGPT Work or Claude, or you’ve built workflows in Power Automate, Langflow or n8n, there’s a decent chance you’ve already got small “agents” making decisions inside your business, often without anyone having a single clear view of what they’re doing or what they can access.
That’s the exact blind spot this DeepSeek story points to.
Four Things Worth Doing Now
Know what agents you’ve actually got. Before anything else, ask your team for a simple list: which agents, bots and automations exist across your Microsoft 365 and Azure environment, what data they can reach, and what login details or permissions they’re using. If nobody can answer that clearly, you don’t really have visibility into your own AI use.
Control what agents can do, not just where they’re allowed to run. A policy that says “Copilot’s approved for these tasks” is a fine starting point, but it doesn’t control what an agent actually does once it’s working. You need rules about which systems it can touch, what changes it can make on its own, and when a person needs to approve something first.
Fix the boring stuff before worrying about the exotic stuff. The DeepSeek attack didn’t need anything fancy. It needed exposed systems, known bugs and missed patches. Keeping software updated, locking down admin access and enforcing multi-factor authentication will do more for your security than almost anything else on this list.
Put someone in charge of agent safety. Just like you’ve got someone accountable for cyber security and data protection, someone needs to own AI agent risk too. That doesn’t mean building a new department. It means security, IT and your AI leads working from the same view, with regular updates going to whoever owns risk at the top.
A good place to start is a straightforward AI inventory paired with clear risk levels, so you know what’s running before deciding how tightly to control it.
The Real Takeaway
This isn’t the moment AI started hacking the world on its own. It’s a clear example of something worth planning for: attackers using agents the same way businesses are starting to, to do more, faster, with less direct oversight.
If you’re building out Copilot, testing agents, or running more than one AI tool inside your business already, the plan doesn’t need to change. Keep going. Just make sure your governance shows up in the systems your agents actually touch, not just in a document nobody reads after the first sign-off.
If you’d like a second opinion on how your agents, Copilot rollout or AI workflows are set up, that’s a conversation I’m always happy to have.

About the Author
Carlos Garcia is the Founder and Managing Director of CG TECH, where he leads enterprise digital transformation projects across Australia.
With deep experience in business process automation, Microsoft 365, and AI-powered workplace solutions, Carlos has helped businesses in government, healthcare, and enterprise sectors streamline workflows and improve efficiency.
He holds Microsoft certifications in Power Platform and Azure and regularly shares practical guidance on Copilot readiness, data strategy, and AI adoption.

Sources
- Chinese Hacker Commands DeepSeek via Telegram to Launch Autonomous Attacks – The Hacker News
- Chinese Hacker Used DeepSeek Model To Attack 460 Systems On Autopilot – Forbes
- AI Struggles to Respect the Employee Handbook – Unite.AI
- Drata Extends Trust Management Platform to Govern AI Agents
- Airlock Digital Unveils Agentic AI Control & Governance
- Zero Networks Launches Least Agency Enforcement
- Microsoft Agent 365: The Control Plane for Agents