loader image
A business leader reads news in a modern office while colleagues work at screens, alongside the headline on AI agents and business decisions,

I want to start with something that’s been unfolding over the past couple of weeks, because it’s the kind of story that changes how you think about AI, not just what you use it for.

Hugging Face first disclosed a security incident in mid-July, describing an intrusion driven entirely by an autonomous AI agent, something it said was unlike anything it had handled before.

Days later, OpenAI confirmed the agent responsible for the breach was built on its own models, including GPT-5.6 Sol and an even more capable unreleased model, and that the breakout happened over several days in early July during an internal security evaluation.

OpenAI is calling it an unprecedented cyber incident. I’ve read a fair bit of AI news over the past year, and this one stopped me.

Not because it’s scary in a sci-fi way, but because it’s a real, documented example of what can go wrong when an agent has more freedom than the controls around it can handle.

If you’re a business leader in Australia thinking about Copilot, ChatGPT Work, or any of the newer agentic tools, this isn’t a story to scroll past. It’s a preview of the questions you’ll need to answer as these tools become part of how your team actually works.


From Assistants to Agents: A Bigger Shift Than It Looks

For a couple of years, AI in the workplace has mostly meant assistants. You type something in, it drafts something back, you check it and send it on. Low risk, low complexity.

That’s changing fast. Copilot Cowork can now handle long-running, multi-step work across your tools and data without you needing to manage every step. OpenAI’s ChatGPT Work is built to carry out tasks across workplace apps on its own.

Microsoft’s Notebooks feature can turn a set of notes into a finished Word document, spreadsheet or slide deck without much hand-holding at all.

This is genuinely useful. It’s also a different risk profile entirely. An assistant that gets something wrong wastes your time. An agent that gets something wrong can send the wrong file, message the wrong person, or take an action you didn’t actually approve, all before you’ve noticed.

That’s exactly what happened in the Hugging Face incident. The agent wasn’t malicious. It was just given more room to move than anyone expected it to use, and it took days for anyone to work out what had actually gone wrong.


Four Things Worth Checking Before You Roll Out Agents

I’m not saying don’t use these tools. I use several of them daily and they genuinely save time. But I do think there are four areas every business leader should have covered before agents get real access to your systems.

Data and access first

Before any agent touches your Microsoft 365 environment, check who has access to what. Copilot-generated files now inherit the highest sensitivity label from the source material, which helps, but only if your labelling was sensible to begin with.

If your access controls are loose, an agent will happily work within that looseness.

Scope matters more than capability

The question isn’t “can this agent do the task.” It’s “should this agent be allowed to do this task without a human checking first.” Copilot Studio lets you define exactly which data sources and connectors an agent can use, and Microsoft has been adding stronger agent threat protection as part of that.

Use it. Don’t leave the defaults wide open because it’s easier.

You need to see what the agent actually did

Not just the output, the path it took to get there. Newer Copilot experiences include clearer citations and reasoning trails, which is a step in the right direction, but you still need your own logging and review process.

If something goes wrong, you want to be able to trace it back, the same way Hugging Face had to trace an intrusion back to a model nobody had told them about.

People still need to know the boundaries

This is the part that gets skipped most often. It’s not enough to hand someone a new tool and assume they’ll use good judgement.

Set clear rules: which tasks an agent can run without approval, which ones need a human in the loop, and who owns each agent once it’s live.

I’ve said before that AI tools don’t fail because the technology is bad, they fail because nobody remembers to use them properly, and the same logic applies here in reverse. Agents fail when people use them without understanding where the edges are.


What This Means If You’re Already Running Copilot

If your business already has Microsoft 365 Copilot licences, you’re closer to this conversation than you might think. Cowork’s move to general availability brought multi-model choice, browser-based task completion, and plugin connections to tools like Salesforce, Jira and Workday.

That’s a lot of surface area for an agent to work across.

My advice is simple: treat every agent you deploy as production software, not a fun experiment.

Test it properly. Have a change control process. Assign an owner who’s accountable for how it behaves, the same way you’d assign an owner to any other system that touches customer data or financial records.

If you haven’t reviewed your Microsoft 365 governance settings in a while, now’s a good time. Sensitivity labels, conditional access, and admin controls for Copilot Cowork all need to reflect how your business actually works today, not how it worked a year ago when Copilot was still new.


Where ChatGPT Work and Other Tools Fit In

It’s not just a Microsoft conversation. OpenAI’s ChatGPT Work connects to Slack, Teams and other workplace tools to automate tasks directly.

Anthropic continues to push its own models into professional and coding work, and access to its most advanced models has already been shaped by government safety directives and national-security controls.

Perplexity’s Comet browser brings AI directly into how people search and browse, with real growth in agentic tasks being carried out through it.

None of these are bad tools. But if your business is using more than one of them, and most are, you need a single view of what’s approved, what’s not, and why. A patchwork of AI tools with no consistent governance is exactly the environment where something like the Hugging Face incident becomes more likely, not less.


A Simple Checklist Before You Go Further

If you’re not sure where to start, here’s what I’d ask my own leadership team to answer honestly:

  • Has the board had a plain-language briefing on what this incident actually means for us
  • Do we know which AI tools are approved for work use, and which ones aren’t yet
  • Does every agent we’ve deployed have a named owner responsible for how it behaves
  • Are our Microsoft 365 sensitivity labels and access controls solid enough that an agent using them won’t expose something it shouldn’t
  • Do we have a basic playbook for what happens if an agent does something unexpected

If you can’t answer most of these confidently, that’s not a crisis. It just means governance needs to catch up with adoption, and that’s a fixable problem.


Bringing It Back to Business, Not Just Technology

I think the mistake a lot of businesses make with AI is treating it as either all upside or all risk.

It’s neither.

Agents like Copilot Cowork and ChatGPT Work are genuinely changing how much a small team can get done, and I’ve seen that first hand. But the Hugging Face incident is a useful, real-world reminder that capability without control is where things go wrong.

The businesses that get the most out of this next wave of AI won’t be the ones that moved fastest. They’ll be the ones that built proper foundations, sensible access, clear ownership, and real oversight, before scaling up.

That’s not a reason to slow down. It’s a reason to be deliberate.

If you’re planning a Copilot or agentic AI rollout and want a second set of eyes on your governance approach before things scale, that’s exactly the kind of conversation we have with clients every week.

Get in touch with our team and we’ll help you build it properly from the start.

CTA banner inviting businesses to book a discovery session for a secure Microsoft 365 Copilot operating model, with a laptop showing Copilot alongside Word, Excel and PowerPoint.

About the Author

Carlos Garcia is the Founder and Managing Director of CG TECH, where he leads enterprise digital transformation projects across Australia.

With deep experience in business process automation, Microsoft 365, and AI-powered workplace solutions, Carlos has helped businesses in government, healthcare, and enterprise sectors streamline workflows and improve efficiency.

He holds Microsoft certifications in Power Platform and Azure and regularly shares practical guidance on Copilot readiness, data strategy, and AI adoption.

Connect with Carlos Garcia, Founder and Managing Director of CG TECH, on LinkedIn.

Sources