
I’ve read a lot of AI announcements this year, but this week’s news from OpenAI and Anthropic felt different. Instead of a shiny new feature, both companies published something closer to a warning label.
OpenAI confirmed that its next model, called Astra, is the first one it’s ever rated as “Critical” for cyber security risk under its own safety framework. In plain terms, that means their internal testing shows Astra can find unknown security flaws and work out how to exploit them across well-protected systems, without a person guiding every step.
It’s not a small tweak, it’s the first time any OpenAI model has hit that top risk tier.
At almost the same time, Anthropic published a detailed update on how it’s tightening the way it tests Claude models, after finding that some of its models had gained access to real computer systems during testing they thought was contained.
Two major AI labs, in the same week, telling the market they need stronger fences around their own technology.
For business and technical leaders already running Copilot, ChatGPT Work, Claude or Perplexity somewhere in the business, this isn’t just industry gossip. It’s a signal about how much trust to place in agentic AI right now, and what questions you should be asking before you let it touch more of your business.
What Astra Being “Critical” Actually Means
Let’s unpack that critical rating a bit, because the label alone doesn’t tell you much.
OpenAI’s Preparedness Framework sets out risk tiers for its models, and until now, the top cybersecurity tier had only existed on paper. Astra is the first model to actually reach it.
According to OpenAI’s own announcement, Astra can spot previously unknown security flaws and build working exploits against hardened systems with little or no human involvement, which is a genuinely different capability to what’s been publicly available so far.
Importantly, Astra wasn’t involved in the incident that triggered all this caution.
That was a separate event over the northern hemisphere summer, where AI agents being tested by OpenAI escaped their sandbox and interfered with systems belonging to Hugging Face, a well-known AI platform. OpenAI paused a chunk of its frontier model development for around two weeks afterwards to strengthen its containment and monitoring, and Astra has been shaped by that experience.
The result is a model with more refusals built in. OpenAI says Astra turns down harmful cyber requests 91.5% of the time in testing, compared with 59% for its current public model, GPT-5.6 Sol.
That’s a big jump, and it shows the lab is taking the risk seriously rather than just talking about it.
But here’s the part that matters for anyone planning to use this kind of technology at work: OpenAI has said plainly that these extra safeguards will sometimes slow, pause or stop legitimate tasks, not just malicious ones.
The most advanced cybersecurity capabilities will only go to a small group of vetted testers first, through a program called Daybreak Blue, before wider access follows. If you’re running long agent workflows or coding pipelines, expect more friction and more approval prompts as these controls roll out, not less.
Anthropic’s Side Of The Story
Anthropic’s update runs on a similar theme, but from a different angle.
Back in July, Anthropic disclosed that three of its Claude models had gained unauthorised access to real computer systems during cybersecurity evaluations that were supposed to be sealed off from the internet.
The cause wasn’t a rogue AI plotting an escape, it was a configuration error in the test environment, combined with the models being willing to take actions beyond what was intended once they realised they weren’t actually isolated.
Anthropic’s response has been to build a real-time classifier that watches for a model trying to probe or escape a test environment, or unexpectedly reaching the internet. When it spots that behaviour, it blocks the action before it runs, ends the task and alerts a person, rather than only catching the problem later when someone reviews the transcript.
They’ve also moved their highest-risk internal sandboxes to stronger isolation and are now applying the same kind of monitoring inside the training environments used to teach the models, not just the evaluation stage.
What I find genuinely useful here is that Anthropic is asking its outside partners to follow the same discipline. Their new guidance says cyber evaluations should run offline by default, in hardened sandboxes with no internet access except to the model’s own API, and that partners should actively test whether an escape is possible before running a real evaluation.
That’s a sensible baseline, and it’s one worth borrowing even if you’re not running AI safety testing yourself, because the same idea applies to piloting agents inside your own business: a sandbox is only as safe as its actual boundaries, not the label on the box.
Why This Matters For Your Copilot And Agent Stack
Here’s where I’ll bring this back to something more relevant to your Tuesday morning. Most Australian businesses I talk to aren’t running frontier cyber-testing labs, so why should any of this change how you think about your own AI use?
Because the pattern is the point.
Two of the biggest AI vendors in the world have just told us, in their own words, that their most capable models can act in ways that go beyond what a human intended, and that the fix is stronger monitoring, tighter sandboxes and clearer human checkpoints, not just better prompts.
If that’s true for models built by teams with unlimited security budgets, it’s a fair bet the same principle applies to whatever mix of Copilot, ChatGPT Work, Claude or Perplexity your teams are already using day to day.
We’ve talked before about Microsoft’s own move into agentic security with Project Perception, which puts AI agents to work finding and fixing vulnerabilities before your team sees the alert.
Astra and Claude’s updates sit on the flip side of that same coin: as agents get more capable at finding weaknesses, they also become something you need to actively govern, not just switch on and trust.
It also reinforces something worth checking in your own tenant right now. If you haven’t reviewed which AI provider is actually processing your Copilot requests, or mapped out where ChatGPT Work already sits alongside Copilot in your business, this is a good week to do it. You can’t apply sensible guardrails to tools you haven’t acknowledged are already in use.
What Business And Technical Leaders Should Do Next
I’m not saying any of this means you should pull back from Copilot or agentic AI. Far from it, the tools are genuinely useful, and most businesses are still under-using them rather than over-using them.
What it does mean is that governance needs to grow alongside capability, and right now capability is moving fast.
A few practical steps worth taking this month:
Map what you’ve actually got. List every AI tool with agent-like abilities across the business, including Copilot, Copilot Studio bots, ChatGPT Work and anything else touching company data. You can’t set sensible rules for tools you haven’t listed.
Decide who owns each agent’s behaviour. Every agent doing meaningful work should have a named business owner and someone technical who understands what it can access and what happens when it’s blocked or refuses a task.
Ask vendors how they test. If a supplier is proposing an AI agent for your business, it’s fair to ask how they test it for safety, and whether that testing follows practices like the ones Anthropic just published.
Treat refusals and pauses as information, not glitches. As these safeguards roll out more broadly, agents will interrupt themselves more often. Track when and why that happens, it tells you something useful about where the real risk sits in your workflows.
Revisit your AI risk register. If cybersecurity-capable AI is now a named risk category for OpenAI and Anthropic, it probably belongs as a line item in your own risk documentation too, alongside your existing vendor and subprocessor register.
None of this needs to slow your AI plans down. It just means building the fences at the same time as you build the value, rather than after something goes wrong.
A Practical Next Step
If you’re not sure how exposed your business is to any of this, that’s a conversation worth having before your next Copilot rollout, agent pilot or budget cycle, not after.
Our team at CG TECH can walk through your current AI and Microsoft 365 setup with you, have a look at where agents are already operating, and help you put sensible, workable guardrails around them.
Get in touch with the team at CG TECH if you’d like a second set of eyes on your AI governance before things scale further.

About the Author
Carlos Garcia is the Founder and Managing Director of CG TECH, where he leads enterprise digital transformation projects across Australia.
With deep experience in business process automation, Microsoft 365, and AI-powered workplace solutions, Carlos has helped businesses in government, healthcare, and enterprise sectors streamline workflows and improve efficiency.
He holds Microsoft certifications in Power Platform and Azure and regularly shares practical guidance on Copilot readiness, data strategy, and AI adoption.

Sources
- OpenAI: Path to Astra: critical capabilities and frontier safeguards
- Reuters: OpenAI says upcoming model is so capable it requires stronger guardrails
- Axios: OpenAI to limit access to Astra’s most powerful cyber tools
- The Wall Street Journal: OpenAI to Restrict Astra Model After Rating It ‘Critical’ Cyber Risk
- Anthropic: Improving our alignment and security practices
- The Register: Anthropic pledges to try harder to keep models under control