BLOG · MODELS AND INFRA

OpenAI slowed work on its own model because it cannot rule out that the model will find and exploit flaws by itself

Models and infra
  • #modele-i-infra
  • #bezpieczenstwo-agentow
  • #ryzyko-cyberbezpieczenstwa-ai
  • #preparedness-framework
  • #operator-lens

On 7 August OpenAI announced it is treating its upcoming Astra model as the first model in the company’s history at a "Critical" level of cyber capability — and paused the internal work that does not meet its tightened security requirements. The reasoning is more careful than the headline: preliminary evaluations do not allow that level to be ruled out. What is new is not what the model did, because it did nothing. What is new is that a vendor took an operational decision on an unresolved evaluation — and the list of safeguards it published along the way reads like a specification for anyone running agents in production.

Adam WszendybyłAI operator-architect

What happened

On 7 August OpenAI published a note about an upcoming model called Astra and did something it had never done before: it classified one of its own, still-unreleased models as its first "Critical" model for cybersecurity under its preparedness framework.

The reasoning is worth reading literally, because it is more careful than the headlines that followed it. The company did not say the model crossed the threshold. It said preliminary evaluations show performance strong enough that the Critical level cannot be ruled out at this time. That is a precautionary classification, not a measurement.

The threshold itself is narrowly defined. A model is Critical when it can, without human intervention, identify and develop functional zero-day exploits against hardened real-world critical systems — or devise and execute end-to-end novel cyberattack strategies against a hardened target given only a high-level goal.

The response was not a statement but a set of constraints. OpenAI paused the internal work on Astra that does not meet the tightened requirements, and it lists what it added: isolated test environments, restricted network and tool access, stronger model-weight protection and encryption, and monitoring across all agentic uses that halts high-risk activity. External testing is to be carried out by government agencies and select AI safety organizations; the UK's AI Safety Institute is named.

One caveat from the same note is easy to miss and changes how you read it: Astra was not involved in July's Hugging Face incident, which we covered when writing about blast radius. That is a separate matter and a separate model.

Our read

Every recent story of this kind was a report on something that had already happened. This one concerns something that has not happened — and that is precisely why it is more serious.

A vendor took an operational decision on an evaluation it had not resolved. Not "we confirmed the threat, so we are acting", but "we cannot rule it out, so stricter rules apply until it is settled". That is exactly the decision companies do not take, because in a risk committee the line "there is no evidence" wins. Note who took the other side of that line this time: not a regulator, but the producer itself, at the cost of its own schedule.

The second point is more practical. The list of safeguards OpenAI applied to itself is public, and it reads like a specification. Four items are familiar to anyone who has run agents: isolation, restricted network and tool access, secret protection, monitoring. The fifth is different, and it is the one to look at — monitoring that halts high-risk activity rather than merely recording it.

That difference, between an alert and a stop, is where most agentic deployments actually end. An alert reaches a person who will read it in the morning. The agent works at night. The lab with the most exposure in this market has just decided that recording alone is not enough for it — a clearer signal than any benchmark.

There is one further conclusion that goes beyond what is confirmed, so we label it plainly: what follows is speculation, not a finding. A capability one lab describes in its own systems rarely stays with one lab for long, and we do not know how much time separates this note from the moment something of comparable strength is working on the other side. What is clear is that the window between "a vendor says it can do this" and "this is widely available" tends to be the only preparation time a defender gets. If you treat AI cybersecurity risk as a line in your risk register, this is the week to refresh it — not because something happened, but because it stopped being a hypothesis.

What not to take from this story: that you need to buy something. There is no product to deploy here, and nobody selling a solution to this problem.

Why it matters

Private Equity

For a fund this is one extra line in the portfolio review, and a cheap one. If the attacker's cost of finding a flaw really does fall, the most exposed companies will be those with a large surface and a thin security bench — the typical company after two add-on acquisitions, with inherited environments nobody on the team remembers any more. The question for the next review is not "are we secure", because the answer is always yes. It is: who here owns getting a security patch onto externally visible systems within a defined time, and how do we know it is happening. A company with a name and a number of days attached to that is in a different place from a company with a policy. How we structure that review on the fund side: owner and cadence first, tooling second.

Enterprise

A large organization has both monitoring and an inventory. What is new is carrying the "record versus halt" distinction across to its own agentic deployments. If an agent in your organization reaches for tools that change system state rather than only read it, it is worth knowing what happens when it behaves unusually at two in the morning. "It gets written to the log" is an answer — just not the one the model vendor chose for itself.

The second thread is procurement. At the next conversation with a vendor, it is worth asking something nobody used to ask: what capability thresholds are defined, who measures them, and what happens to your access when one is crossed. Astra shows this is not a theoretical question — a model's schedule can move for a reason that has nothing to do with the product or the market.

SMB / mid-market

A mid-sized company will not build a security team and does not need to. What it does have to change is one thing that needs no budget: the calendar. We covered the inventory of what is exposed to the internet in late July — if that list exists at your company, the other half of the job is speed. How many days pass between a patch being published and it landing on what is visible from outside? If the answer is "whenever someone remembers", that is the variable a cheaper adversary hits first, because it does not have to pick a target — it just scans. This needs no tool, only a date in the calendar and one person who keeps it.

One move this week

Take one agent running in production that has access to state-changing tools — it sends, it writes, it calls someone else's API. Answer one question: what specifically stops it mid-action if it starts doing something you did not anticipate. Not what records it — what stops it. If the answer is a person, add the second part: what hours does that person watch. The whole diagnostic takes fifteen minutes and usually ends in one line of configuration rather than a project. Describe your case: mailto:[email protected]?subject=Rozmowa%20z%20Aurora%20AI.

Reading us regularly? Set us as a preferred source.

In your Google search settings you can add aurora-ai.pl as a preferred source — our analysis will then surface more often in your results.

LET'S START

Bring the process, not the slides.

If you read our blog and spot an area you want to improve in your own organization — write to us. We start every conversation from something concrete.