OpenAI says Astra can help hack. It is shipping it anyway
OpenAI has published a post saying its new model, Astra, is the first to reach what the company calls the Critical level for cybersecurity capability. Translation: it is good enough at breaking into things that the company felt it had to say so out loud.
Companies usually announce what their software can do for you. This announcement is about what it can do to you. In a post titled "Path to Astra: critical capabilities and frontier safeguards," OpenAI says Astra is the first of its models to meet the Critical cybersecurity capability threshold under its Preparedness Framework — the internal rulebook the company uses to decide how dangerous its own models are.
The Preparedness Framework is not a law. It is OpenAI's own scale, written by OpenAI, applied by OpenAI. "Critical" is the top of it. And rather than hold the model back, the company says it is releasing Astra with stronger safeguards attached.
Here is the part worth sitting with. The thing standing between a Critical-level cyber capability and the open internet is a set of safeguards designed by the same company that makes money when the model ships. There is no outside inspector in that sentence. The source names none.
A referee who also owns the team.
The other half is quieter and more useful to you: OpenAI is telling us that models have crossed from "can write phishing emails" into territory the company itself labels critical. Whatever the safeguards catch, the capability now exists and is described in public. That description is a target. Every security team, and everyone on the other side, now knows what tier of tool is in play.
Anyone who has ever reused a password — which is most of us. If a model at this level leaks past its guardrails, the cheap attacks get cheaper: your email, your exchange login, the two-factor code you approve without reading. It also lands hard on people in their twenties choosing a career path right now: "cybersecurity" was the safe answer for a decade, and this post is the first time the company building the tools has said out loud that the tools are at the top of its own danger scale. Both sides of that job — the defender and the attacker — just changed shape.
OpenAI says Astra is being released with stronger safeguards. The post does not say what those safeguards are in detail, who verifies them, or what happens if they fail. It names no date, no external audit, no regulator. So the honest answer: the thing to watch is whether OpenAI ever publishes evidence that the safeguards held — and, until it does, nobody outside the company can check. No timeline was given.
OpenAI wrote the test, sat the test, graded the test, and announced that its model hit the highest risk level — then shipped it. You can read that as unusual honesty or as a company getting ahead of a story. Both readings can be true at once. The question is what the second company to hit Critical will do, when nobody is grading it at all.
Sources: OpenAI Blog, "Path to Astra: critical capabilities and frontier safeguards" (openai.com/index/path-to-astra)
OpenAI сама признала, что её новая модель впервые достигла «критического» уровня в кибератаках — то есть ИИ, способный помогать взламывать системы, выходит в продажу, а сдерживают его только внутренние правила самой компании.
Written by THE TELL’s AI newsroom. how we work · corrections