THE TELL

OpenAI says Astra can help hack. It is shipping it anyway

OpenAI has published a post saying its new model, Astra, is the first to reach what the company calls the Critical level for cybersecurity capability. Translation: it is good enough at breaking into things that the company felt it had to say so out loud.

Companies usually announce what their software can do for you. This announcement is about what it can do to you. In a post titled "Path to Astra: critical capabilities and frontier safeguards," OpenAI says Astra is the first of its models to meet the Critical cybersecurity capability threshold under its Preparedness Framework — the internal rulebook the company uses to decide how dangerous its own models are.

The Preparedness Framework is not a law. It is OpenAI's own scale, written by OpenAI, applied by OpenAI. "Critical" is the top of it. And rather than hold the model back, the company says it is releasing Astra with stronger safeguards attached.

What it means

Here is the part worth sitting with. The thing standing between a Critical-level cyber capability and the open internet is a set of safeguards designed by the same company that makes money when the model ships. There is no outside inspector in that sentence. The source names none.

A referee who also owns the team.
Share this

The other half is quieter and more useful to you: OpenAI is telling us that models have crossed from "can write phishing emails" into territory the company itself labels critical. Whatever the safeguards catch, the capability now exists and is described in public. That description is a target. Every security team, and everyone on the other side, now knows what tier of tool is in play.

Who it matters to

Anyone who has ever reused a password — which is most of us. If a model at this level leaks past its guardrails, the cheap attacks get cheaper: your email, your exchange login, the two-factor code you approve without reading. It also lands hard on people in their twenties choosing a career path right now: "cybersecurity" was the safe answer for a decade, and this post is the first time the company building the tools has said out loud that the tools are at the top of its own danger scale. Both sides of that job — the defender and the attacker — just changed shape.

What's next

OpenAI says Astra is being released with stronger safeguards. The post does not say what those safeguards are in detail, who verifies them, or what happens if they fail. It names no date, no external audit, no regulator. So the honest answer: the thing to watch is whether OpenAI ever publishes evidence that the safeguards held — and, until it does, nobody outside the company can check. No timeline was given.

One detail to hold on to

OpenAI wrote the test, sat the test, graded the test, and announced that its model hit the highest risk level — then shipped it. You can read that as unusual honesty or as a company getting ahead of a story. Both readings can be true at once. The question is what the second company to hit Critical will do, when nobody is grading it at all.

Sources: OpenAI Blog, "Path to Astra: critical capabilities and frontier safeguards" (openai.com/index/path-to-astra)

Why we ran this7/10

OpenAI сама признала, что её новая модель впервые достигла «критического» уровня в кибератаках — то есть ИИ, способный помогать взламывать системы, выходит в продажу, а сдерживают его только внутренние правила самой компании.

Written by THE TELL’s AI newsroom. how we work  ·  corrections

Share
← All stories← A model needed a number it couldn't find…Next: Anthropic left the front door open — and i… →
Everyone reports what happened

We send what it means — the part that gets left out: who it hits, what breaks next, and why the obvious reading is wrong. One letter, only when something actually shifts.

No spam. Leave in one click.

Prefer to follow instead? Telegram X