OpenAI just released GPT-6, called Astra, and on raw intelligence it is on a different level from anything before it.
The benchmark everyone is quoting
It hit 99.9% on ARC-AGI, a test built specifically so that AI cannot prepare for it. For context on how fast this moved: six months ago the best model in the world scored 8% on that same test.
That is the number that will end up in every headline. It is genuinely remarkable. It is also not the part that will change your day.
The upgrade that actually matters
Computer use. AI controlling your computer is not new, but until now it has been slow, clumsy, and honestly close to unusable for real work. You would watch it fumble a form and take over yourself.
Astra is roughly twice as fast at finishing tasks, and it holds together on longer jobs. It fills out forms, updates spreadsheets, and can build an entire website and then test the thing itself. Sam Altman called it the best model in the world for computer use and coding.
Speed is what turns a demo into a tool. A thing that does your task in two minutes gets used. The same thing at six minutes gets abandoned.
The part worth sitting with
OpenAI's president said that when people look back and ask when AGI actually started, this might be the model they point to. Take that with the appropriate amount of salt, given who is saying it, but it is a notable thing to say out loud.
More concretely: this is the first model OpenAI has ever rated critical for cybersecurity. During testing it found two previously unknown security flaws in Google's software, on its own, without being pointed at them.
That cuts both ways, and everyone in the industry knows it. A model good enough to find real vulnerabilities unprompted is a model good enough to find them for the wrong reasons.
What to do with it
- Retry the computer-use tasks you gave up on. If you wrote off agent workflows six months ago, that judgment is out of date.
- Give it a full job, not a step. The gains show up on long multi-step work, not on single questions.
- Watch the first few runs. Faster and more capable still is not the same as correct.
Every model release gets called historic. This one is worth paying attention to for an unglamorous reason: it is fast enough that agents stop being a demo.
The benchmark number will get quoted everywhere. The computer use speed is the part that changes your week.
– Anir
If this is the kind of AI thinking you want working on your own business, apply to work with me. And to catch the next breakdown, join the newsletter.
Anir Suren