OpenAI just made its biggest move of the year. On July 9 they launched ChatGPT Work, an agent that lives on your computer and works like an actual employee, and paired it with GPT-5.6, a new family of models that undercuts Claude's flagship at half the price. Here's everything that shipped, in plain English, and my honest take on the ChatGPT vs Claude question at the end.
What ChatGPT Work actually is
Think of it as the difference between an assistant you talk to and an employee you delegate to. ChatGPT Work runs on your computer, opens your local files and apps, and executes whole workflows in the background while you do something else. You can schedule tasks so it works without you, and it has its own built-in browser, so it can navigate websites, use online tools, and pull from files on the web the way a person would.
The output isn't just chat. From a single prompt it produces finished work: complex spreadsheets, product mockups and images, and full slide decks. And once a task is running on your desktop, you can monitor and steer the whole thing from your phone.
Plugins and Sites
Two pieces make it feel like a real coworker instead of a chatbot. Plugins connect the systems you already work in, so ChatGPT Work can act inside your actual stack instead of asking you to copy things in and out. And a new beta called Sites lets it build interactive web apps for you: live dashboards, reports, internal tools, made and hosted from a prompt.
The new desktop app (and where Codex went)
All of this ships in a brand new ChatGPT desktop app for Mac and Windows. The old app is being renamed ChatGPT Classic, and OpenAI is merging the standalone Codex app into the new one, so developers keep the coding features (inline diff editing, PR review, multi-repo support) inside the same app that now runs their spreadsheets and decks. One app, one agent, everything on your machine.
GPT-5.6: the engine underneath
The whole thing is powered by GPT-5.6, which also launched into general availability the same day. It comes in three sizes:
- Sol, the flagship, at $5 in / $30 out per million tokens
- Terra, the balanced everyday model, at $2.50 / $15
- Luna, the fast cheap one, at $1 / $6
The number that matters: Sol costs half of what Anthropic charges for Claude Fable 5 ($10 / $50). And on agent coding tests it actually pulls ahead, scoring 88.8% on Terminal-Bench 2.1 against Fable 5's 83.4%. Worth knowing: OpenAI hasn't published a Sol score on SWE-bench Pro, the brutal benchmark where Fable 5 holds the crown at 80%, so the full picture isn't in yet.
So did ChatGPT just kill Claude?
Honest answer: it's the closest race AI has ever had, and anyone giving you a clean winner is selling something. Here's my read. OpenAI just won on price and reach: GPT-5.6 is generally available to everyone today at half Claude's flagship cost, and ChatGPT Work does on your computer what Claude Cowork has been doing, in a package hundreds of millions of people already use. Anthropic still holds the capability high ground: Fable 5 is the strongest model on the hardest published benchmark, and Claude Code remains the tool serious builders reach for.
Both companies have shipped generational stuff in the span of a month. The smart move hasn't changed: don't pick a team, pick the best tool for the job in front of you. This week, for agent work on a budget, that answer just got a lot more interesting.
How to try it
- Go to openai.com/chatgpt-work and download the new desktop app for Mac or Windows.
- Sign in, connect the tools you use through Plugins, and give it a real multi-step task, not a question. Something like "build me a spreadsheet comparing our last 3 months of sales and turn it into a 5-slide deck."
- Then leave it alone and watch it work. That part still feels illegal.
For two years the argument was about which AI writes better answers. That era just ended. The new argument is about which AI does better work, and both sides are shipping like the company depends on it. Because it does.
Download it, give it one real task from your actual job, and judge it on the finished work. That test tells you more than every benchmark chart combined.
– Anir
If this is the kind of AI thinking you want working on your own business, apply to work with me. And to catch the next breakdown, join the newsletter.
Anir Suren