RADIO’S REPORT · AGENTIC AI · WATCH IN 60 SECONDS
Anthropic dropped Opus 5.5 last week and Sonnet 5.5 yesterday. The cheaper model lands two points behind the flagship on real-world work, at half the price. So Radio grabbed the mic and took the story desk by desk.
⚡ THE 20-SECOND VERSION
- 1844 vs 1846. Sonnet 5.5 vs Opus 5.5 on GDPval-AA, a benchmark built from real professional tasks.
- Half the price. $2 in / $10 out per million tokens, against Opus 5.5’s $4 / $20.
- Coding went nuclear. Terminal-Bench 4.0 jumped from 10.3% to 70.6% in one generation.
- It can drive a computer. OSWorld 2.1 climbed from 57% to 80.1%.
- The move: everyday work on Sonnet 5.5, the hard judgment calls on Opus 5.5.

Every few months some AI release gets called a game changer. This one brought receipts. On GDPval-AA, Anthropic’s new Claude Sonnet 5.5 scored 1844. Opus 5.5, the company’s top model, released last week, scored 1846.
Two points. That’s the gap between the mid-priced model and the flagship. And the mid-priced one costs half as much.

The scoreboard
| What changed | Before | Claude 5.5 |
|---|---|---|
| Real-world work (GDPval-AA) | Sonnet 5: 1449 | Sonnet 5.5: 1844 · Opus 5.5: 1846 |
| Coding (Terminal-Bench 4.0) | Sonnet 5: 10.3% | Sonnet 5.5: 70.6% |
| Using a computer (OSWorld 2.1) | Sonnet 5: 57% | Sonnet 5.5: 80.1% |
| Speed | Sonnet 5 | 30%+ faster output |
| API price, per million tokens | Opus 5.5: $4 in / $20 out | Sonnet 5.5: $2 in / $10 out |
| Opus, generation over generation | Opus 5 | Opus 5.5: 20% cheaper per token, about 30% faster and 40% cheaper per task |

And it’s not just lab scores. Anthropic says early customers saw real gains: Box processed work 2.4× faster with 12% fewer tokens, Zendesk handled tickets 20% faster, and Lovable needed about a third fewer tool calls on coding tasks. Those are customer results as reported by Anthropic, not our own tests.
Smarter, faster and cheaper, all at once. Radio took the story around the newsroom.
💜 Bella’s desk: media production

“For media teams, that’s scripts, storyboards and edit notes at half the cost. Same quality. More time for the creative.” — Bella
Most of a production is words long before it’s pictures: treatments, scripts, shot lists, edit notes, captions, a version for every platform. That everyday writing is exactly where Sonnet 5.5 now sits two points off the flagship. Same budget, roughly twice the drafting.

🩵 Kai’s desk: VR and AR

“Coding scores jumped from ten percent to seventy. That means agentic AI that can really help build the virtual worlds we explore.” — Kai
Virtual worlds are made of code. On Terminal-Bench 4.0, a test of real command-line coding work, Sonnet 5.5 went from 10.3% to 70.6% in a single generation. Agentic AI that can write, run and fix its own code is how small teams start shipping the scenes, filters and AR experiences Kai covers.

🩷 Peaches’ desk: the smart home

“And it can use a computer. Eighty percent on OSWorld. So your smart-home AI can finally run the apps, not just the lights. And it costs less.” — Peaches
Your “smart” home is really a pile of apps: lights in one, the thermostat in another, then the camera, the vacuum and the shades. OSWorld 2.1 measures how well an AI can operate a computer on its own, and Sonnet 5.5 jumped from 57% to 80.1%. An assistant that can actually use those apps gets a lot more useful. And cheaper models are what make an always-on assistant affordable to run.

💰 The part everyone feels: the bill

Which Claude should do the job?
| The job | Send it to |
|---|---|
| Email, research, content, social posts | Sonnet 5.5 |
| Routine code, bug fixes, support tickets | Sonnet 5.5 |
| Long, messy projects that need sustained judgment | Opus 5.5 (Anthropic says it stays clearly stronger here) |
| Quick lookups and tiny edits | A smaller, cheaper model |
Do this week: pull up your AI spend. If you’re paying flagship prices for everyday tasks, you’re overpaying.


The bottom line
Opus 5.5 last week. Sonnet 5.5 this week. Anthropic’s 5.5 family is the biggest price-to-performance jump we’ve covered this year. Smarter. Faster. Cheaper. That’s the 5.5 story.
Which desk gets the most out of it: Bella’s edit bay, Kai’s virtual worlds, or Peaches’ smart home? Tell us on our socials, and share the video with the one friend still paying flagship prices.

Sources: Anthropic: Introducing Claude Sonnet 5.5 · VentureBeat · Claude blog: What a task costs on Opus 5.5. Benchmarks and customer results are as published by Anthropic; Zip AI did not run independent tests.
Disclosure: Radio, Bella, Kai, Peaches and Emily are AI-generated characters with synthetic voices. All images are AI-generated illustrations; the home technology shown is concept imagery. Editorial oversight is human. Independent coverage with no sponsorship from Anthropic.
