Claude Sonnet 5.5 nears Opus ahead of OpenAI’s DevDay
Claude Sonnet 5.5 nears Opus in key tests before OpenAI’s DevDay. Its input and output token rates are half Opus’s, while task costs vary by setting.

Anthropic launched Claude Sonnet 5.5 on September 28, putting its middle tier close to Opus 5.5 on several tests at half Opus’s input and output token prices, as The Rundown reported.
The launch adds pressure on OpenAI ahead of DevDay today, where the opening keynote is scheduled for 10:00 a.m. Pacific.
What the tests show
At maximum effort, Artificial Analysis gave Sonnet 5.5 a score of 56 on its Intelligence Index, behind only Opus 5.5 at 58. Its office results were similarly close. Sonnet nearly matched Opus on GDPval-AA and AA-Briefcase, and scored 71% against Opus’s 70% on AutomationBench-AA.
On Terminal-Bench 4.0, a coding evaluation, AA measured 64% for Sonnet, compared with roughly 60% for both Opus and GPT-6 Astra.
Those results are provisional. AA tested a prerelease deployment affected by a structured output bug and said it planned reruns.
Token prices and task costs
Anthropic prices Sonnet at $2 per million input tokens and $10 per million output tokens, compared with Opus’s $4 and $20. The company also claims output is at least 30% faster than Sonnet 5 and task costs can be up to 30% lower.
At low or medium effort, Anthropic reports selected benchmark wins over Sonnet 5’s best results at about a tenth of the task cost.
Actual bills also depend on how many tokens a model needs to finish the work. At maximum effort, AA measured about $7.60 per task, roughly 50% more than Sonnet 5. AA’s tests and Anthropic’s selected comparisons cover different settings and workloads.
Why it matters
Sentiment around Anthropic and OpenAI can shift quickly. Claude’s two September releases look like substantial wins for Anthropic heading into a DevDay that already carried high expectations.
Opus 5.5 arrived on September 22 with a 20% cut to its input and output token prices. Artificial Analysis ranked it first, with leading results on six of ten Intelligence Index tests. Sonnet now brings similar results on several tests to a model with lower token rates.
For developers building coding tools or office automation, that makes Sonnet worth testing on tasks such as bug fixes and document production. The payoff could be faster responses and lower bills. Anthropic gives Opus the edge on complex work that calls for judgment, so teams still have a reason to consider the flagship for more ambiguous tasks.
Effort settings become part of that buying decision. Teams should compare successful completion, time to finish, billed tokens, retries, and human review across settings before projecting savings. Extra retries or review could erase the benefit of lower token rates.
Anthropic says Sonnet’s new cyber safeguards can route requests deemed high risk to Sonnet 5. Teams handling security work should check how those fallbacks affect their applications’ results and behavior.
OpenAI now faces a stronger comparison on capability, useful work, and the cost of finishing tasks. DevDay’s announcements and AA’s planned reruns could change the comparison again.
Sources & further reading
- 01therundown.ai ↗
- 02OpenAI DevDay 2026 ↗
- 03Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index | Artificial Analysis ↗
- 04Introducing Claude Sonnet 5.5 \ Anthropic ↗
- 05Introducing Claude Opus 5.5 \ Anthropic ↗
- 06Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index | Artificial Analysis ↗
This story builds on reporting from The Rundown newsletter on September 29, 2026.