Onset
The whole story

Update

Early testers find low-effort gains, but not on every coding task

Sonnet 5.5’s useful gains aren’t confined to its expensive maximum-effort setting. Every says it built a working clone of the publication’s open-source document editor, Proof, at low effort in Kieran Klaassen’s testing. Opus 5 had failed that test; Opus 5.5 succeeded and produced a stronger result. Every cautions that a working clone is not a production-ready application.

But the upgrade did not win across the board. In Klaassen’s separate 15-task coding test, Sonnet 5.5 passed nine tasks at each of low, medium and high effort. Sonnet 5 passed ten at low. That small test does not overturn the broader benchmark gains, but it is a reason to test an upgrade on your own work rather than assume more reasoning will help. These results come from Every’s public test summaries; the full review is subscriber-only.

Anthropic’s developer guide likewise recommends starting at medium effort for well-specified coding and multi-step tool tasks, raising it for harder work. Effort controls how much reasoning the model does, and Anthropic says the levels have been recalibrated: an old Sonnet 5 setting will not produce the same amount of thinking in Sonnet 5.5. It recommends extra-high or max only where your own evaluations show a quality gain.

More from this story

Update

Sonnet 5.5 reaches Claude’s free tier

The stronger model is now what people get on Claude’s free plan, Simon Willison notes. That matters more widely than its API price: anyone can try the upgrade without subscribing, though Anthropic’s free-plan message limit still applies and resets every five hours. Claude’s previous free default was Sonnet 5.

Anthropic’s new side-by-side demo shows the difference on one creative coding task: asked for a fall-foliage simulator, Sonnet 5 makes a sparse tree in a plain field; Sonnet 5.5 adds a landscaped scene and a timeline that sheds leaves. It’s a handpicked example, not evidence that every free user will get a finished app in one prompt.

@claudeai · Watch on 𝕏