The stronger model is now what people get on Claude’s free plan, Simon Willison notes. That matters more widely than its API price: anyone can try the upgrade without subscribing, though Anthropic’s free-plan message limit still applies and resets every five hours. Claude’s previous free default was Sonnet 5.
Anthropic’s new side-by-side demo shows the difference on one creative coding task: asked for a fall-foliage simulator, Sonnet 5 makes a sparse tree in a plain field; Sonnet 5.5 adds a landscaped scene and a timeline that sheds leaves. It’s a handpicked example, not evidence that every free user will get a finished app in one prompt.
Early testers find low-effort gains, but not on every coding task
Sonnet 5.5’s useful gains aren’t confined to its expensive maximum-effort setting. Every says it built a working clone of the publication’s open-source document editor, Proof, at low effort in Kieran Klaassen’s testing. Opus 5 had failed that test; Opus 5.5 succeeded and produced a stronger result. Every cautions that a working clone is not a production-ready application.
But the upgrade did not win across the board. In Klaassen’s separate 15-task coding test, Sonnet 5.5 passed nine tasks at each of low, medium and high effort. Sonnet 5 passed ten at low. That small test does not overturn the broader benchmark gains, but it is a reason to test an upgrade on your own work rather than assume more reasoning will help. These results come from Every’s public test summaries; the full review is subscriber-only.
Anthropic’s developer guide likewise recommends starting at medium effort for well-specified coding and multi-step tool tasks, raising it for harder work. Effort controls how much reasoning the model does, and Anthropic says the levels have been recalibrated: an old Sonnet 5 setting will not produce the same amount of thinking in Sonnet 5.5. It recommends extra-high or max only where your own evaluations show a quality gain.