New independent results strengthen GPT-6.1 Sol’s case as a cheaper alternative to Astra. On ARC-AGI-3, which tests agents learning unfamiliar interactive games, it makes a large jump over the previous Sol.
With both models at maximum reasoning effort in ARC’s Standard harness, Sol 6.1 scores 52.7%, up from Sol’s 4.6%. Astra still leads at 62.7%. These scores measure performance relative to human action efficiency, not the percentage of ordinary work an agent can complete.
The software around the model makes another big difference. The Standard setup lets it carry forward chosen notes; the Provider Adapter also preserves its internal reasoning state and compresses long conversations so it can reuse earlier work. With that adapter, at maximum effort, Sol 6.1 reaches 96.2% versus Astra’s 98.6%, at listed costs of $3,800 versus $17,300, about 78% less. Both evaluation setups are verified by ARC Prize.
Separately, Sol 6.1 enters third on Arena’s WebDev leaderboard, behind Opus 5.5 and Astra, with 1,264 votes. Users compare generated web apps, such as a chess game, and choose the better one. Its lead over fourth-place Fable 5.1 falls within overlapping uncertainty ranges, but the improvement over the previous Sol, now seventh, is clearer.
Sol’s input and output token rates are one-fifth of Astra’s on Arena’s table. That is not a measured fivefold saving per web app: longer answers and retries affect the bill. Still, the two evaluations add evidence that this cheaper model is catching up on both interactive reasoning and web-app building.


