New independent results give Gemini 4 Argon a clearer place among the leading models: first in Arena’s text-response ranking, but level with GPT-6 Astra, and behind Claude Opus 5.5, on Artificial Analysis’s broader test suite.
Arena’s ranking reflects which responses people prefer. Argon leads with 1,525 points from 4,942 votes, ahead of Claude Opus 4.6 at 1,505. That preference does not carry over equally to building websites: Arena places Argon eighth in its separate WebDev ranking.
Artificial Analysis gives Argon 53 at high reasoning, its highest available setting, matching Astra at maximum effort. Opus 5.5 scores 58 at maximum effort with default fallback models. The index combines tests of knowledge, reasoning, coding and multi-step work; it is a different measure from Arena’s user votes.
Argon’s cost advantage over Astra depends on the launch discount. Artificial Analysis puts its weighted cost per test task at $1.99, versus Astra’s $3.26. At Google’s standard rates, the evaluator calculates $3.98, more than Astra. Argon generates about 62,000 output tokens per task, compared with Astra’s 27,000, so cheaper tokens do not translate directly into cheaper work.
GPT-6.1 Sol remains cheaper on this suite: $0.72 per task for a score of 52 at maximum effort. These are benchmark costs, not quotes for every workload. Argon still is not publicly available, and the introductory discount has no confirmed end date.


