Onset
The whole story

Update

Argon leads Arena chat, ties Astra in broader tests

New independent results give Gemini 4 Argon a clearer place among the leading models: first in Arena’s text-response ranking, but level with GPT-6 Astra, and behind Claude Opus 5.5, on Artificial Analysis’s broader test suite.

Arena’s ranking reflects which responses people prefer. Argon leads with 1,525 points from 4,942 votes, ahead of Claude Opus 4.6 at 1,505. That preference does not carry over equally to building websites: Arena places Argon eighth in its separate WebDev ranking.

Artificial Analysis gives Argon 53 at high reasoning, its highest available setting, matching Astra at maximum effort. Opus 5.5 scores 58 at maximum effort with default fallback models. The index combines tests of knowledge, reasoning, coding and multi-step work; it is a different measure from Arena’s user votes.

Argon’s cost advantage over Astra depends on the launch discount. Artificial Analysis puts its weighted cost per test task at $1.99, versus Astra’s $3.26. At Google’s standard rates, the evaluator calculates $3.98, more than Astra. Argon generates about 62,000 output tokens per task, compared with Astra’s 27,000, so cheaper tokens do not translate directly into cheaper work.

GPT-6.1 Sol remains cheaper on this suite: $0.72 per task for a score of 52 at maximum effort. These are benchmark costs, not quotes for every workload. Argon still is not publicly available, and the introductory discount has no confirmed end date.

Artificial Analysis’s weighted task costs, including reasoning and cached input. Argon uses high reasoning; Astra uses max. Argon’s standard-price figure is calculated, not a price increase already in effect.
More from this story

Update

Bloomberg reports internal doubts about Gemini 4’s coding; Google disputes them

Some Google employees say Gemini 4’s strong benchmark results do not consistently carry over to their coding work, according to Bloomberg’s September 30 reporting. The report, citing people with direct access to the effort, adds an internal disagreement to the mixed outside results accompanying Argon’s launch.

Google disputes the characterization. It told Bloomberg that describing Gemini 4 as underperforming in coding would be inaccurate. Bloomberg also spoke to an employee familiar with model development who said internal testing supported its standing among the leading models and rejected claims that it struggles with messy coding tasks. Other employees believe rivals are advancing faster.

Bloomberg separately reports that Google abandoned Gemini 3.5 Pro, which it had planned to release in June. That is a reported cancellation of an earlier model, not a halt to Argon’s rollout.

The public evidence supports a narrower conclusion than either a clean victory or a failed model. Arena’s WebDev ranking places Argon eighth, behind leading Claude and GPT models. Google, meanwhile, describes specialized engineering successes, including a Rust video decoder it says Argon made 2.7 times faster than the previous Rust version.

Those are different kinds of coding work. The employee accounts raise a practical question about consistency, but do not establish a measured failure rate. Access still starts with selected cyber defenders; Google has not given a date for broader availability.

Update

Argon leads on taking corrections in early Agent Arena results

Gemini 4 Argon ranks eighth overall but first at responding to user corrections in Arena’s new results from real agent sessions. At high reasoning, it places ahead of Gemini 3.8 Flash in nineteenth, with a median task cost of $0.62 versus Flash’s $0.58.

Unlike Arena’s head-to-head chat votes, Agent Arena randomly assigns models to handle users’ tasks, then measures feedback and tool behavior. Its steerability signal checks whether a user accepts the agent’s fix after correcting it. Argon also ranks second on explicit user-confirmed success.

The overall score combines those signals with praise versus complaints, recovery from failed commands and attempts to call nonexistent tools. Argon’s +7.92% net improvement is an estimated gain over Arena’s baseline, not the percentage of jobs it completes. Claude Fable 5.1 at maximum effort leads the overall ranking.

The result is preliminary: Argon has 3,417 sessions, and its overall score carries a 95% confidence interval of ±3.24 percentage points. Its $0.62 median uses Google’s introductory $2/$10 input/output rates per million tokens, so it is not a permanent-price estimate.

The useful early signal is about collaboration: Argon looks particularly good at getting back on track when a person intervenes. That is different from demonstrating reliable, unattended work.

Arena’s cost-performance chart places Argon among the best trade-offs at its introductory price. Costs are median per task; cheaper models sit farther right. Source: Arena.

Update

Argon makes fewer wrong guesses on a factual-knowledge test

Gemini 4 Argon shows a sharp improvement over Gemini 3.8 Flash on Artificial Analysis’s hallucination measure: 15% versus 55%, with both models set to high reasoning. But the accompanying accuracy results show what kind of improvement this is: Argon gets fewer questions fully right, too.

The test, AA-Omniscience, asks 6,000 factual questions across areas including law, health and software engineering, without browsing or other tools. It rewards correct answers and penalizes wrong guesses, while leaving abstentions unpenalized.

The denominator matters. That 15% counts outright incorrect answers as a share of responses that were incorrect, partial or not attempted. It does not mean Argon is 85% accurate. On the separate accuracy measure, Argon answers 50% of all questions correctly, compared with Flash’s 55%.

Together, the results point to a more cautious model: substantially fewer outright wrong answers, alongside more partial answers or abstentions, rather than a jump in factual recall. That can be valuable when a missing answer is better than a false one. It is not evidence that hallucinations have been eliminated in everyday use.

Argon’s hallucination rate is much lower than Flash’s. This measures wrong answers among non-correct responses, not errors across all answers. Source: Artificial Analysis.