Some Google employees say Gemini 4’s strong benchmark results do not consistently carry over to their coding work, according to Bloomberg’s September 30 reporting. The report, citing people with direct access to the effort, adds an internal disagreement to the mixed outside results accompanying Argon’s launch.
Google disputes the characterization. It told Bloomberg that describing Gemini 4 as underperforming in coding would be inaccurate. Bloomberg also spoke to an employee familiar with model development who said internal testing supported its standing among the leading models and rejected claims that it struggles with messy coding tasks. Other employees believe rivals are advancing faster.
Bloomberg separately reports that Google abandoned Gemini 3.5 Pro, which it had planned to release in June. That is a reported cancellation of an earlier model, not a halt to Argon’s rollout.
The public evidence supports a narrower conclusion than either a clean victory or a failed model. Arena’s WebDev ranking places Argon eighth, behind leading Claude and GPT models. Google, meanwhile, describes specialized engineering successes, including a Rust video decoder it says Argon made 2.7 times faster than the previous Rust version.
Those are different kinds of coding work. The employee accounts raise a practical question about consistency, but do not establish a measured failure rate. Access still starts with selected cyber defenders; Google has not given a date for broader availability.


