Gemini 4 Argon shows a sharp improvement over Gemini 3.8 Flash on Artificial Analysis’s hallucination measure: 15% versus 55%, with both models set to high reasoning. But the accompanying accuracy results show what kind of improvement this is: Argon gets fewer questions fully right, too.
The test, AA-Omniscience, asks 6,000 factual questions across areas including law, health and software engineering, without browsing or other tools. It rewards correct answers and penalizes wrong guesses, while leaving abstentions unpenalized.
The denominator matters. That 15% counts outright incorrect answers as a share of responses that were incorrect, partial or not attempted. It does not mean Argon is 85% accurate. On the separate accuracy measure, Argon answers 50% of all questions correctly, compared with Flash’s 55%.
Together, the results point to a more cautious model: substantially fewer outright wrong answers, alongside more partial answers or abstentions, rather than a jump in factual recall. That can be valuable when a missing answer is better than a false one. It is not evidence that hallucinations have been eliminated in everyday use.


