The Rundown AI homepage
Artificial intelligence/News & analysis

Google unveils Gemini 4 Argon with strong scores and limited access

Google’s Gemini 4 Argon posts strong benchmark scores, but access is limited to vetted cyber teams. Broader testing will show how those gains hold up.

By The Rundown Editorial TeamReviewed by Kelly Pitts2 min read
Gemini 4 Argon looks to return Google to the frontier — newsletter story image
Image source: Google

Google unveiled Gemini 4 Argon on September 30, saying it beats GPT-6 Astra and Claude Opus 5.5 on 13 of 19 benchmarks in the company’s comparison. As The Rundown reported, the new frontier model is initially limited to select vetted cybersecurity teams.

Access runs through Google’s Fairwind Program for trusted cyber defenders. Google plans to expand access first to paid API customers and Google AI Ultra subscribers, but has not announced a date.

The scores and the price

Argon’s high setting took first place on Arena’s September 30 text leaderboard, based on 4,942 votes. Arena marked the result preliminary.

In Artificial Analysis’s evaluation, Argon scored 53 at high reasoning, matching GPT-6 Astra at its maximum reasoning setting. Gemini 3.1 Pro Preview scored 30 on the same index.

Google also reports 77.9% on DeepSWE v1.1, a benchmark for software engineering, alongside leading scores on tests of knowledge work, long documents, and reading charts and video.

The announced API rates start at $2 per million input tokens and $10 per million output tokens. After the introductory period, those rates rise to $4 and $20, respectively. Google has not said when the discount ends.

Why it matters

Argon’s scores put Google back in the frontier conversation. A broader rollout still has to show whether those gains translate into dependable work.

Google launched Gemini 3 on November 18, 2025 and said Gemini 3 Pro led LMArena at the time. Gemini 3.1 Pro followed on February 19, 2026. Artificial Analysis identifies Argon as Google’s first new proprietary model above the Flash line in more than seven months.

For developers choosing models for coding or document work, Argon now deserves a closer look. Trials should measure how often it completes their usual tasks, how much review the results need and how long the work takes. Those measures would help a team decide whether a switch improves its daily work.

Google describes employees working with Argon and substantial code migrations, along with auditing and review before production rollout. That gives the launch some practical support from inside the company. Broader customer trials would test whether those experiences carry over to other teams and tasks.

Price adds another test. Artificial Analysis estimates a cost of $1.99 per Intelligence Index task for Argon at its introductory rates, compared with $3.26 for Astra. At Argon’s standard rates, that estimate rises to $3.98.

Those figures describe AA’s evaluation workload. A team considering a switch should budget at standard rates and track the cost per result it can accept. The review needed to reach that point belongs in the calculation too.

Sources & further reading

This story builds on reporting from The Rundown newsletter on October 1, 2026.