The fkra blog
· 11 min

Gemini 4 Argon can write a million tokens, and you cannot call it yet

Google's first flagship since Gemini 3 wins most of its own benchmark table, comes last on the two agentic-coding tests, and goes to cyber defenders first.

Google announced Gemini 4 Argon on 30 September 2026, one day after OpenAI's DevDay. It is Google's first new flagship generation since Gemini 3 last November, it arrives in place of the Gemini 3.5 Pro that Google promised in May and never shipped, and until launch day most coverage called it Gemini 4 Pro, the name the leaks used. Google's own pages never say "Pro". A Google spokesperson told Reuters the model is "larger in size than Google's previous line of top-tier 'Pro' models."

You cannot use it yet. Argon is going first to a subset of the cyber-defence partners in Google's Fairwind Program, and Google says paid API customers and Google AI Ultra subscribers come next, with no date. On 1 October the model's page in the Gemini API docs still returned a 404.

I read the announcement, the DeepMind model page with its benchmark table, the five-page evaluation methodology, the Fairwind terms, and the independent numbers from Artificial Analysis and Arena, then compared them with the models Argon is priced against.

What Google shipped, and to whom

Gemini 4 Argon

Announced

30 September 2026, by Koray Kavukcuoglu, who took charge of Gemini development at Google DeepMind in August

Who has it now

Teams inside Google, and a set of Fairwind partners, who get it "without cyber guardrails"

Next in line

Paid API customers and Google AI Ultra subscribers. No dates

Price

$2 per million input tokens and $10 per million output as an introductory price, then $4 and $20. Cached input is 95% off

Output limit

1M tokens, up from 64K

Input window

Not stated by Google. Artificial Analysis and Arena list 1M

Not published

A model card, the knowledge cutoff, speed, the Frontier Safety Framework levels it reached, and when the introductory price ends

Fairwind launched on 2 September with Gemini 3.8 Flash Cyber and has more than 650 partners, and only some of them get Argon. The terms are strict: phishing-resistant MFA, access limited to internal security, incident-response and penetration-testing teams, logged employee use, background checks on applicants, and no resale. Governments, critical infrastructure operators and core technology platforms are first in line. Google says it is also taking part in the US government's voluntary process for pre-release access.

Google gives one concrete example of what a defender gets. Wiz, which Google bought for $32 billion in March, used Argon to find "a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide." Ars Technica noted that Google gave no specifics, and I found none anywhere else.

The table Google published

Google compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 18 tests. Argon wins 12 outright, ties one and loses five.

Sorted by Argon's margin over the best of the three rivals on each test. Rival scores are the providers' self-reported numbers; the margins are my arithmetic from Google's table.

Test

Gemini 4 Argon

Best rival

Rival

Gap, points

Harvey's Legal Agent Benchmark

19.6%

6.7%

Fable 5.1

+12.9

GraphWalks 256k to 1M (F1)

84.2

71.8

Astra

+12.4

AutomationBench

51.3%

42.5%

Opus 5.5

+8.8

Vals Finance Agent v2

65.4%

58.9%

Fable 5.1

+6.5

LVBench

91.7%

87.5%

Astra

+4.2

RiemannBench

76.0%

72.0%

Astra

+4.0

DeepSWE v1.1

77.9%

74.2%

Opus 5.5

+3.7

LABBench 2

88.8%

85.4%

Astra

+3.4

Vals Index

68.9%

67.0%

Opus 5.5

+1.9

Vibe Code Bench

91.9%

90.3%

Fable 5.1, Opus 5.5

+1.6

Agent's Last Exam

39.5%

38.2%

Opus 5.5

+1.3

GraphWalks up to 128k (F1)

99.7

98.7

Astra

+1.0

Chartography

71.6%

71.0%

Astra

+0.6

CWE-bench v1

68.0%

68.0%

Astra

0

OSWorld-2.0 (offline subset)

69.2%

72.6%

Astra

−3.4

PostTrainBench

45.3%

49.3%

Opus 5.5

−4.0

Terminal-bench 4.0

57.4%

66.4%

Opus 5.5

−9.0

Terminal-Bench Science 0.1

57.6%

68.1%

Astra

−10.5

FrontierSWE v2

55.0%

65.5%

Astra

−10.5

The big leads are in knowledge work and long context. Argon scores 19.6% on Harvey's legal agent benchmark against 6.7% for the next model, 84.2 F1 on graph traversal across 256k to 1M tokens against 71.8, and 51.3% on AutomationBench against 42.5. Five of the other wins are under two points, which is inside the run-to-run noise of most of these tests.

The losses are in agentic coding and science. Argon comes last of the four on FrontierSWE v2, 10.5 points behind Astra, and last on Terminal-bench 4.0, 9 points behind Opus 5.5. Those are the two tests closest to a coding agent's working day: a repository, a task, a terminal, and no one to finish it for you. The New Stack also pointed out that 19.6% on Harvey means Argon fully completes about one legal task in five. That is nearly three times the next model and still a low number.

The methodology PDF adds caveats that change how far you can lean on some rows:

  • Google computed Argon's DeepSWE and Terminal-bench scores itself. The rivals' scores come from leaderboards and system cards.

  • On LVBench, a long-video test, Argon watched at one frame per second, while GPT-6 Astra got 800 frames, Fable 5.1 got 300 and Opus 5.5 got 600, "due to API limitations".

  • Unless noted otherwise, every rival score is the provider's own number at its maximum reasoning setting.

The CWE-bench row needs one more note. That leaderboard pairs each model with its own harness, so Argon ran in Google's Antigravity, Astra in Codex and Claude in Claude Code. The row measures a model and its tooling together, and Grok 4.7 in opencode tied at 68% as well.

Google's full table, 18 tests

Benchmark

Gemini 4 Argon

GPT-6 Astra

Claude Fable 5.1

Claude Opus 5.5

Vals Index

68.9%

63.1%

65.8%

67.0%

AutomationBench

51.3%

41.4%

31.4%

42.5%

Vals Finance Agent v2

65.4%

53.5%

58.9%

58.6%

Harvey's Legal Agent Benchmark

19.6%

5.4%

6.7%

3.8%

DeepSWE v1.1

77.9%

74.1%

67.4%

74.2%

FrontierSWE v2

55.0%

65.5%

56.3%

62.3%

Vibe Code Bench

91.9%

89.6%

90.3%

90.3%

Terminal-bench 4.0

57.4%

58.2%

57.9%

66.4%

PostTrainBench

45.3%

44.3%

40.2%

49.3%

Terminal-Bench Science 0.1

57.6%

68.1%

52.6%

63.3%

LABBench 2

88.8%

85.4%

68.6%

73.1%

RiemannBench

76.0%

72.0%

65.6%

69.6%

GraphWalks up to 128k, BFS (F1)

99.7%

98.7%

91.4%

90.6%

GraphWalks 256k to 1M, BFS (F1)

84.2%

71.8%

65.0%

66.8%

Agent's Last Exam

39.5%

34.2%

-

38.2%

OSWorld-2.0 (offline subset)

69.2%

72.6%

-

-

Chartography

71.6%

71.0%

46.2%

66.3%

LVBench

91.7%

87.5%

79.7%

83.7%

CWE-bench v1

68.0%

68.0%

58.0%

67.0%

What the independent numbers say

Artificial Analysis ran Argon through its own index of ten evaluations. It scores 53, level with GPT-6 Astra and one point ahead of GPT-6.1 Sol, and behind Claude Opus 5.5 at 57.6 and Claude Sonnet 5.5 at 56. That is 23 points above Gemini 3.1 Pro Preview, Google's last model above the Flash class, and Artificial Analysis concluded that Google "is now back to being one of the top three labs in intelligence achieved."

Intelligence against cost per task
Artificial Analysis Intelligence Index v4.3.2 against the cost of running one index task at list price, fetched 1 October 2026. At the standard price, which Artificial Analysis estimates at $3.98 a task, Argon would sit just right of GPT-6 Astra.

Two of its measurements pull in opposite directions. Argon's hallucination rate on AA-Omniscience is 15%, the lowest of any model scoring 45 or more on the index, against 51% for Astra and 59% for Opus 5.5. Its accuracy on the same test is 50%, which is 13 points below Astra's 63%. My reading is that Argon declines to answer more often when it does not know. For legal and finance work, where a confident wrong answer costs more than a blank one, I would take that trade.

Arena put it first on the text leaderboard at 1525, 20 points clear, marked preliminary on 4,942 votes. On the WebDev board it is eighth, 139 points behind Claude Opus 5.5. Two independent boards and Google's own table agree about coding.

Bloomberg reported that some Google employees say the model "does less well when employees actually put it to work" and "struggles to handle certain coding tasks." Google called that inaccurate. The article is paywalled, and the quotes reached me through posts citing it.

A million output tokens

Maximum output per response, by Gemini model, thousands of tokens
From the Gemini API model pages; Argon from Google's announcement, which gives "1M" without an exact integer. Every model listed also accepts 1,048,576 input tokens in the API.

Every recent Gemini model, from 2.5 Pro through 3.8 Flash, stops at 65,536 output tokens. Argon can write a million. Google's argument is that "when the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go." Artificial Analysis ran it with a new API feature called Long Decode Continuation, which pauses a long response and resumes it over follow-up calls so that one answer can outlast a request timeout.

Google's internal examples are the kind of job that needs the room. Argon agents replaced 32,000 lines of SIMD code in libgav1, and the result runs 2.7 times faster than the existing Rust port with identical video output. Rust migrations are under way on code up to the 800,000-line Zircon kernel in Fuchsia, still being audited before anything reaches production.

The model uses the room. Across the Artificial Analysis index Argon averaged about 62,000 output tokens per task, against 27,000 for GPT-6 Astra. It still costs less per task, $1.99 against $3.26, because its token prices are lower, and Artificial Analysis says so directly: the advantage "is driven by lower token prices, rather than reduced token use." When the introductory price ends, the same task costs $3.98, about 1.2 times Astra.

Worked example: what the price change does to a real job

Take an agent job that reads 200,000 tokens and writes 60,000, close to Argon's average output on the index.

  • At the introductory price: 0.2M × $2 = $0.40 for input, plus 0.06M × $10 = $0.60 for output, so $1.00 a job.

  • At the standard price: $0.80 plus $1.20, so $2.00 a job.

At a thousand jobs a day that is $1,000 against $2,000. Output is 60% of the bill in both cases, so the number to watch in your logs is output tokens per call, and a verbose model gets more expensive faster than its price per token suggests.

Why cyber defenders get it first

Google's reason for the staging is the cyber capability itself. When it launched Fairwind it argued that early access gives defenders "a vital adaptation window to harden their systems before bad actors have a chance to exploit new capabilities." The cyber numbers are the largest jump in the announcement over Google's own previous model: 85.8% against 71.0% for 3.8 Flash Cyber on finding vulnerabilities in source code across 20 languages, and 70.9% against 58.2% on Wiz's penetration test, which works without source code.

Google also published its strongest safety result. On Gray Swan's indirect prompt-injection test, attackers succeeded against Argon 0.7% of the time over 15 attempts, the lowest of the 13 models on the chart. Claude Opus 5.5 and Fable 5.1 are at 1.0%, GPT-6 Astra at 8.5%.

Model

Gray Swan attack success rate, k=15 (lower is better)

Gemini 4 Argon

0.7%

Claude Opus 5.5

1.0%

Claude Fable 5.1

1.0%

Claude Opus 5

4.6%

Gemini 3.8 Flash

5.5%

GPT-6 Astra

8.5%

GPT-6 Sol

10.1%

Google did not publish the two documents that would let an outsider check the staging decision: a model card, and which levels of its own Frontier Safety Framework Argon reached. The model-card URL returned a 404 on 1 October. The announcement describes four safeguard areas in prose: refusals for misuse backed by monitoring of the model's internal activations, prompt-injection resistance, monitoring of chain-of-thought and actions that can stop execution, and sealed sandboxes for high-risk training runs. It also asks the rest of the industry to "preserve reasoning transparency."

The launch came a day after Sundar Pichai said Google had signed the White House Accord on Super Intelligence, and in the same week OpenAI said it would not release a planned GPT-6.1 Astra over safety concerns. A staged release is a defensible call for a model this good at finding vulnerabilities. Until the card exists, though, nobody outside Google can check the table or the safety claims.

The year it took

Google's newest model on the Artificial Analysis index, by release date
Best-scoring listed setting of each release. Artificial Analysis data, fetched 1 October 2026. Gemini 3.5 Pro, promised for June, never shipped.

At I/O on 19 May, Google said Gemini 3.5 Pro was already in internal use and would roll out "next month". It never shipped. Business Insider reported that it was postponed several times internally over poor performance, and the Financial Times reported three missed deadlines. Google itself has never said it was cancelled. What shipped instead was four Flash releases in under four months, each a few points better than the last, and a leadership change on 5 August that put Kavukcuoglu in charge of Gemini development and made Demis Hassabis chair of Google DeepMind.

The Gemini app passed one billion monthly users on 11 August, in the middle of the months when Google did not have a top model. Argon's 11.7-point step over 3.8 Flash is larger than the three Flash steps before it combined.

If you build in Arabic

Google published nothing about Argon's languages. There is no multilingual benchmark in the table or the methodology, no list of supported languages, and no statement about regional availability. The announcement exists in English and Canadian French. None of the Fairwind partners Google names is in this region.

So nothing above tells you how Argon handles Arabic. When it reaches the API we will run it on our own Arabic generation set before we believe any row of that table, the same fifty-examples rule this journal keeps coming back to.

What I think

On the work Google chose to measure, Argon puts Google back at the frontier: long documents, legal and finance agents, retrieval across a million tokens, and finding vulnerabilities. On the work most developers would try first, coding in a terminal against a real repository, Google's own table and both independent boards put it behind Anthropic and OpenAI. The price is the most competitive part of the launch, it is introductory, and nobody has said when it ends.

For now Argon is a model announced to the public and deployed to a few hundred security teams. I would not plan anything around it until it has an API page, a model card and a date. When it gets them, budget at $4 and $20, and measure it on your own work.

When it reaches the API

  • Budget at the standard $4 and $20, not the introductory $2 and $10

  • Log output tokens per call, since Argon writes long and output is the expensive side

  • Test a coding agent on your own repository before switching it over

  • Run your own Arabic set, because Google published no multilingual numbers

  • Pin the model version once your thresholds are tuned

  • Read the model card when it appears, including the safety framework levels

0
15 views
AILLMsGeminiGoogleBenchmarksCybersecurity
MA
Mosab Alrasheed
Get the next entry

One email when we publish. Research, product decisions, and what teams report back.