Artificial Intelligence • • 5-8 minutes

Gemini 4 Argon: Google's Frontier Model That Writes 1 Million Tokens and Reaches Defenders First

Diego Cortés
Diego Cortés
Full Stack Developer & SEO Specialist
Share:
Gemini 4 Argon: Google's Frontier Model That Writes 1 Million Tokens and Reaches Defenders First
Image generated with AI

Google unveiled Gemini 4 Argon, its new frontier model: it can write up to one million tokens in a single response and starts at $2 per million input tokens. Today only cyber defenders in the Fairwind Program have access to it.

What Google Announced and What "Announced" Means Here

On Wednesday, September 30, 2026, Google DeepMind published Gemini 4 Argon under the banner of "our next era of frontier intelligence". The important word in that phrase is the last one: this is a limited-availability announcement, not an open commercial launch. Between the headline and what you can do today there is a gap worth measuring before you change anything in production.

September 30, Google DeepMind and the First Model of the Gemini 4 Generation

The announcement is signed by Koray Kavukcuoglu, senior vice president at Google DeepMind. The company presents it as the first model of the Gemini 4 generation — the starting point of a family, not an interim release of the previous line. The post dedicates an explicit section to strengthening safeguards before broad availability, and that detail explains almost everything that follows.

"Our Next Era of Frontier Intelligence": What Argon Is Built For

Google points to three specific uses: long-horizon software engineering, professional knowledge work in legal and finance, and cyber defense. "Long horizon" means tasks that chain many steps together and that a normal model has to break into dozens of calls. That is the clue to the technical headline.

What You Still Cannot Do: There Is No Public API Endpoint

As of October 1, 2026 there is no gemini-4-argon identifier you can call from the public API. The service's model page does not list it, and several write-ups report errors when trying to invoke it. Google says it will widen availability "as soon as possible", without a date. Any migration plan that depends on this model is, for now, a plan built on thin air.

The Change That Actually Matters: One Million Output Tokens

From 64,000 to 1,000,000 Tokens in a Single Trajectory

The central figure in the announcement is not intelligence, it is the production ceiling: Argon can generate up to one million output tokens in a single trajectory, against the 64,000 previous models in the family allowed. That is a factor of sixteen, and it changes the kind of job you can hand over in one go.

What a Long Output Unlocks (and What It Makes Expensive)

Within that ceiling you can deliver an entire module, a long report or a full migration in one pass, instead of chaining responses that lose the thread. The flip side is the bill: at almost every provider output costs several times more than input, and here the ratio is five to one. A million output tokens is not a marketing number; it is the most expensive line item in the workflow.

Input Context and Max Output: Why Aggregators Disagree

Do not conflate two numbers that sound alike. The context window is how much the model can read at once; the output ceiling is how much it can write in a single response. The announcement refers to the second. Third-party summaries contradict each other on the first, so only the figure Google publishes as the output limit is used here.

The Launch Numbers

Pricing: $2 and $10 per Million, Cached Input at 95% Off, and the Later List Rate

According to the announcement as reported by specialist coverage, Argon carries introductory pricing of $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%, which leaves a million cached tokens at roughly ten cents. Price aggregators also cite a later list rate of $4 and $20; that second number is not a Google promise and should be confirmed before you budget twelve months ahead.

Lined up, frontier pricing now looks like this:

  • Gemini 4 Argon: $2 input and $10 output per million during the introductory period.
  • GPT-6.1 Sol: $2 and $10 per million since its September 29 launch.
  • GPT-6 Astra: $10 and $50 per million, per its own announcement.

The interesting reading is not "it's cheap": it's that the frontier stopped being priced as a luxury. The price war between Anthropic and OpenAI had already settled that band; Google has just stepped into it.

Vals Index, DeepSWE v1.1, CWE-bench and Terminal-Bench: Where It Leads

According to aggregators and launch coverage, Argon takes first place on the Vals Index at 68.9%, leads DeepSWE v1.1 at 77.9%, ties for first on CWE-bench v1 at 68% for vulnerability remediation, and lifts Terminal-Bench 4.0 to 57.6%, against 19.0% for Gemini 3.8 Flash. On the Artificial Analysis intelligence index it lands around 53 points, level with GPT-6 Astra.

Where It Doesn't Win: FrontierSWE v2 Against GPT-6 Astra and Claude Opus 5.5

The picture is not a clean sweep. On FrontierSWE v2, according to one third-party analysis, Argon would place third, behind GPT-6 Astra and Claude Opus 5.5. A model leading on some suites and losing on others is normal; the anomaly would be believing a single table.

Who Measures What: Self-Computed Benchmarks and No Outside Verification

The model's own evaluation PDF acknowledges that the DeepSWE v1.1 results are self-computed with an in-house harness, and that competitor figures are taken from third-party leaderboards. Several analyses stress that no independent lab has verified those numbers. It is like a restaurant awarding itself its own stars: it may be true and still not be comparable. That is why this article always names who measured.

Tiered Access: Fairwind, Ultra and the Paid API

What the Fairwind Program Is and Who It Lets In

The Fairwind Program is a Google DeepMind initiative that gives trusted partners early access against AI-assisted cyber threats. Argon starts there: vetted defenders and internal teams, ahead of any commercial customer. The program launched on September 2, 2026, and on September 30 it was extended to this model.

Next: Google AI Ultra Subscribers and Paid Customers, With No Date

The stated order is defenders first, then paid API customers and Google AI Ultra subscribers. There is no calendar. For a team planning quarters, that makes Argon something to track, not a piece to design today's architecture around.

Frontier Safeguards: What Google Says Before Opening Up

Google says it is strengthening safeguards before broad availability and describes a model built to refuse harmful requests in cyber offense and in chemical, biological, radiological and nuclear risk, while preserving legitimate defensive use. The trade-off is plain: until that filter is proven, the door stays shut.

The 2026 Pattern: Frontier Models Now Ship With a Gatekeeper

Mythos 5.1 for Vetted Defenders Only, and the Astra OpenAI Never Shipped

Argon is not an isolated case. Anthropic released Claude Fable 5.1 openly while reserving Claude Mythos 5.1 for vetted access programs. OpenAI, for its part, cancelled the launch of GPT-6.1 Astra after finding safety and alignment problems in internal testing, as reported by the Wall Street Journal and other outlets. The result of the year is a frontier that opens in layers.

Why Restricted Access Is Also a Market Decision

Restricting access does more than protect: it orders demand. The provider chooses who stress-tests the expensive model first, collects real usage signals and keeps room to adjust price before general release. Being the most capable model says nothing about when you will be able to bill against it.

What to Do About It Today

Don't Plan a Migration Around a Model You Can't Call

With no public endpoint there is nothing to migrate. Before rewriting a workflow, check in the documentation that the identifier exists and responds; until then, treat Argon as an announcement.

Benchmark With Your Own Tasks and Your Own Bill, Not the Vendor's Table

An index point does not translate into your product. Take ten real tasks from your team, run them on the models you can actually call and record input tokens, output tokens and cost per completed task. That is where you see whether long output saves you money or doubles it.

A Provider Layer That Doesn't Tie Your Hands

The lesson from a year of staggered launches is architectural: isolate the model call behind your own layer, and keep the provider name as configuration rather than code. When Argon becomes available, switching should be one line.

What Is Still Unknown

There is no date for the paid API or for Ultra subscribers, no certainty that introductory pricing will hold, and no external verification of the published results. It also remains to be seen whether the one-million-token ceiling holds up on real tasks or only in lab tests.

Conclusion

Gemini 4 Argon is a serious announcement with minimal availability: an output ceiling sixteen times higher, pricing that sits inside the frontier band and a door that for now only Fairwind defenders walk through. The useful move today is not to rush a migration, but to understand what changes and to keep your stack ready for when it opens. If you want the run-up to this moment, OpenAI's September 29 DevDay and the Gemini security incident in a test complete the month's map.

Categories