Gemini 4 Argon: Inside Google’s 1M-Token Output Model, Benchmarks, Pricing and Release Date

Google’s newest frontier model arrives with an unusual headline: Gemini 4 Argon can produce up to one million output tokens in a single trajectory. That is not just a bigger chatbot reply. Google is positioning the model for jobs that may involve hours of reasoning, tool use, code changes, document analysis and repeated decisions before the task is finished.

The early numbers are equally ambitious. Gemini 4 Argon leads the comparison set on long-context reasoning, several professional-work evaluations, scientific benchmarks and multimodal tests. But it does not sweep coding. GPT-6 Astra beats it on FrontierSWE and Terminal-Bench Science, while Claude Opus 5.5 has a sizeable lead on Terminal-Bench 4.0 and PostTrainBench.

There is another catch. Argon was announced on September 30, 2026, but it is not yet a generally available Gemini API model. Google is starting with trusted cybersecurity testers through its Fairwind Program, then plans to expand access to developers, enterprises and consumers.

That makes this less of a conventional model launch and more of a preview of where frontier AI is heading: longer autonomous workflows, deeper reasoning budgets and models expected to finish substantial jobs rather than merely answer prompts.

1. Gemini 4 Argon At A Glance

Gemini 4 Argon: Key Specs, Pricing and Availability

FeatureGemini 4 Argon
AnnouncedSeptember 30, 2026
Current release statusLimited rollout through the Fairwind Program
Broader availabilityPlanned for developers, enterprises and consumers
First broader-access groupsPaid API customers and Google AI Ultra subscribers
Introductory API price$2 / 1M input, $10 / 1M output
Cached input95% below normal input price during introductory pricing
Reported standard price$4 / 1M input, $20 / 1M output
Maximum outputUp to 1M tokens in a trajectory
Main focusLong-horizon reasoning, coding, professional knowledge work, cybersecurity
Standout results77.9% DeepSWE, 51.3% AutomationBench, 84.2% long-context GraphWalks, 91.7% LVBench
Public Gemini 4 Argon API todayNot yet broadly available

Google describes Argon as a model built to sustain reasoning across complex, long-running workflows, particularly software engineering, finance, legal work and defensive cybersecurity. Introducing Gemini 4 Argon

2. What Is Gemini 4 Argon?

Gemini 4 Argon infographic showing a looping long-horizon agent workflow versus a single short answer
Gemini 4 Argon infographic showing a looping long-horizon agent workflow versus a single short answer

Gemini 4 Argon is Google DeepMind’s new frontier model, but the interesting change is not simply another increase in benchmark scores.

Its design target is the long-horizon workflow.

A normal model interaction might involve reading a prompt, reasoning briefly and producing an answer. A long-horizon agent may need to inspect a repository, run tools, interpret failures, modify files, test the changes, revisit assumptions and continue that loop many times.

Google says Argon combines coding, reasoning and multimodal capabilities with the ability to sustain these multi-step tasks. Internally, Google engineers are already using it for work ranging from everyday debugging to major code migrations and algorithm design. Introducing Gemini 4 Argon

That distinction matters. The frontier-model race is increasingly becoming less about who gives the smartest single response and more about which model stays useful after step 50, step 500 or step 5,000.

3. Gemini 4 Argon Release Date And API Availability

The Gemini 4 Argon release date is September 30, 2026, but “released” needs qualification.

The first rollout is through Google’s Fairwind Program, with access going to a limited group of trusted cyber defenders. Google says it is also participating in the U.S. government’s voluntary process for pre-release model access while it gathers feedback and strengthens safeguards.

There is currently no confirmed date for general Gemini 4 Argon API availability.

Google says the broader rollout will eventually cover developers, enterprises and consumers, beginning with paid API customers and Google AI Ultra subscribers. Introducing Gemini 4 Argon

So developers searching for a model ID they can drop into production today will have to wait. The launch announcement describes pricing, but access remains phased.

4. Gemini 4 Argon Pricing Compared With GPT-6 Astra And Claude Opus 5.5

Gemini 4 Argon pricing is aggressive, particularly during its introductory period.

Gemini 4 Argon Pricing Compared With GPT-6.1 Sol, Claude Opus 5.5 and GPT-6 Astra

ModelInput / 1M TokensCached Input / 1MOutput / 1M Tokens
Gemini 4 Argon, introductory$2$0.10$10
Gemini 4 Argon, standard$4Not yet directly specified by Google$20
GPT-6.1 Sol$2$0.10$10
Claude Opus 5.5$4$0.20$20
GPT-6 Astra$10$1$50

Google directly confirms the introductory $2 input and $10 output rate, plus a 95% discount for cached input. Introducing Gemini 4 Argon Venturebeat reports that Google plans to raise Argon to $4 input and $20 output after the introductory period, although Google has not announced when that period ends. Venturebeat

For comparison, OpenAI lists GPT-6.1 Sol at $2/$10 and GPT-6 Astra at $10/$50, while Anthropic prices Claude Opus 5.5 at $4/$20.

Per-token pricing, however, is only half the bill.

If Model A costs twice as much per token but solves a task in one attempt while Model B burns through four retries, Model A can still be cheaper. Agentic workloads magnify that difference because reasoning tokens, tool calls, repeated context and failed trajectories all cost money.

For builders, cost per successfully completed task will matter more than the sticker price.

5. Gemini 4 Argon’s 1M Output Tokens Explained

Gemini 4 Argon infographic comparing a short context-window input with a one-million-token output budget
Gemini 4 Argon infographic comparing a short context-window input with a one-million-token output budget

The easiest detail to misunderstand about Argon is also its most unusual.

A 1M-token output limit is not the same thing as a 1M-token context window.

Google says Argon’s maximum output has increased from 64K to one million tokens, giving the model enough generation headroom to reason and produce hundreds of thousands of tokens during a single trajectory. Introducing Gemini 4 Argon

Think of context as what the model can have available to read and reason over. Output capacity is how much it can generate while working through the task.

Google’s separate model evaluation does show strong long-context behavior. Its GraphWalks testing includes 200 problems with context lengths between 256K and 1M tokens, on which Argon scores 84.2% F1.

That provides evidence for operation over very large contexts, but it should not be used to casually turn “1M output tokens” into a claimed production “1M context window.” They are different specifications.

The more interesting question is whether applications can use that huge output budget productively. Most users certainly do not need a million-token answer. Agents performing thousands of intermediate reasoning and tool-use steps might.

6. Gemini 4 Argon Benchmarks: Full Results

Here is Google’s complete published comparison. The evaluation spans knowledge work, coding, ML engineering, science, long context, computer use, multimodal understanding and cybersecurity.

DomainBenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Knowledge workVals Index68.9%63.1%65.8%67.0%
Knowledge workAutomationBench51.3%41.4%31.4%42.5%
Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%
Knowledge workHarvey’s Legal Agent Benchmark19.6%5.4%6.7%3.8%
Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%
Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%
Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%
Agentic codingTerminal-Bench 4.057.4%58.2%57.9%66.4%
ML engineeringPostTrainBench45.3%44.3%40.2%49.3%
Science and mathTerminal-Bench Science 0.157.6%68.1%52.6%63.3%
Science and mathLABBench 288.8%85.4%68.6%73.1%
Science and mathRiemannBench76.0%72.0%65.6%69.6%
Long contextGraphWalks, up to 128K99.7%98.7%91.4%90.6%
Long contextGraphWalks, 256K to 1M84.2%71.8%65.0%66.8%
Computer useAgent’s Last Exam39.5%34.2%N/A38.2%
Computer useOSWorld 2.0, offline partial69.2%72.6%N/AN/A
MultimodalChartography71.6%71.0%46.2%66.3%
MultimodalLVBench91.7%87.5%79.7%83.7%
CybersecurityCWE-bench v168.0%68.0%58.0%67.0%

The pattern is more useful than counting bold cells. Argon looks especially strong where a task requires maintaining information and making decisions across a long chain of work.

6.1 Is Gemini 4 Argon Actually Better At Coding?

Not universally.

DeepSWE is impressive at 77.9%, and Argon also leads Vibe Code Bench at 91.9%. But FrontierSWE tells a very different story: GPT-6 Astra scores 65.5%, Claude Opus 5.5 scores 62.3%, and Argon falls to 55.0%.

Terminal-Bench 4.0 is even more revealing. Argon reaches 57.4%, while Claude Opus 5.5 scores 66.4%.

PostTrainBench also goes to Opus, at 49.3% versus Argon’s 45.3%.

So the useful interpretation is not “Argon wins coding.” Different coding benchmarks stress repository navigation, terminal execution, patching, engineering judgment and agent harnesses differently.

The best model for your codebase may depend heavily on what you actually ask it to do.

6.2 Where Argon Looks Most Distinctive

The long-context result stands out. Argon scores 84.2% on GraphWalks between 256K and 1M tokens, 12.4 percentage points above Astra.

Professional workflows are another strong area. Argon posts 51.3% on AutomationBench, 65.4% on Vals Finance Agent v2 and 19.6% on Harvey’s Legal Agent Benchmark. Google also reports 91.7% on LVBench for long-video understanding.

Science is mixed but strong overall. Argon leads LABBench 2 and RiemannBench, while Astra has a clear advantage on Terminal-Bench Science.

One warning applies to all of these numbers: being first on a leaderboard does not mean the underlying task is solved. A 39.5% pass rate can lead Agent’s Last Exam while still failing most attempts.

7. Gemini 4 Argon Cybersecurity And Fairwind

Cybersecurity is not merely another benchmark category in this launch. It is central to the rollout strategy.

Google says Argon can autonomously find, validate and patch software vulnerabilities. Trusted defenders and internal Google teams are receiving access through Fairwind, while Wiz is testing the model through its Scan for Good program. Introducing Gemini 4 Argon

On CWE-bench v1, Argon scores 68% and ties GPT-6 Astra, rather than winning outright.

Google also reports stronger vulnerability discovery than Gemini 3.8 Flash Cyber on internal evaluations, including testing across codebases covering 20 programming languages and a Wiz black-box penetration-testing benchmark. Introducing Gemini 4 Argon

The company is simultaneously working on misuse safeguards, prompt-injection resistance, agent monitoring and hardened sandbox environments before wider access.

8. What Gemini 4 Argon Is Already Doing Inside Google

The most interesting evidence in Google’s announcement may not be the benchmark table at all.

Google describes Argon agents that analyzed data-center telemetry and identified memory optimizations which have already freed more than 300 TiB of memory, with estimated potential savings of 500 TiB to 1 PiB.

Other internal projects include C and C++ to Rust migrations ranging from tens of thousands of lines to more than 800,000 lines for the Fuchsia Zircon kernel.

On libgav1, Argon agents replaced 32,000 lines of SIMD code in an existing Rust port. Google says the resulting decoder is 2.7 times faster while producing identical video output. Another example involved quantum-computing optimization, where Argon reportedly improved a published baseline by 40%.

These examples answer an important question: can Argon do more than benchmark puzzles?

Inside Google, apparently yes.

But there is still an important boundary. These projects demonstrate carefully engineered agent workflows within Google. They do not prove that an ordinary Gemini 4 Argon API call will automatically reproduce the same autonomy, tooling or reliability.

9. How Much Should You Trust The Benchmark Claims?

Google provides more methodology detail than a benchmark graphic alone suggests.

Argon results are generally pass@1, with the model run through the Gemini API using its highest thinking settings. Google also says many non-Gemini numbers come from competitors’ published results or public leaderboards.

The evaluation is not one perfectly controlled tournament.

DeepSWE and Terminal-Bench results for Argon are self-computed. PostTrainBench and LABBench were run across all models by Google. GraphWalks was also self-computed, while several other scores come from official third-party leaderboards. gemini_4_argon_model_evaluation gemini_4_argon_model_evaluation

Even multimodal conditions differ. For LVBench, Google used different frame counts across Gemini, GPT-6 Astra and the Claude models because of API limitations. gemini_4_argon_model_evaluation

None of that makes the results meaningless. It means the table should be treated as strong launch evidence, not the final independent verdict.

Public reproducibility is the missing piece.

10. What Developers Should Test When Argon Opens Up

When broader Gemini 4 Argon API access arrives, reproducing Google’s leaderboard should not be the first priority.

Test the workloads you actually pay for.

Give it a real repository and measure completed issues, not attractive patches. Test how often a 30-minute or three-hour agent run recovers from mistakes. Feed it hundreds of thousands of tokens and check whether important details remain usable near the end. Track retries, tool calls, output tokens, latency and human review time.

Then calculate cost per accepted result.

That comparison will be particularly interesting against GPT-6.1 Sol, which matches Argon’s introductory $2/$10 token pricing, and Claude Opus 5.5, which matches Argon’s reported eventual $4/$20 rate.

The model that charges the fewest dollars per million tokens is not automatically the model that costs the least to get real work finished.

11. Gemini 4 Argon Is Really A Bet On Longer AI Work

Gemini 4 Argon has strong benchmark numbers, but its bigger idea is the move from short model interactions toward persistent computational work.

The one-million-token output ceiling gives agents far more room to keep going. GraphWalks suggests Argon can maintain useful reasoning across extremely large contexts. Its internal Google deployments show the kind of projects that become possible when those abilities are wrapped in tools, testing and engineering infrastructure.

There are still reasons to wait before declaring the race settled. Argon loses meaningful coding and science benchmarks, much of the public cannot test it yet, and several headline results depend on evaluation setups that differ across models.

For now, Gemini 4 Argon looks less like “Gemini with a bigger score” and more like Google’s attempt to make frontier models useful for jobs that are too long, messy and interconnected for a single clever answer.

Binary Verse AI will be tracking Argon’s wider API release, independent testing and real cost-per-task results as access expands. If you care about what these models can actually do beyond launch-day charts, follow Binary Verse AI for the next round of testing.

1. Is Gemini 4 Argon out now?

Yes, but only in limited release. Google is initially providing Gemini 4 Argon to trusted cybersecurity defenders through its Fairwind Program. Broader availability is planned for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. Google has not announced a date for general availability. I

2. When is the Gemini 4 Argon release date for everyone?

Google announced Gemini 4 Argon on September 30, 2026, but has not provided a firm date for broad public access. The rollout is phased, with trusted testers first and paid API customers and Google AI Ultra subscribers expected to receive access before wider consumer availability.

3. How much does Gemini 4 Argon cost?

Gemini 4 Argon’s introductory API price is $2 per million input tokens and $10 per million output tokens. Cached input receives a 95% discount, equivalent to $0.10 per million tokens at the introductory rate. After the introductory period, pricing rises to $4 input and $20 output per million tokens; Google has not specified when the introductory period ends.

4. Does Gemini 4 Argon really have a 1 million token output limit?

Yes. Google says Gemini 4 Argon increases the maximum output from 64K to 1 million tokens, allowing very long reasoning and generation trajectories. This should not be confused with the model’s input context window: maximum output and context capacity are separate specifications.

5. Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?

There is no universal benchmark winner. Argon leads several evaluations, including DeepSWE v1.1, AutomationBench, Vals Index, LABBench 2, RiemannBench and LVBench, but GPT-6 Astra leads some tests such as FrontierSWE v2 and OSWorld 2.0, while Claude Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench. The useful question is therefore which model performs best for a particular workload, not which model has the highest overall collection of scores.

Leave a Comment