Gemini 4 Argon: What Google’s New Flagship AI Model Actually Changes
Google’s Gemini 4 Argon targets long-running coding, cybersecurity, and enterprise AI workflows, with a one-million-token output limit and a phased rollout.
Google has finally moved its flagship Gemini line forward, but Gemini 4 Argon is arriving in a very different AI market from the one that greeted earlier Gemini generations. Coding agents are now expected to work across entire repositories, companies are experimenting with AI on legal and financial workflows, and cybersecurity models are beginning to find and repair vulnerabilities with far less human direction.
Announced on September 30, Gemini 4 Argon is Google DeepMind’s new frontier model for those longer, more complicated jobs. Google says it was designed around software engineering, professional knowledge work, and defensive cybersecurity, rather than simply improving the conversational experience of the Gemini chatbot. For now, however, most users cannot actually try it. Google is initially providing Argon to selected cybersecurity defenders through its Fairwind Program while it conducts additional safety work ahead of a wider release.
That restricted launch matters because much of what we know about Argon comes from Google’s own testing. The early numbers are strong, but the more interesting story is how Google appears to be designing the model for much longer periods of autonomous work.
Coding Is One of Argon’s Strongest Early Claims
Google reports that Gemini 4 Argon scored 77.9% on DeepSWE v1.1, a benchmark built around longer software-engineering tasks across real codebases. Google describes the result as state-of-the-art and says its engineers are already using Argon for work ranging from routine debugging to large-scale code migrations and algorithm development.
Some of the company’s internal examples are considerably more ambitious than ordinary code completion. Google says Argon agents have been helping migrate C and C++ projects to Rust, including work touching hundreds of thousands of lines in the Fuchsia Zircon kernel. In another project involving the open-source libgav1 video decoder, Google says agents replaced roughly 32,000 lines of SIMD code while producing a Rust implementation that ran 2.7 times faster than the previous Rust port with identical video output. Those are Google’s reported results, not independent evaluations, but they indicate the kind of engineering workload the company wants Argon to handle.
The 77.9% DeepSWE figure deserves more caution than a leaderboard position might suggest. Epoch AI reviewed DeepSWE v1.1 in September and classified the benchmark as flawed after finding problems in at least 23 of its 113 tasks. The issues included false negatives in which valid model behaviour could still be scored as a failure because of how the benchmark’s verifier handled modified tests.
That does not make Argon’s result meaningless, nor is there evidence that Google manipulated the benchmark. It does illustrate why a few percentage points between frontier models should not automatically be interpreted as a definitive ranking. Coding-agent performance depends not only on the underlying model but also on the tools, prompts, execution environment and software harness surrounding it.
One Million Output Tokens Is the More Unusual Upgrade
Argon’s most striking technical specification may not be a benchmark score. Google has raised the model’s maximum output from 64,000 tokens to one million tokens, giving a single run dramatically more room to reason, generate code and continue through extended tasks.
That is an output allowance, not simply a one-million-token input context window. The distinction matters. Large context windows let a model read enormous amounts of information. At the same time, a huge output budget gives an agent room to keep working through a lengthy trajectory without being forced to stop after a relatively short response.
For coding, that could matter during repository-wide migrations, repeated test-and-fix cycles or tasks involving many dependent files. In research, it could let an agent collect information, analyse it, revisit earlier conclusions, and produce a much larger body of work in one continuous process.
Longer generation does not automatically make a model more reliable, however. Autoregressive models still produce their outputs sequentially, and errors made early in a complicated workflow can influence later decisions. A million-token allowance therefore creates room for deeper work, but it also gives the model a much longer runway for mistakes to accumulate. The real question is not whether Argon can produce a million tokens, but whether useful reasoning remains coherent over the portion of that budget developers actually use.
Google Is Positioning Argon as More Than a Coding Model
Software engineering is only part of Google’s pitch. The company says Argon leads the Vals Index, which evaluates economically relevant work across areas including finance, legal services, coding and taxation. Google also reports a 51.3% result on Zapier’s AutomationBench and a 91.7% score on LVBench, an evaluation focused on understanding long-form video.
This broader positioning matters because the competition among frontier models is increasingly moving beyond chatbot quality. Businesses want models that can operate inside workflows: review several documents, use tools, make decisions, and continue through a task rather than returning one polished answer.
Google says thousands of its employees already use Argon internally. One group reportedly used agents to identify memory optimisations across Google’s data-centre fleet, freeing more than 300 TiB of memory after deployment and potentially identifying more savings. Another Google research team used Argon while optimising quantum-computing algorithms, where the company says it improved on a published baseline.
Those examples are difficult to compare with traditional AI benchmarks, but they may ultimately matter more if they translate into repeatable productivity gains.
Cybersecurity Explains the Unusual Rollout
Google is being particularly careful with Argon because it says the model has substantially improved cybersecurity capabilities. According to the company, Argon can autonomously discover, validate, and patch software vulnerabilities, and selected defensive-security partners are receiving access through the Fairwind Program before the model becomes generally available.
Google says cloud-security company Wiz has already used Argon through its Scan for Good initiative and that the model found a critical vulnerability affecting healthcare software that previous frontier models had missed. On CWE-bench v1, an evaluation focused on fixing security weaknesses, Google reports that Argon scored 68%, tying the highest reported result.
More capable cyber models create an obvious dual-use problem. The same ability to inspect software deeply enough to discover vulnerabilities can potentially be useful to attackers. Google says it is testing Argon against misuse, indirect prompt injection, and behaviour that could move outside a user’s intended boundaries. The company is also participating in the U.S. government’s voluntary process for pre-release model access before broadening availability.
Google Has Distribution Its Rivals Cannot Easily Replicate.
Argon also arrives with an advantage that has little to do with benchmarks: Google already operates one of the largest AI distribution networks in the industry.
During Alphabet’s second-quarter 2026 earnings update, CEO Sundar Pichai said Google’s model APIs were processing roughly 22 billion tokens per minute, up from 16 billion in the previous quarter. Google also reported more than nine million developers building monthly with its models and developer products, while its Antigravity agentic development platform had reached more than 2.4 million weekly active users.
OpenAI still has substantial developer momentum. In June, it said Codex had surpassed five million weekly active users, showing just how competitive the market for AI-assisted development has become. Google’s advantage is that it can eventually distribute Gemini models across Cloud, Workspace, Android, Search and its developer ecosystem rather than relying on a single coding product to establish adoption.
Pricing reinforces that strategy. Google says Argon will initially cost $2 per million input tokens and $10 per million output tokens, with cached input priced at a 95% discount. Those introductory rates are temporary; after that period, Google says standard pricing will rise to per million input tokens and $ per million output tokens
Argon Looks Powerful, but the Real Test Has Barely Started
Gemini 4 Argon is an important release for Google because it shows that the company is still competing aggressively at the frontier rather than concentrating only on smaller, cheaper Flash models. The combination of long-running agent capabilities, software-engineering performance, cybersecurity specialisation, and a one-million-token output ceiling makes Argon considerably more ambitious than a routine Gemini refresh.
The early benchmark numbers should not be treated as a final verdict. DeepSWE has documented flaws, much of Argon’s performance data currently comes from Google itself, and the model has not yet been released widely enough for developers to test it across the messy range of workloads that reveal a model’s real strengths and weaknesses.
That broader testing will matter more than whether Argon sits a few points above another model on a chart. If Google can make the model reliable over genuinely long tasks while keeping the economics workable at the scale of its existing ecosystem, Gemini 4 Argon could become more consequential for its ability to keep working than for any single benchmark score.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0