Key Notes
- Gemini 4 Argon is Google's most powerful AI model yet, with an output limit of one million tokens.
- Introductory API rates are $2 per million input tokens and $10 per million output tokens, with higher rates to follow.
- Access starts with trusted cyber defenders before expanding to paid API customers and Google AI Ultra subscribers.
Google unveiled Gemini 4 Argon on September 30, introducing its most powerful AI model yet with a focus on demanding software projects, professional research, and cybersecurity. The company is initially releasing it to selected defensive security partners, while preparing broader access for paid API customers and Google AI Ultra subscribers.
In its announcement, Google positioned Argon as a model capable of following complex tasks through many steps. Its most striking specification is a one-million-token output limit, alongside introductory API prices of $2 per million input tokens and $10 per million output tokens.
The launch gives Alphabet a new contender in the race to turn AI systems into dependable workers for developers and businesses. It also introduces a distinction that matters for anyone hoping to try the model immediately: Google has announced Argon and begun a restricted rollout, but has not opened general access or named a firm date for it.
A Flagship Built for Longer Tasks
Argon’s expanded output budget is intended to support extended reasoning and substantial work within a single run. Google says the ceiling has risen from 64,000 tokens to one million. An output limit describes how much a model can generate; it should not be confused with the amount of source material it can accept as input.
For a coding agent, additional room could support a longer sequence of investigation, implementation, and testing before it reaches a generation limit. For a research system, it could allow a more extensive chain of analysis. Those are potential benefits of the larger budget, rather than a guarantee that longer answers will be better or every task will require it.
Google reports a 77.9% result on DeepSWE v1.1, which evaluates extended software engineering work, and 51.3% on Zapier’s AutomationBench for business task execution. It also reports 91.7% on LVBench, a test of understanding long videos. Together, those results illustrate the mix of coding, document work, and visual reasoning Google wants Argon to handle.
How Argon Compares on Professional Work
The Vals Index provides an external view of the model’s enterprise performance. Its published leaderboard places Gemini 4 Argon first at 68.90%, followed by Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%. GPT-6 Astra scores 63.13% on the same index.
Vals combines finance, software development, legal, and tax evaluations, weighting sectors by their share of US economic output. That makes the index useful for comparing a particular collection of professional tasks. It does not establish that one model is better at every job, or that the same ranking will hold for a company’s private documents and software.
The cost and timing columns add another qualification. Vals lists Argon’s average cost per test at $15.68 and duration at 46 minutes and 33 seconds, using $4 input and $20 output pricing. Those are evaluation results under its setup, not the introductory API rates. Businesses still need to compare accuracy, task duration, and total spending against their own requirements.
The competitive pressure is already visible in AIstify’s coverage of Sonnet 5.5, where Anthropic emphasized everyday work and lower task costs. Argon enters a market in which completing a difficult assignment reliably matters as much as winning a headline benchmark.
Google Points to Work Inside Its Own Systems
Google’s internal examples offer a more concrete picture of its ambitions. The company says Argon agents identified memory improvements that freed more than 300 TiB across its data centers after deployment. It also describes work on moving substantial C and C++ codebases to Rust, including projects involving hundreds of thousands of lines.
These examples remain company-reported outcomes, and Google says major rewrites undergo automated checks, manual auditing, and other review before production use. That detail is significant: generating a large patch and deploying a trustworthy replacement are separate engineering milestones. Existing behavior, performance, and compatibility all have to survive the transition.
For organizations assessing coding agents, the practical question is therefore broader than how quickly they can produce code. A useful system must help engineers understand changes, reproduce results, and catch regressions. Argon’s internal deployments are evidence of the kinds of projects Google is attempting, while broader customer experience will determine how well those gains transfer.
Introductory Pricing Comes With a Later Increase
Google’s launch pricing puts Argon at $2 per million input tokens and $10 per million output tokens. Cached input receives a 95% discount, which works out to $0.10 per million tokens at the introductory input rate. Caching can reduce spending when an application repeatedly supplies eligible, previously processed material.
| Token type | Introductory rate | Later rate |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
A footnote in Google’s announcement says input and output rates will double after the introductory period. The announcement does not give that period an end date. Buyers should therefore distinguish launch pricing from the longer-term rates when estimating the cost of a service built around Argon.
The larger output allowance also makes token consumption important. At the advertised introductory rate, one million billed output tokens would cost $10 before input and any other applicable charges. That is a simple illustration, not a typical task estimate: reaching the maximum is optional, and actual costs depend on how an application uses the model.
Why Cybersecurity Partners Receive Access First
The initial release runs through Google’s Fairwind Program. The program prioritizes organizations protecting public services and critical infrastructure, including government agencies, healthcare providers, and telecommunications networks. Its rationale is to give defenders access to stronger tools before comparable capabilities become widely available to attackers.
Fairwind participation is controlled. Google describes organizational vetting, restrictions on sharing access, and requirements around authentication and tracking employee use. Authorized security research can involve examining the same weaknesses an attacker might exploit, so access conditions are part of how the program separates legitimate investigation from misuse.
Google says trusted defenders will receive Argon without cyber guardrails so they can use its full defensive capabilities. That provision applies to the restricted release; it is not a promise of unrestricted access for ordinary consumers. The company says it is also participating in the US government’s voluntary process for pre-release model access.
For security performance, Google reports a 68% score on CWE-bench v1, tying the leading result. The benchmark asks agents to audit repositories and repair vulnerabilities. Its methodology requires a fix to block the exploit while preserving existing functionality, rather than merely producing a plausible explanation of a problem.
That distinction helps explain why a strong score still leaves work for security teams. Passing a controlled benchmark does not establish complete coverage of an unfamiliar production system. Review, testing, and deployment controls remain necessary, particularly when a patch affects software that supports essential services.
Wider Availability Depends on Further Testing
Google plans to expand Argon beyond the initial cohort after additional testing, starting with paid API customers and Google AI Ultra subscribers. Developers, enterprises, and consumers are all part of the intended audience, but the announcement does not provide a public rollout calendar.
The company describes work on misuse prevention, prompt injection resistance, monitoring for actions outside a user’s intentions, and stronger isolation for testing environments. Its earlier security roadmap frames agent protection as several layers working together, including controlled permissions and safeguards that remain useful when model behavior is imperfect.
Argon also broadens Google’s model strategy beyond the conversational experience highlighted in its Gemini Live releases. The emphasis here is on sustained work: managing a substantial code change, completing professional research, or investigating a security weakness through to a verified result.
The next test will come as access expands. Argon’s launch establishes a new flagship, published pricing, and promising performance evidence. Its wider impact will depend on whether customers can reproduce useful results at an acceptable cost, with enough visibility and control to trust the work being done.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.