Creuto is now an OpenAI Select Partner Read More
Gemini 4 Argon scores 77.9% on DeepSWE v1.1 and ships at $2/$10 per million tokens - introductory. Plus the build Google sends out without cyber guardrails.

Google shipped Gemini 4 Argon on 30 September 2026 with a version that has the cyber guardrails removed. Not for everyone — for a vetted list of defenders. The rest of the announcement is the usual frontier release: 77.9% on DeepSWE v1.1, 91.7% on LVBench, a one million token output limit, and introductory pricing of $2 per million input and $10 per million output. The guardrail decision is the part that will still matter in six months.
Every figure below is from Google's own announcement, dated 30 September 2026. There is no published model card for Argon yet, which means the evaluation harnesses, attempt counts and scaffolds behind these numbers are not public. Treat them as vendor-reported.
| Benchmark | Score | Google's claim |
|---|---|---|
| DeepSWE v1.1 | 77.9% | Real-world software engineering tasks |
| LVBench | 91.7% | State of the art, long video understanding |
| CWE-bench v1 | 68% | Ties for first place, security vulnerabilities |
| AutomationBench | 51.3% | Ranks first |
Google also reports leading results on the Vals Index, Vals Finance Agent v2, Harvey's legal agent benchmark and the Gray Swan IPI benchmark for prompt injection robustness, but publishes no number against any of them. A claim without a figure is a claim you cannot check, so we are not repeating those as results.
The number worth pausing on is AutomationBench at 51.3%. First place on a benchmark where the leader fails roughly half the tasks tells you the category is young, not that the model is unreliable. It is the wrong figure to put in a business case for unattended automation.
CWE-bench v1 at 68% is a security benchmark, and 68% on finding vulnerabilities in curated repositories is not 68% on your codebase. We see the same gap in the systems we build: a model that clears a public coding benchmark can still stall on an internal service with undocumented conventions and a ten-year-old migration history. The published score sets an upper bound on what to expect, never a floor.
Introductory pricing is $2 per million input tokens and $10 per million output tokens. Google states the standard rate afterwards will be $4 per million input and $20 per million output — a straight doubling. Cached input is priced at 95% off the input token price, which works out to $0.10 per million at the introductory rate and $0.20 once standard pricing lands.
Google has not published a date for the changeover. That is the single most important omission in the announcement, because a budget built on $2/$10 becomes a budget for $4/$20 on a day Google chooses. If you model Argon spend this quarter, model both columns.
Worth noting where Argon lands against the obvious comparison. OpenAI lists gpt-6.1-sol at $2 input and $10 output per million tokens as its standing rate, not an introductory one. At introductory pricing Argon matches it. At standard pricing it is twice the price. Anyone reading the launch post as a price cut has the arithmetic backwards — it is a price cut with an expiry Google has not disclosed.
The output limit is the more interesting change: one million output tokens, up from 64K. That removes the chunk-and-stitch machinery most long-generation pipelines carry, but it also means a single runaway call can now bill $10 of output on its own at introductory rates, or $20 at standard. Put a max output token cap in your client before you put Argon in production.
As of 1 October 2026, Google's Gemini API pricing page carries no Gemini 4 row at all — the most recent entries are the Gemini 3.x Flash family. The prices above exist only in the launch blog post. If you are putting figures in a procurement document this week, cite the blog post and the date, because the pricing page cannot corroborate them.
Argon is rolling out first through Google's Fairwind Program, which gives governments, national cyber authorities, critical infrastructure operators and core technology platforms early access. The programme page names Argon as its frontier model, pairs it with CodeMender for vulnerability discovery and patching, and reports over 650 partners.
The launch post goes further. It states that for trusted defenders and Google's own internal teams, Google will be releasing Argon without cyber guardrails. That is a two-tier capability split, not a two-tier price list: the same weights, different refusal behaviour, allocated by who you are.
Here is the honest complication. The Fairwind programme page does not mention an unguardrailed build anywhere. It describes restricted dual-use tasks — authorised threat simulation, reverse engineering and malware analysis — under phishing-resistant MFA, access limited to named cybersecurity teams, employee access tracking and background checks on applicants. Two primary sources, two different descriptions of the same distribution. Until Google reconciles them, the only defensible statement is that Google has said it will ship a version without cyber guardrails and has not yet documented what that version can do that the general one cannot.
The case against Google here is straightforward: capability gating by customer identity is security through paperwork. Background checks and MFA do not stop a vetted organisation's compromised laptop, and the offensive capability, once released, exists. That argument is correct and Google has not answered it publicly.
The case for it is that the capability exists either way, and the choice is whether defenders get it before or after attackers reconstruct it. Google is betting on before. You do not have to agree with the bet to plan around it — if your threat model assumes frontier offensive tooling stays behind refusals, that assumption expired on 30 September.
On published numbers, nobody can tell you. Argon's 77.9% on DeepSWE v1.1 has no directly comparable figure from OpenAI on the same harness, and cross-benchmark comparison is the mistake that produces confident wrong answers about model choice. The decision that actually pays is the one we push clients toward in every AI engineering engagement: build a small evaluation set from your own traffic, twenty to fifty real tasks with known-good outputs, and run the candidates against it.
Three practical notes for teams weighing Argon now:
If you are choosing between frontier models rather than adopting one on reputation, our comparison of GPT-6 Astra against GPT-5.6 sets out the method: cost per successful task on your own evaluation set, not headline benchmark scores. Run Argon through that once it is generally available. Until then, the only Argon figure you can act on is the price — and that one has an expiry date nobody has published.
Gemini 4 Argon launched at introductory pricing of $2 per million input tokens and $10 per million output tokens, with cached input at 95% off the input price. Google states the standard rate afterwards is $4 input and $20 output per million, but has not published the changeover date.
Google reports Gemini 4 Argon at 77.9% on DeepSWE v1.1 for real-world software engineering tasks and 68% on CWE-bench v1 for security vulnerabilities, where it ties for first. Both figures are vendor-reported, and no public model card documents the evaluation harness.
Published figures cannot answer that. Google has not reported Argon on the same harnesses OpenAI uses, and comparing scores across different benchmarks produces confident wrong answers. Build an evaluation set from your own traffic and measure cost per successful task on both.
Google's launch post states it will release Argon without cyber guardrails to trusted defenders and its own internal teams, while the general release keeps them. Access runs through the Fairwind Program, which vets governments, critical infrastructure operators and core technology platforms.
Argon is rolling out first to a set of trusted cyber defenders through Google's Fairwind Program, which reports over 650 partners. Google says the broader rollout reaches paid API customers and Google AI Ultra subscribers first, with developers, enterprises and consumers following.
Google reports an output limit of one million tokens for Gemini 4 Argon, up from 64K previously. That removes chunking machinery from long-generation pipelines, but it also lets one unbounded call bill $10 of output at introductory rates, so set an explicit maximum output cap.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand