Gemini 4 Argon Arrives as a Gated Release With Frontier Numbers

Google published frontier numbers for Gemini 4 Argon and opened access to a vetted group of cyber defenders first. The benchmarks are real and self-reported, and the gate is the part that sh

Netics editorial card on Gemini 4 Argon, contrasting the published frontier benchmarks with the gated rollout to vetted defenders.
Netics editorial card on Gemini 4 Argon: frontier benchmark numbers published on launch day, and a rollout that starts with vetted cyber defenders.

TL;DR

  • Google announced Gemini 4 Argon on 2026-09-30 and is releasing it to a set of trusted cyber defenders through its Fairwind Program, with wider access starting with paid API customers and Google AI Ultra subscribers.
  • The published numbers are specific: 77.9% on DeepSWE v1.1, 51.3% on Zapier's AutomationBench, 91.7% on LVBench, a tie for first at 68% on CWE-bench v1, and an output limit raised from 64K tokens to 1M.
  • The benchmark claims are Google's own measurements, which is normal and worth saying out loud before anyone rebuilds a roadmap on them.
  • Google's illustration of the model's value is its own fleet: agents applied memory optimisations that freed over 300 TiB, and Rust migrations from tens of thousands of lines up to an 800K-line kernel.
  • The build decision belongs to what your team can call today, so the practical move is to keep the routing layer you already have and pilot the gated tier when access opens.

What Google announced on 2026-09-30

The launch post describes a model built for long-horizon work: sustained deep reasoning across complex workflows, with emphasis on software engineering, enterprise knowledge work in areas such as legal and finance, and cybersecurity defence. Google frames the release as phased and cautious. Argon is going first to a vetted group of cyber defenders through the Fairwind Program, Google says it is engaging with the United States government's voluntary pre-release process, and wider availability will start with paid API customers and Google AI Ultra subscribers — with no date attached to that step.

That combination is worth reading carefully, because two things are being said at once. The capability is being presented at frontier level, and the access is being rationed. A model that a security team can use to find and fix vulnerabilities is also a model that a security team can use to find and break things, and the vendor has decided the second risk dominates the first until its own evaluations come back.

Official Google image from the Gemini 4 Argon launch post: the model's key art with the Gemini 4 Argon wordmark.
Official Google image from the Gemini 4 Argon launch post (blog.google, 2026-09-30, retrieved 2026-10-04): the model announced on 2026-09-30 and released first to vetted cyber defenders.

The numbers Google attached to Argon

The figures in the post are concrete, and they cover the ground a technical buyer cares about. On DeepSWE v1.1, which Google describes as measuring real-world long-horizon software engineering, Argon scores 77.9%. On Zapier's AutomationBench, which measures end-to-end execution across core business functions, it ranks first at 51.3%. On LVBench, for long video understanding, Google reports 91.7% as state of the art. On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Argon ties for first with a top score of 68%, building on the performance Gemini 3.8 Flash Cyber established on CWE-bench v0.

Two more published details matter for how a system gets designed. The output token limit rises to an industry-leading 1 million tokens, up from 64,000 — a five-digit multiple, and one that changes what a single request can produce in a long migration or a large document workflow. Pricing after the introductory period is set at $4 per million input tokens and $20 per million output tokens, which is the number to plan against rather than the launch rate.

Every one of those figures is Google's own measurement. That is ordinary practice at a launch, and it is also the reason a cautious team treats a launch post as a hypothesis. A benchmark result tells you what a model did on a benchmark under a vendor's configuration. It does not tell you what it will do on your contract templates, your legacy migration or your incident queue.

Netics editorial metric grid: 68% on CWE-bench v1 tied first, 77.9% on DeepSWE v1.1, 51.3% on AutomationBench in first place, and a 1 million token output limit.
Netics editorial metric grid built from the figures in Google's launch post (blog.google, 2026-09-30): four published numbers that describe capability and capacity rather than access.

The engineering examples in the announcement

The most interesting passage in the post is not a benchmark. Google describes agents analysing fleet-wide profiling telemetry to identify and apply memory optimisations across its data centres, freeing over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB of total savings. It also describes agents migrating C and C++ codebases to Rust across the company, scaling from tens of thousands of lines in core libraries such as re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel, and notes that these rewrites pass automated and manual auditing, emulation testing and review before reaching production. In one example, agents replaced 32,000 lines of SIMD code in the libgav1 video decoder and produced a memory-safe result running 2.7 times faster than the previous Rust port with identical video output.

Read that as a deployment pattern rather than a capability boast. The agents act on production fleets and kernel-adjacent code, and every example is wrapped in auditing, emulation and review gates. Google also states that it is hardening its sandboxed environments by isolating and sealing them before high-risk training or evaluations begin, in line with its agent control roadmap, and that it plans to share those practices with partners.

There is an awkward implication for anyone selling agent deployments to a regulated client. If the reference implementation of frontier agents runs inside the vendor's own estate, behind audit and emulation gates, then the controlling design question is not which model is smartest. It is who owns the gate.

Why a gated release changes the build decision

A gated rollout splits the audience in two, and the split is what a team has to plan around. Reviewers can read the numbers but cannot run them on their own tasks. A defence team inside a partner programme can run them and report back. Everyone else is deciding whether to wait, to build against a model they can call today and swap later, or to route by task.

Netics' position is the boring one, and it has not changed across the last few launches. Build against the model you can call, keep the call behind an interface, and treat a gated tier as a pilot decision with a date on it rather than as an architecture. Teams that pinned production workflows to a preview model have already paid for that lesson twice this year, and the cost was not the inference bill; it was the migration weekend.

The pricing detail deserves the same treatment. An introductory rate is a temporary number, and the published post-introductory rate of $4 per million input tokens and $20 per million output tokens is the one to put in the business case. If a workflow only returns its value at the launch price, it does not return its value.

Netics editorial comparison: the benchmarks published with Argon, including DeepSWE v1.1 at 77.9%, AutomationBench at 51.3%, LVBench at 91.7%, CWE-bench v1 tied first at 68% and the 64K to 1M output change, against the access opened so far through the Fairwind programme, with paid API customers and Google AI Ultra subscribers to follow.
Netics editorial comparison built from the launch post: the published capability figures describe what the model does in Google's evaluations, while the access column describes what a given team can actually test.

What a team should do this week

Three concrete moves, none of them expensive. First, pick the two or three tasks in your estate where a long-horizon model would visibly help — a migration assessment, a contract review, a patch backlog triage — and write down the pass criteria before you touch a vendor page. Second, keep your evaluations in the repository so that a candidate model is judged on your data by the same script every time; Argon and its successors will each arrive with a different leaderboard, and yours is the one that transfers. Third, register interest in the gated tiers, because programme access is how a regulated team gets to evaluate a restricted model ahead of general availability.

If the model layer is the part you want to keep interchangeable, that is the architecture behind the self-hosted AI agent platform we build for clients: the loop, the tools and the evaluation set stay in the client's estate, and the model is a value in configuration. Our earlier piece on Gemini 3.8 Flash for defenders covers the model in the same family that a team can deploy today.

Official Google chart from the Gemini 4 Argon launch post: the CWE-bench v1 leaderboard, with equivalent configurations highlighted, where Argon ties for first place at 68%.
Official Google chart from the Gemini 4 Argon launch post (blog.google, 2026-09-30, retrieved 2026-10-04): on CWE-bench v1, which measures vulnerability remediation, Argon ties for first with a top score of 68%.

The uncomfortable trade-off, stated plainly: frontier capability increasingly arrives behind a gate, and the teams that get early access are the ones with a partner relationship, a compliance story and an evaluation suite ready to run. Preparing those three things is a week of work that pays off the next time a model like this one is announced.

Sources

Source: Introducing Gemini 4 Argon — blog.google, 2026-09-30 (long-horizon frontier model framing across software engineering, enterprise knowledge work and cybersecurity defence; phased rollout to a set of trusted cyber defenders through the Fairwind Program; engagement with the US government's voluntary pre-release process; wider availability starting with paid API customers and Google AI Ultra subscribers; 77.9% on DeepSWE v1.1; first place at 51.3% on Zapier's AutomationBench; state of the art at 91.7% on LVBench; a tie for first at 68% on CWE-bench v1 building on Gemini 3.8 Flash Cyber's CWE-bench v0 performance; output token limit raised from 64K to 1M; post-introductory pricing of $4 per million input tokens and $20 per million output tokens; agents freeing over 300 TiB of memory across Google's data centres with an estimated 500 TiB to 1 PiB of total savings; C and C++ to Rust migrations from tens of thousands of lines up to more than 800,000 lines for the Fuchsia Zircon kernel under automated and manual auditing, emulation testing and review; the libgav1 decoder example replacing 32,000 lines of SIMD code for a result 2.7 times faster than the previous Rust port with identical output; sandboxed environments isolated and sealed before high-risk training or evaluations). Internal linkage: Gemini 3.8 Flash Cyber Gives Defenders a Faster Patch Loop. More on Netics' work at neticslabs.com.

Source: Introducing Gemini 4 Argon — blog.google, 2026-09-30. Figures: official Google images from the launch post, retrieved 2026-10-04.