Google announced Gemini 4 Argon, first released to cyber defenders

Google announced Gemini 4 Argon on 2026-09-30, calling it “our next era of frontier intelligence.” Unusually for a flagship launch, it did not go to the public: Argon is “rolling out to a set of trusted cyber defenders through our Fairwind Program,” and Google says it is taking part in the U.S. government’s voluntary process for pre-release model access while it widens availability. For trusted defenders and Google’s own teams, Argon is released “without cyber guardrails.” Developers, enterprises and consumers get it later, starting with paid API customers and Google AI Ultra subscribers.

The headline technical change is output length: the output token limit rises to 1 million tokens, from the previous 64,000. Google reports 77.9 percent on DeepSWE v1.1, first place on Zapier’s AutomationBench at 51.3 percent, 91.7 percent on LVBench (long video understanding) and a tie for first at 68 percent on a vulnerability-remediation benchmark. Internal examples include Argon agents freeing over 300 TiB of memory across Google data centers through profiling-driven optimizations, and migrating C/C++ code to Rust, including a libgav1 decoder that runs 2.7x faster than the earlier Rust port. Introductory pricing is 2 dollars per million input tokens and 10 dollars per million output, with cached input 95 percent off; after the introductory period the price becomes 4 and 20 dollars.

On safety, Google lists CBRN and cyber refusal safeguards, prompt-injection hardening (it says Argon leads Gray Swan’s indirect prompt injection benchmark), misalignment mitigations that monitor the model’s chain of thought and actions and can stop execution, and sealed sandboxes for high-risk training and evaluation.

Why it matters: Argon is another frontier model shipped first, and for now only, to vetted defenders, which makes the staged “cyber-first” release a pattern across labs rather than a one-off. The 1M-token output limit also raises the ceiling on single-trajectory agent work by more than an order of magnitude. What it does not show: every number is Google’s own, several comparisons are against Google-chosen baselines, and the general public cannot yet use the model, so independent evaluation is not possible at launch.