Gemini 4 Argon Benchmarks, Specifications & Availability
Explore Gemini 4 Argon from Google: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Gemini 4 Argon with other models →Explore data coverage
Published specifications
- Provider
- Access
- Proprietary
- License
- Proprietary
- Context window
- 1M
- Total parameters
- Not published
- Active parameters
- Not published
- Released
- 2026-09-30
- Modalities
- text, image, video, file
- Family
- Gemini 4
Model card · Announcement · Website
Model notes
Frontier phased rollout via Fairwind Program. Google API intro pricing $2/$10 per 1M tokens (cached input 95% off), $4/$20 after. Not on OpenRouter yet.
Pricing · OpenRouter
No OpenRouter pricing is recorded for this model. Missing rates are not free.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Vals Index | 68.9%max effort | All 1 recorded result & sources68.9% · raw 68.9 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Vendor-reported Vals AI figure; Vals primary run lives in evaluators.json | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| AutomationBench v1.0.6 | 51.3%max effort | All 1 recorded result & sources51.3% · raw 51.3 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Private set via official Zapier public leaderboard; ranks #1 | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| Vals Finance Agent v2 | 65.4%max effort | All 1 recorded result & sources65.4% · raw 65.4 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Sourced from Vals AI per methodology | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Harvey Legal Agent Benchmark | 19.6%max effort | All 1 recorded result & sources19.6% · raw 19.6 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Sourced from Vals AI per methodology | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| DeepSWE v1.1 | 77.9%max effort | All 1 recorded result & sources77.9% · raw 77.9 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed, mini-swe agent harness; state of the art | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| FrontierSWE v2 | 55%max effort | All 1 recorded result & sources55% · raw 55 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Proximal official public leaderboard | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Vibe Code Bench | 91.9%max effort | All 1 recorded result & sources91.9% · raw 91.9 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Vals AI public leaderboard | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Terminal-Bench 4.0 | 57.4%max effort | All 1 recorded result & sources57.4% · raw 57.4 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed; peers from official public leaderboard (highest thinking level) | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| PostTrainBench | 45.3%max effort | All 1 recorded result & sources45.3% · raw 45.3 % Headline · max effort · Own vendor Source/record date: 2026-09-30 v1.1 OpenCode harness, 10h H100 budget, weighted aggregate | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Terminal-Bench Science 0.1 | 57.6%max effort | All 1 recorded result & sources57.6% · raw 57.6 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed with 6x verifier timeout | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LABBench 2 | 88.8%max effort | All 1 recorded result & sources88.8% · raw 88.8 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed with Linux terminal, bioinfo tools, internet | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| RiemannBench | 76%max effort | All 1 recorded result & sources76% · raw 76 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Surge public leaderboard | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks up to 128k BFS F1 | 99.7 f1max effort | All 1 recorded result & sources99.7 f1 · raw 99.7 f1 Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed, 650 items up to 128k tokens, BFS F1 | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GraphWalks 256k-1M BFS F1 | 84.2 f1max effort | All 1 recorded result & sources84.2 f1 · raw 84.2 f1 Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed, 200 problems 256k-1M tokens, BFS F1 | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Agents' Last Exam | 39.5%max effort | All 1 recorded result & sources39.5% · raw 39.5 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed ALE-Claw harness 5h, binary pass rate, safety filters on | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| OSWorld 2.0 | 69.2%max effort | All 1 recorded result & sources69.2% · raw 69.2 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed, offline subset partial score, max of 3 runs, Gemini CUA harness | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Chartography | 71.6%max effort | All 1 recorded result & sources71.6% · raw 71.6 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Without tools, Surge public leaderboard | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| LVBench | 91.7%max effort | All 1 recorded result & sources91.7% · raw 91.7 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed without tools at 1FPS; state of the art | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| CWE-bench v1 | 68%max effort | All 1 recorded result & sources68% · raw 68 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Official public leaderboard pass@1; ties for first | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| Gray Swan IPI | 0.7%max effort | All 1 recorded result & sources0.7% · raw 0.7 % Headline · max effort · Own vendor Source/record date: 2026-09-30 K=15 attack success rate (lower better); sourced from Gray Swan; leads chart | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 53high effort | All 1 recorded result & sources53 · raw 53 index Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gemini-4-argon |
| Cost per Intelligence Index task · Artificial Analysis | $1.99high effort | All 1 recorded result & sources$1.99 · raw 1.99 USD Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/gemini-4-argon |
| Vals Index · Vals AI | 68.9%high effort | All 1 recorded result & sources68.9% · raw 68.9 % Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $15.68. https://www.vals.ai/benchmarks/vals_index |
| Vibe Code Bench v1.1 · Vals AI | 91.9%high effort | All 1 recorded result & sources91.9% · raw 91.91 % Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $8.11. https://www.vals.ai/benchmarks/vibe-code |
| Money gain · Andon Labs | $13218.16unknown effort | All 1 recorded result & sources$13218.16 · raw 13218.16 $ Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked. https://andonlabs.com/evals/vending-bench-2 |
| Blueprint Bench · Andon Labs | 54.4%unknown effort | All 1 recorded result & sources54.4% · raw 54.400000000000006 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100. https://andonlabs.com/evals/blueprint-bench-2 |
Read how we select and source scores or the comparison guide.