Examenos

Gemini 3.8 Flash Benchmarks, Specifications & Availability

Explore Gemini 3.8 Flash from Google: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Gemini 3.8 Flash with other models →Explore data coverage

Published specifications

Provider
Google
Access
Proprietary
License
Proprietary
Context window
1M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-03
Modalities
text, image, video, audio, file
Family
Gemini 3

Model card · Announcement · Website · OpenRouter

Model notes

Flash-tier workhorse for long-horizon SWE and agentic workflows.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object][object Object]2026-10-03 · OpenRouter source
Recorded pricing notes

min-healthy endpoint minima 2026-10-01; discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); min-healthy endpoint minima 2026-10-01; discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); min-healthy endpoint minima 2026-10-02; discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); min-healthy endpoint minima 2026-10-03; discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied); discount flag 0.5 on Google AI Studio (undocumented, not applied)

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
AA Coding Agent Index v1.461.2unknown effort
All 1 recorded result & sources

61.2 · raw 61.2 index

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra table

https://openai.com/index/gpt-6-astra/
BioMysteryBench88.8%unknown effort
All 1 recorded result & sources

88.8% · raw 88.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

BioMysteryBench human-solvable

https://deepmind.google/models/model-cards/gemini-3-8-flash
CharXiv Reasoning86.2%unknown effort
All 1 recorded result & sources

86.2% · raw 86.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

CharXiv Reasoning no tools

https://deepmind.google/models/model-cards/gemini-3-8-flash
DeepSWE v1.173.7%unknown effort
All 1 recorded result & sources

73.7% · raw 73.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

DeepMind model card (prefer over OpenAI 73.8 if reconciling)

https://deepmind.google/models/model-cards/gemini-3-8-flash
FrontierCode 1.1 Extended56.3%unknown effort
All 1 recorded result & sources

56.3% · raw 56.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra table (Gemini 3.8 Flash col)

https://openai.com/index/gpt-6-astra/
FrontierCode 1.1 Main43.6%unknown effort
All 1 recorded result & sources

43.6% · raw 43.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported in OpenAI GPT-6 Astra table

https://openai.com/index/gpt-6-astra/
GDP.PDF35%unknown effort
All 1 recorded result & sources

35% · raw 35 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

GDP.PDF all-pass rate

https://deepmind.google/models/model-cards/gemini-3-8-flash
GDPval-AA v21545unknown effort
All 1 recorded result & sources

1545 · raw 1545 Elo

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

Elo

https://deepmind.google/models/model-cards/gemini-3-8-flash
GPQA Diamond95.3%unknown effort
All 1 recorded result & sources

95.3% · raw 95.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
Harvey Legal Agent10%unknown effort
All 1 recorded result & sources

10% · raw 10 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

All pass rate

https://deepmind.google/models/model-cards/gemini-3-8-flash
HealthBench Professional52.1%unknown effort
All 1 recorded result & sources

52.1% · raw 52.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/
HLE-Verified54.9%unknown effort
All 1 recorded result & sources

54.9% · raw 54.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

HLE-Verified

https://deepmind.google/models/model-cards/gemini-3-8-flash
LABBench286.2%unknown effort
All 1 recorded result & sources

86.2% · raw 86.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

LABBench2

https://deepmind.google/models/model-cards/gemini-3-8-flash
LVBench87.8%unknown effort
All 1 recorded result & sources

87.8% · raw 87.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

LVBench long video (agentic)

https://deepmind.google/models/model-cards/gemini-3-8-flash
OSWorld 2.059%unknown effort
All 1 recorded result & sources

59% · raw 59 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

Partial score; batch tool enabled

https://deepmind.google/models/model-cards/gemini-3-8-flash
Terminal-Bench 2.189.4%unknown effort
All 1 recorded result & sources

89.4% · raw 89.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://deepmind.google/models/model-cards/gemini-3-8-flash
Terminal-Bench 4.019.1%unknown effort
All 1 recorded result & sources

19.1% · raw 19.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://deepmind.google/models/model-cards/gemini-3-8-flash
Vals Finance Agent v261.4%unknown effort
All 1 recorded result & sources

61.4% · raw 61.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

Vals Finance Agent v2

https://deepmind.google/models/model-cards/gemini-3-8-flash
AA Intelligence Index (vendor-cited)59unknown effort
All 1 recorded result & sources

59 · raw 59 index

Headline · unknown effort · Peer vendor

Source/record date: Not recorded

from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); Gemini 3.8 Flash (high)

https://huggingface.co/inclusionAI/Ling-3.0-flash-VL
Gray Swan IPI5.5%unknown effort
All 1 recorded result & sources

5.5% · raw 5.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon Gray Swan chart (K=15 attack success rate)

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis41high effort
All 3 recorded results & sources

41 · raw 41 index

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gemini-3-8-flash

40 · raw 40 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gemini-3-8-flash-medium

33 · raw 33 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gemini-3-8-flash-low
Cost per Intelligence Index task · Artificial Analysis$1.24high effort
All 2 recorded results & sources

$1.24 · raw 1.24 USD

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gemini-3-8-flash

$0.93 · raw 0.93 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/gemini-3-8-flash-medium
Output speed · Artificial Analysis249 tok/shigh effort
All 1 recorded result & sources

249 tok/s · raw 249 tok/s

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/gemini-3-8-flash
Vals Index · Vals AI54.8%unknown effort
All 1 recorded result & sources

54.8% · raw 54.83 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $5.73.

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench18 fixeshigh effort
All 1 recorded result & sources

18 fixes · raw 18 fixes

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Antigravity CLI; effort high; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
CUDA board % of roofline · KernelBench (community board)16.7%high effort
All 1 recorded result & sources

16.7% · raw 16.7 %

Headline · high effort · Independent evaluator

Source/record date: 2026-09-23

Gemini 3.8 Flash (High) board; CUDA 3/4; Mega 2.74×

https://kernelbench.com/models/gemini-3.8-flash-high
GDPval-AA Elo · Artificial Analysis1412high effort
All 1 recorded result & sources

1412 · raw 1412 Elo

Headline · high effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; Gemini 3.8 Flash (high)

https://artificialanalysis.ai/evaluations/gdpval-aa
ARC-AGI-1 · ARC Prize98.5%high effort
All 3 recorded results & sources

98.5% · raw 98.5 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/google-gemini-3-8-flash

97.5% · raw 97.5 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/google-gemini-3-8-flash

90.5% · raw 90.5 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/google-gemini-3-8-flash
ARC-AGI-2 · ARC Prize89.2%high effort
All 3 recorded results & sources

89.2% · raw 89.2 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/google-gemini-3-8-flash

82.9% · raw 82.9 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/google-gemini-3-8-flash

77.5% · raw 77.5 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/google-gemini-3-8-flash
ARC-AGI-3 (Standard) · ARC Prize10.4%high effort
All 3 recorded results & sources

10.4% · raw 10.37 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry). Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/google-gemini-3-8-flash

4% · raw 3.99 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/google-gemini-3-8-flash

6% · raw 5.99 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Standard (notes carry).

https://arcprize.org/results/google-gemini-3-8-flash
ARC-AGI-3 · ARC Prize35%high effort
All 3 recorded results & sources

35% · raw 35 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/google-gemini-3-8-flash

24.2% · raw 24.2 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/google-gemini-3-8-flash

15.1% · raw 15.07 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Harness: Provider Adapter.

https://arcprize.org/results/google-gemini-3-8-flash
Money gain · Andon Labs$4593.79unknown effort
All 1 recorded result & sources

$4593.79 · raw 4593.79 $

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Vending-Bench 2 net gain = final_value in the page public vb2 data module minus $500 starting balance. Full 66-model source checked.

https://andonlabs.com/evals/vending-bench-2
Blueprint Bench · Andon Labs38.6%unknown effort
All 1 recorded result & sources

38.6% · raw 38.6 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Blueprint-Bench 2 connectivity similarity; published fractional score multiplied by 100.

https://andonlabs.com/evals/blueprint-bench-2
Average Score · WeirdML v37.4%high effort
All 1 recorded result & sources

7.4% · raw 7.38 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Gemini 3.8 Flash (high); harness gemini_cli 0.59.0; values from prepared data JSON; raw 0.073787; official 80/20 aggregate (area 500k-50M tokens + final best)

https://htihle.github.io/weirdml.html
Final Best Score · WeirdML v314.6%high effort
All 1 recorded result & sources

14.6% · raw 14.57 %

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Gemini 3.8 Flash (high); harness gemini_cli 0.59.0; values from prepared data JSON; raw 0.145741; mean final best effective score

https://htihle.github.io/weirdml.html
Cost / Run · WeirdML v3$3.94high effort
All 1 recorded result & sources

$3.94 · raw 3.94 USD

Headline · high effort · Independent evaluator

Source/record date: 2026-10-02

WeirdML variant Gemini 3.8 Flash (high); harness gemini_cli 0.59.0; values from prepared data JSON; mean API cost per run, same task weighting as scores

https://htihle.github.io/weirdml.html
Vibe Code Bench v1.1 · Vals AI78.7%unknown effort
All 1 recorded result & sources

78.7% · raw 78.65 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $6.87.

https://www.vals.ai/benchmarks/vibe-code

Read how we select and source scores or the comparison guide.