Model guide / Google

Gemini 3.6 Flash

current

Gemini 3.6 Flash is Google's fast workhorse for coding, multimodal knowledge work, and multi-step agents. It improves quality and token efficiency over 3.5 Flash while cutting the output-token price, though it remains much more expensive than Flash-Lite for simple, high-volume jobs.

Released Jul 21, 2026ProprietaryClosed weightsVerified Jul 21, 2026

Context window

1.05M tokens

Maximum output

66K tokens

Input price

$1.50 / 1M

Output price

$7.50 / 1M

Best for

  • Multi-step coding and agent workflows
  • Multimodal document, chart, and spatial analysis
  • Computer-use and tool-calling applications
  • Production workloads that need strong quality with lower token use

Not ideal for

  • The cheapest possible bulk extraction or classification
  • Workloads that require deterministic sampling controls
  • Native image, audio, or video generation
  • Applications that cannot tolerate occasional model latency or timeouts
PromptHQ evaluation

PromptHQ 100-word challenge

We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.

Gemini 3.6 Flash100 words

Write a science-fiction mystery in exactly 100 words.

A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:

"Do not open the greenhouse door."

The door is opened anyway.

Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.

Output only the story. Do not explain it.

The comms console chimed with a station message routed through Europa’s orbital relay. Due to the system’s clock offset, the timestamp read eleven minutes in the future: "Do not open the greenhouse door."

Jon looked at the empty third bunk. "Elena is still inside servicing the pressure seal, right?"

"Just a software loop," Marcus scoffed, punching the manual release.

The heavy airlock unsealed. Instantly, explosive decompression tore through the station, sweeping Marcus into the freezing vacuum.

The message was not a prophetic warning to protect them from what lurked inside, but an automated safety lockout triggered when the glass shattered.

Run the same prompt yourself

Open Gemini 3.6 Flash in Multichat with the evaluation prompt prefilled.

Try this prompt

What the response revealed

Exact word countpass

The response contains exactly 100 whitespace-delimited words.

Constraint followingpass

It outputs only the story and avoids aliens, time travel, hallucinations, and simulations.

Causal coherencepass

The orbital relay and clock offset explain the apparent future timestamp, while a breached greenhouse and sealed door support the decompression warning.

Foreshadowingpass

The relay route, clock offset, empty bunk, pressure-seal work, and manual release all prepare the resolution without naming it outright.

Endingpass

The final sentence reframes the message from a warning about something inside the greenhouse to an automated lockout protecting the rest of the station.

Writingpass

The prose is economical and vivid, with a complete setup, consequence, and recontextualizing payoff inside the word limit.

The model receives the prompt without web access or external tools. Gemini 3.6 Flash is run through OpenRouter at medium reasoning effort. We preserve the response as generated apart from display rendering.

Performance

Gemini 3.6 Flash benchmarks

Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.

DeepSWE

49%

agentic software engineering · Google3

MLE-Bench

63.9%

machine-learning engineering · Google3

OSWorld-Verified

83%

computer use · Google3

GDPval-AA v2

1,421 Elo

knowledge work · Google3

Family position

Gemini 3.6 positioning

Gemini 3.6 Flash is Google's workhorse model for agentic execution, coding, knowledge work, and multimodal reasoning. Google says it uses fewer reasoning steps, turns, tool calls, and output tokens than 3.5 Flash while improving quality on several launch evaluations.

Gemini 3.6 Flash

This model

Efficient agentic workhorse

$1.50 input

$7.50 output

Gemini 3.5 Flash

Previous Flash generation

$1.50 input

$9 output

Gemini 3.5 Flash-Lite

High-volume efficiency

$0.30 input

$2.50 output

API pricing

Per million text tokens

Input

$1.50

Cached input

$0.15

Cache write

$1.50

Output

$7.50

Published prices are standard text-token rates and output pricing includes thinking tokens. Batch, flex, priority inference, cache storage, grounding, and other tools have separate rates.

Where Gemini 3.6 Flash stands out

More efficient agent loops

Google reports 17% fewer output tokens on the Artificial Analysis Index than 3.5 Flash, alongside fewer reasoning steps and tool calls in multi-step workflows.

Stronger coding and computer use

Launch results show gains on DeepSWE, MLE-Bench, and OSWorld-Verified, while Google highlights fewer unwanted edits and execution loops.

Broad multimodal tool support

The model accepts text, images, video, audio, and PDFs, with built-in support for code execution, search grounding, file search, structured outputs, function calling, and preview computer use.

Limitations to know

Still priced above Flash-Lite

At $1.50 per million input tokens and $7.50 per million output tokens, simple high-volume work can cost substantially more than with Gemini 3.5 Flash-Lite.

Sampling controls are changing

Google deprecates temperature, top_p, and top_k for this generation, so migrations may need stronger system instructions and updated request configuration.

General model limitations remain

Google's model card notes possible hallucinations and occasional slowness or timeouts. A large context window also does not guarantee perfect recall at maximum length.

Capabilities and specifications

Knowledge cutoff

Mar 1, 2026

Inputs

text, image, video, audio, PDF

Reasoning

Supported

Default effort

medium

Supported API features

Streaming
Function calling
Structured outputs
Web search
File search
Code interpreter
Computer use
Tool search
Web and file search minimal, low, medium, high effort

Frequently Asked Questions

Sources

  1. 1
    Gemini 3.6 Flash

    Google AI for Developers · official documentation

  2. 2
    Gemini Developer API pricing

    Google AI for Developers · official pricing documentation

  3. 3
  4. 4
    Gemini 3.6 Flash model card

    Google DeepMind · official model card

  5. 5
    PromptHQ model registry

    PromptHQ · internal product configuration

Compare leading AI assistants

See how Gemini 3.6 Flash handles your own work.

Try Gemini 3.6 Flash in Multichat

Available on PromptHQ Plus, Max