Gemini 3.6 Flash
currentGemini 3.6 Flash is Google's fast workhorse for coding, multimodal knowledge work, and multi-step agents. It improves quality and token efficiency over 3.5 Flash while cutting the output-token price, though it remains much more expensive than Flash-Lite for simple, high-volume jobs.
Context window
1.05M tokens
Maximum output
66K tokens
Input price
$1.50 / 1M
Output price
$7.50 / 1M
Best for
- Multi-step coding and agent workflows
- Multimodal document, chart, and spatial analysis
- Computer-use and tool-calling applications
- Production workloads that need strong quality with lower token use
Not ideal for
- The cheapest possible bulk extraction or classification
- Workloads that require deterministic sampling controls
- Native image, audio, or video generation
- Applications that cannot tolerate occasional model latency or timeouts
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
The comms console chimed with a station message routed through Europa’s orbital relay. Due to the system’s clock offset, the timestamp read eleven minutes in the future: "Do not open the greenhouse door."
Jon looked at the empty third bunk. "Elena is still inside servicing the pressure seal, right?"
"Just a software loop," Marcus scoffed, punching the manual release.
The heavy airlock unsealed. Instantly, explosive decompression tore through the station, sweeping Marcus into the freezing vacuum.
The message was not a prophetic warning to protect them from what lurked inside, but an automated safety lockout triggered when the glass shattered.
Run the same prompt yourself
Open Gemini 3.6 Flash in Multichat with the evaluation prompt prefilled.
What the response revealed
The response contains exactly 100 whitespace-delimited words.
It outputs only the story and avoids aliens, time travel, hallucinations, and simulations.
The orbital relay and clock offset explain the apparent future timestamp, while a breached greenhouse and sealed door support the decompression warning.
The relay route, clock offset, empty bunk, pressure-seal work, and manual release all prepare the resolution without naming it outright.
The final sentence reframes the message from a warning about something inside the greenhouse to an automated lockout protecting the rest of the station.
The prose is economical and vivid, with a complete setup, consequence, and recontextualizing payoff inside the word limit.
The model receives the prompt without web access or external tools. Gemini 3.6 Flash is run through OpenRouter at medium reasoning effort. We preserve the response as generated apart from display rendering.
Performance
Gemini 3.6 Flash benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.
DeepSWE
49%
agentic software engineering · Google3
MLE-Bench
63.9%
machine-learning engineering · Google3
OSWorld-Verified
83%
computer use · Google3
GDPval-AA v2
1,421 Elo
knowledge work · Google3
Family position
Gemini 3.6 positioning
Gemini 3.6 Flash is Google's workhorse model for agentic execution, coding, knowledge work, and multimodal reasoning. Google says it uses fewer reasoning steps, turns, tool calls, and output tokens than 3.5 Flash while improving quality on several launch evaluations.
Gemini 3.6 Flash
This modelEfficient agentic workhorse
$1.50 input
$7.50 output
Gemini 3.5 Flash
Previous Flash generation
$1.50 input
$9 output
Gemini 3.5 Flash-Lite
High-volume efficiency
$0.30 input
$2.50 output
API pricing
Per million text tokens
Input
$1.50
Cached input
$0.15
Cache write
$1.50
Output
$7.50
Where Gemini 3.6 Flash stands out
More efficient agent loops
Google reports 17% fewer output tokens on the Artificial Analysis Index than 3.5 Flash, alongside fewer reasoning steps and tool calls in multi-step workflows.
Stronger coding and computer use
Launch results show gains on DeepSWE, MLE-Bench, and OSWorld-Verified, while Google highlights fewer unwanted edits and execution loops.
Broad multimodal tool support
The model accepts text, images, video, audio, and PDFs, with built-in support for code execution, search grounding, file search, structured outputs, function calling, and preview computer use.
Limitations to know
Still priced above Flash-Lite
At $1.50 per million input tokens and $7.50 per million output tokens, simple high-volume work can cost substantially more than with Gemini 3.5 Flash-Lite.
Sampling controls are changing
Google deprecates temperature, top_p, and top_k for this generation, so migrations may need stronger system instructions and updated request configuration.
General model limitations remain
Google's model card notes possible hallucinations and occasional slowness or timeouts. A large context window also does not guarantee perfect recall at maximum length.
Capabilities and specifications
Knowledge cutoff
Mar 1, 2026
Inputs
text, image, video, audio, PDF
Reasoning
Supported
Default effort
medium
Supported API features
Frequently Asked Questions
Sources
- 1Gemini 3.6 Flash
Google AI for Developers · official documentation
- 2Gemini Developer API pricing
Google AI for Developers · official pricing documentation
- 3Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google · official launch announcement
- 4Gemini 3.6 Flash model card
Google DeepMind · official model card
- 5PromptHQ model registry
PromptHQ · internal product configuration
Compare leading AI assistants
See how Gemini 3.6 Flash handles your own work.
Try Gemini 3.6 Flash in MultichatAvailable on PromptHQ Plus, Max