Gemini 3.5 Flash
currentGemini 3.5 Flash combines frontier-level agent and coding results with high throughput and broad multimodal input support. It is a compelling general-purpose model when speed and tool use matter, but its output-token price is higher than earlier Flash generations and long-context recall still degrades at the extreme end.
Context window
1.05M tokens
Maximum output
66K tokens
Input price
$1.50 / 1M
Output price
$9 / 1M
Best for
- Fast coding and multi-step agent workflows
- Multimodal analysis of text, images, audio, video, and PDFs
- Computer-use and tool-calling applications
- High-throughput tasks that still need adjustable reasoning
Not ideal for
- The cheapest possible bulk classification or extraction
- Perfect recall across an entire one-million-token prompt
- Native image or audio generation
- Use cases that need a long, current knowledge cutoff without grounding
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
The backup generator hummed, its unsynchronized clock blinking eleven minutes fast. Coughing, Elena ignored her burning throat and forced the emergency valve.
A message from their own automated diagnostic terminal had just arrived, baffling the three-person Europa crew with its future timestamp: Do not open the greenhouse door.
Fearing a botany biohazard, they had hesitated. But with the main
Run the same prompt yourself
Open Gemini 3.5 Flash in Multichat with the evaluation prompt prefilled.
What the response revealed
The response stops after 59 words rather than reaching exactly 100.
It avoids the prohibited concepts but ends mid-sentence and therefore does not provide the requested complete story.
Clock drift and physical symptoms begin an explanation, but the output truncates before connecting the warning and opened door.
The unsynchronized clock, cough, and burning throat are promising clues without a completed payoff.
The response has no ending or recontextualizing final sentence.
The setup is vivid and efficient, but cannot be judged as a complete miniature story.
The model receives the prompt without web access or external tools. Gemini 3.5 Flash is run through OpenRouter at medium reasoning effort. We preserve the response as generated apart from display rendering.
Performance
Gemini 3.5 Flash benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.
Terminal-Bench 2.1
76.2%
terminal and agentic coding · Google DeepMind2
SWE-Bench Pro
55.1%
software engineering · Google DeepMind2
MCP Atlas
83.6%
agentic tool use · Google DeepMind2
OSWorld-Verified
78.4%
computer use · Google DeepMind2
MMMU-Pro, no tools
83.6%
multimodal reasoning · Google DeepMind2
ARC-AGI-2
72.1%
abstract reasoning · Google DeepMind2
Family position
Gemini 3.5 positioning
Gemini 3.5 Flash is the speed-oriented member of the 3.5 generation. Google reports that it outperforms 3.1 Pro on most of its agentic launch suite while responding much faster.
Gemini 3.5 Flash
This modelFast frontier agent
$1.50 input
$9 output
Gemini 3.1 Pro
Deep reasoning preview
$2 input
$12 output
Gemini 3.1 Flash Lite
Low-cost volume
$0.25 input
$1.50 output
API pricing
Per million text tokens
Input
$1.50
Cached input
$0.15
Cache write
$1.50
Output
$9
Where Gemini 3.5 Flash stands out
Built for demanding work
Fast coding and multi-step agent workflows are central to the model's positioning, rather than an incidental capability.
Long-context capacity
The published context window is 1,048,576 tokens, making the model a candidate for large documents, repositories, and sustained agent state.
Reasoning and tools
Reasoning is supported with minimal, low, medium, high provider setting(s), and the model can participate in tool-using workflows through its available API surface.
Limitations to know
Benchmarks are configuration-sensitive
Scores can move substantially with the harness, tool access, effort setting, token budget, and evaluator. Treat the table as evidence, not a universal ranking.
Context size is not guaranteed recall
A large advertised window does not mean every detail is retrieved reliably at maximum length. Validate representative long-context workloads before deployment.
Product access differs from model capability
PromptHQ and gateway limits may expose fewer modalities, tools, or tokens than the provider's first-party API.
Capabilities and specifications
Knowledge cutoff
Jan 1, 2025
Inputs
text, image, video, audio, PDF
Reasoning
Supported
Default effort
medium
Supported API features
Frequently Asked Questions
Sources
- 1Gemini 3.5 Flash
Google AI for Developers · official documentation
- 2Gemini 3.5 Flash model card
Google DeepMind · official model card
- 3Gemini 3.5 Flash API pricing and availability
OpenRouter · gateway model page
- 4PromptHQ model registry
PromptHQ · internal product configuration
Compare leading AI assistants
See how Gemini 3.5 Flash handles your own work.
Try Gemini 3.5 Flash in MultichatAvailable on PromptHQ Plus, Max