Gemini 3.5 Flash-Lite
currentGemini 3.5 Flash-Lite is Google's fastest and lowest-cost 3.5 model, built for high-volume extraction, document processing, translation, routing, and subagent work. It delivers a large quality jump over 3.1 Flash-Lite, but complex tasks may need higher reasoning effort and 3.6 Flash remains the stronger workhorse.
Context window
1.05M tokens
Maximum output
66K tokens
Input price
$0.30 / 1M
Output price
$2.50 / 1M
Best for
- High-volume extraction, classification, and routing
- Low-latency document parsing and translation
- Cost-sensitive subagent and tool-calling workflows
- Multimodal batch processing at production scale
Not ideal for
- The hardest coding or multi-step reasoning tasks
- Native computer-use automation through the Gemini API
- Native image, audio, or video generation
- Workloads that require deterministic sampling controls
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
Station Vanguard on Europa hummed as Jupiter induced heavy magnetospheric static, scrambling local router packets. Three researchers stared at the screen. A message from their own terminal blinked, timestamped eleven minutes ahead.
Do not open the greenhouse door.
Dr Vance nervously wiped sweat, eyeing the erratic digital wall clock. The buffer delay explained the strange timestamp.
Curiosity won. Vance overrode the lock and pulled the heavy seal.
Instantly, emergency klaxons wailed.
The warning had not been about a threat inside the room. It was the automated life support system begging them not to vent their precious oxygen supply right now.
Run the same prompt yourself
Open Gemini 3.5 Flash Lite in Multichat with the evaluation prompt prefilled.
What the response revealed
The response contains exactly 100 whitespace-delimited words.
It outputs only the story and avoids aliens, time travel, hallucinations, and simulations.
Static, packet delay, and an erratic clock can explain the timestamp, but the story does not fully establish what caused life support to issue the door warning before Vance overrode the lock.
The networking interference and clock are useful timestamp clues, while the oxygen-loss resolution receives little advance setup beyond the heavy seal and lock.
The final lines reframe the warning from danger inside the greenhouse to preservation of the station's oxygen supply.
The response is concise, readable, and complete, with an effective turn in its closing sentences.
The model receives the prompt without web access or external tools. Gemini 3.5 Flash-Lite is run through OpenRouter at medium reasoning effort to match PromptHQ's default. We preserve the response as generated apart from display rendering.
Performance
Gemini 3.5 Flash-Lite benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.
Terminal-Bench 2.1
54%
terminal and agentic coding · Google3
GDM-MRCR v2
72.2%
long-context retrieval · Google3
GDPval-AA v2
1,140 Elo
real-world task execution · Google3
SWE-Bench Pro
54.2%
software engineering · Google3
OSWorld-Verified
74%
computer use evaluation · Google3
Family position
Gemini 3.5 positioning
Gemini 3.5 Flash-Lite is the fastest, lowest-cost model in Google's 3.5 family. It targets latency-sensitive, high-throughput production tasks and improves substantially on 3.1 Flash-Lite across agents, coding, long context, and real-world task execution.
Gemini 3.6 Flash
Efficient agentic workhorse
$1.50 input
$7.50 output
Gemini 3.5 Flash
Frontier Flash model
$1.50 input
$9 output
Gemini 3.5 Flash-Lite
This modelHigh-volume efficiency
$0.30 input
$2.50 output
API pricing
Per million text tokens
Input
$0.30
Cached input
$0.03
Cache write
$0.30
Output
$2.50
Where Gemini 3.5 Flash-Lite stands out
High throughput at low cost
Google positions 3.5 Flash-Lite as its fastest 3.5-class model and reports 350 output tokens per second on the Artificial Analysis Index.
Large generational quality gains
Launch comparisons show substantial improvements over 3.1 Flash-Lite on Terminal-Bench 2.1, GDM-MRCR v2, and GDPval-AA v2.
Broad multimodal inputs and tools
The model accepts text, images, video, audio, and PDFs, and supports code execution, search grounding, file search, structured outputs, and function calling.
Limitations to know
Complex work can need more thinking
Google recommends minimal effort for high-volume extraction but medium or high effort for autonomous subagents, tool calls, code execution, and multi-step reasoning.
Computer Use is not exposed
Despite a strong OSWorld-Verified evaluation result, Google's detailed Gemini API model page currently marks Computer Use as unsupported for 3.5 Flash-Lite.
General model limitations remain
Google's model card notes possible hallucinations and occasional slowness or timeouts. A large context window also does not guarantee perfect recall at maximum length.
Capabilities and specifications
Knowledge cutoff
Mar 1, 2026
Inputs
text, image, video, audio, PDF
Reasoning
Supported
Default effort
medium
Supported API features
Frequently Asked Questions
Sources
- 1Gemini 3.5 Flash-Lite
Google AI for Developers · official documentation
- 2Gemini Developer API pricing
Google AI for Developers · official pricing documentation
- 3Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google · official launch announcement
- 4Gemini 3.5 Flash-Lite model card
Google DeepMind · official model card
- 5PromptHQ model registry
PromptHQ · internal product configuration
Compare leading AI assistants
See how Gemini 3.5 Flash-Lite handles your own work.
Try Gemini 3.5 Flash-Lite in MultichatAvailable on PromptHQ Almostfree, Plus, Max