Model guide / Google

Gemini 3.5 Flash-Lite

current

Gemini 3.5 Flash-Lite is Google's fastest and lowest-cost 3.5 model, built for high-volume extraction, document processing, translation, routing, and subagent work. It delivers a large quality jump over 3.1 Flash-Lite, but complex tasks may need higher reasoning effort and 3.6 Flash remains the stronger workhorse.

Released Jul 21, 2026ProprietaryClosed weightsVerified Jul 21, 2026

Context window

1.05M tokens

Maximum output

66K tokens

Input price

$0.30 / 1M

Output price

$2.50 / 1M

Best for

  • High-volume extraction, classification, and routing
  • Low-latency document parsing and translation
  • Cost-sensitive subagent and tool-calling workflows
  • Multimodal batch processing at production scale

Not ideal for

  • The hardest coding or multi-step reasoning tasks
  • Native computer-use automation through the Gemini API
  • Native image, audio, or video generation
  • Workloads that require deterministic sampling controls
PromptHQ evaluation

PromptHQ 100-word challenge

We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.

Gemini 3.5 Flash Lite100 words

Write a science-fiction mystery in exactly 100 words.

A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:

"Do not open the greenhouse door."

The door is opened anyway.

Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.

Output only the story. Do not explain it.

Station Vanguard on Europa hummed as Jupiter induced heavy magnetospheric static, scrambling local router packets. Three researchers stared at the screen. A message from their own terminal blinked, timestamped eleven minutes ahead.

Do not open the greenhouse door.

Dr Vance nervously wiped sweat, eyeing the erratic digital wall clock. The buffer delay explained the strange timestamp.

Curiosity won. Vance overrode the lock and pulled the heavy seal.

Instantly, emergency klaxons wailed.

The warning had not been about a threat inside the room. It was the automated life support system begging them not to vent their precious oxygen supply right now.

Run the same prompt yourself

Open Gemini 3.5 Flash Lite in Multichat with the evaluation prompt prefilled.

Try this prompt

What the response revealed

Exact word countpass

The response contains exactly 100 whitespace-delimited words.

Constraint followingpass

It outputs only the story and avoids aliens, time travel, hallucinations, and simulations.

Causal coherencemixed

Static, packet delay, and an erratic clock can explain the timestamp, but the story does not fully establish what caused life support to issue the door warning before Vance overrode the lock.

Foreshadowingmixed

The networking interference and clock are useful timestamp clues, while the oxygen-loss resolution receives little advance setup beyond the heavy seal and lock.

Endingpass

The final lines reframe the warning from danger inside the greenhouse to preservation of the station's oxygen supply.

Writingpass

The response is concise, readable, and complete, with an effective turn in its closing sentences.

The model receives the prompt without web access or external tools. Gemini 3.5 Flash-Lite is run through OpenRouter at medium reasoning effort to match PromptHQ's default. We preserve the response as generated apart from display rendering.

Performance

Gemini 3.5 Flash-Lite benchmarks

Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.

Terminal-Bench 2.1

54%

terminal and agentic coding · Google3

GDM-MRCR v2

72.2%

long-context retrieval · Google3

GDPval-AA v2

1,140 Elo

real-world task execution · Google3

SWE-Bench Pro

54.2%

software engineering · Google3

OSWorld-Verified

74%

computer use evaluation · Google3

Family position

Gemini 3.5 positioning

Gemini 3.5 Flash-Lite is the fastest, lowest-cost model in Google's 3.5 family. It targets latency-sensitive, high-throughput production tasks and improves substantially on 3.1 Flash-Lite across agents, coding, long context, and real-world task execution.

Gemini 3.6 Flash

Efficient agentic workhorse

$1.50 input

$7.50 output

Gemini 3.5 Flash

Frontier Flash model

$1.50 input

$9 output

Gemini 3.5 Flash-Lite

This model

High-volume efficiency

$0.30 input

$2.50 output

API pricing

Per million text tokens

Input

$0.30

Cached input

$0.03

Cache write

$0.30

Output

$2.50

Published prices are standard token rates and output pricing includes thinking tokens. Batch, flex, priority inference, cache storage, grounding, and other tools have separate rates.

Where Gemini 3.5 Flash-Lite stands out

High throughput at low cost

Google positions 3.5 Flash-Lite as its fastest 3.5-class model and reports 350 output tokens per second on the Artificial Analysis Index.

Large generational quality gains

Launch comparisons show substantial improvements over 3.1 Flash-Lite on Terminal-Bench 2.1, GDM-MRCR v2, and GDPval-AA v2.

Broad multimodal inputs and tools

The model accepts text, images, video, audio, and PDFs, and supports code execution, search grounding, file search, structured outputs, and function calling.

Limitations to know

Complex work can need more thinking

Google recommends minimal effort for high-volume extraction but medium or high effort for autonomous subagents, tool calls, code execution, and multi-step reasoning.

Computer Use is not exposed

Despite a strong OSWorld-Verified evaluation result, Google's detailed Gemini API model page currently marks Computer Use as unsupported for 3.5 Flash-Lite.

General model limitations remain

Google's model card notes possible hallucinations and occasional slowness or timeouts. A large context window also does not guarantee perfect recall at maximum length.

Capabilities and specifications

Knowledge cutoff

Mar 1, 2026

Inputs

text, image, video, audio, PDF

Reasoning

Supported

Default effort

medium

Supported API features

Streaming
Function calling
Structured outputs
Web search
File search
Code interpreter
Tool search
Web and file search minimal, low, medium, high effort

Frequently Asked Questions

Sources

  1. 1
    Gemini 3.5 Flash-Lite

    Google AI for Developers · official documentation

  2. 2
    Gemini Developer API pricing

    Google AI for Developers · official pricing documentation

  3. 3
  4. 4
    Gemini 3.5 Flash-Lite model card

    Google DeepMind · official model card

  5. 5
    PromptHQ model registry

    PromptHQ · internal product configuration

Compare leading AI assistants

See how Gemini 3.5 Flash-Lite handles your own work.

Try Gemini 3.5 Flash-Lite in Multichat

Available on PromptHQ Almostfree, Plus, Max