Gemini 3.1 Pro
currentGemini 3.1 Pro is Google's preview reasoning model for difficult multimodal, coding, and agentic work. Its one-million-token context and broad input support make it especially flexible for mixed media and large corpora, though newer Gemini 3.5 Flash can be faster and stronger on several agent benchmarks.
Context window
1.05M tokens
Maximum output
66K tokens
Input price
$2 / 1M
Output price
$12 / 1M
Best for
- Complex multimodal analysis across text, images, audio, video, and PDFs
- Advanced reasoning and planning
- Large-context document and code analysis
- Structured outputs and tool-rich Gemini API workflows
Not ideal for
- Workloads that require a stable non-preview model identifier
- Tasks where Gemini 3.5 Flash's speed is more valuable
- Native image or audio generation
- Prompts relying on post-January-2025 knowledge without search
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
The console chimed. Maya frowned at the text from their own station ID: Do not open the greenhouse door.
"Timestamp is 11:42," Ren said, fully zipping his heavy vacuum suit.
"It's only 11:31. Network bounce from the orbital relay?"
"Just a server sync error," Maya dismissed, checking a flashing pressure
Run the same prompt yourself
Open Gemini 3.1 Pro in Multichat with the evaluation prompt prefilled.
What the response revealed
The response stops after 50 words rather than reaching exactly 100.
Although it avoids the prohibited explanations, the response is visibly truncated and does not deliver a complete story.
The response ends mid-sentence before explaining the warning, timestamp, or opened door.
The fast timestamp, vacuum suit, and relay suggestion could function as clues, but the missing resolution prevents confirmation.
There is no final sentence or reversal because the output is incomplete.
The opening is economical and intriguing, but the generation terminates before it becomes a story.
The model receives the prompt without web access or external tools. Gemini 3.1 Pro is run through OpenRouter at medium reasoning effort. We preserve the response as generated apart from display rendering.
Performance
Gemini 3.1 Pro benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.
ARC-AGI-2
77.1%
abstract reasoning · Google DeepMind3
SWE-Bench Pro
54.2%
software engineering · Google DeepMind3
MCP Atlas
78.2%
agentic tool use · Google DeepMind3
MMMU-Pro, no tools
80.5%
multimodal reasoning · Google DeepMind3
Humanity's Last Exam
44.4%
academic reasoning · Google DeepMind3
Family position
Gemini 3.1 positioning
Gemini 3.1 Pro was Google's most capable general Gemini model at launch. Gemini 3.5 Flash later surpassed it on many agentic benchmarks while targeting much higher throughput.
Gemini 3.1 Pro
This modelDeep reasoning preview
$2 input
$12 output
Gemini 3.5 Flash
Fast agentic model
$1.50 input
$9 output
Gemini 3.1 Flash Lite
Low-cost volume
$0.25 input
$1.50 output
API pricing
Per million text tokens
Input
$2
Cached input
$0.20
Cache write
$2
Output
$12
Where Gemini 3.1 Pro stands out
Built for demanding work
Complex multimodal analysis across text, images, audio, video, and PDFs are central to the model's positioning, rather than an incidental capability.
Long-context capacity
The published context window is 1,048,576 tokens, making the model a candidate for large documents, repositories, and sustained agent state.
Reasoning and tools
Reasoning is supported with low, medium, high provider setting(s), and the model can participate in tool-using workflows through its available API surface.
Limitations to know
Benchmarks are configuration-sensitive
Scores can move substantially with the harness, tool access, effort setting, token budget, and evaluator. Treat the table as evidence, not a universal ranking.
Context size is not guaranteed recall
A large advertised window does not mean every detail is retrieved reliably at maximum length. Validate representative long-context workloads before deployment.
Product access differs from model capability
PromptHQ and gateway limits may expose fewer modalities, tools, or tokens than the provider's first-party API.
Capabilities and specifications
Knowledge cutoff
Jan 1, 2025
Inputs
text, image, video, audio, PDF
Reasoning
Supported
Default effort
medium
Supported API features
Frequently Asked Questions
Sources
- 1Gemini 3.1 Pro Preview
Google AI for Developers · official documentation
- 2Gemini 3.1 Pro: A smarter model for your most complex tasks
Google · official launch post
- 3Gemini 3.5 Flash model card
Google DeepMind · official model card
- 4Gemini 3.1 Pro API pricing and availability
OpenRouter · gateway model page
- 5PromptHQ model registry
PromptHQ · internal product configuration
Compare leading AI assistants
See how Gemini 3.1 Pro handles your own work.
Try Gemini 3.1 Pro in MultichatAvailable on PromptHQ Plus, Max