Claude 5 Opus
currentClaude 5 Opus is Anthropic's newest Opus model for complex coding, agents, and enterprise work, and it launches at the top of the Artificial Analysis Intelligence Index while charging half of Fable 5's token price. It keeps the one-million-token context and adaptive thinking of Opus 4.8, but early hands-on reviews and independent benchmarking both flag verbosity and slow output as real costs.
Context window
1M tokens
Maximum output
128K tokens
Input price
$5 / 1M
Output price
$25 / 1M
Best for
- Complex coding and repository-wide changes
- Long-horizon agents and computer-use workflows
- Frontier-level reasoning at half Fable 5's token price
- Enterprise analysis across very large document sets
Not ideal for
- Latency-sensitive chat and high-volume production traffic
- Workloads that need terse output without prompt tuning
- Routine tasks where Sonnet 5 or Haiku 4.5 are sufficient
- Self-hosting, fine-tuning, or native audio and video generation
Independent perspective
What experts are saying
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
Kepler logged the message at 04:12. Sent from Station Europa. Timestamp 04:23.
Do not open the greenhouse door.
"Clock drift," Renn said. The comms array had been fighting the ice-shelf interference for weeks. Every log entry ran minutes ahead of the wall panel.
Ahn was already in the greenhouse, harvesting. She hadn't heard.
Kepler opened the door. Warm air, wet soil, Ahn humming.
At 04:23 the pressure alarm sounded — hull breach, sector nine, sealed off from them by the door now standing open.
Renn found the outgoing draft in Kepler's own queue, composed at 04:24.
He hadn't been warned. He'd been confessing.
Run the same prompt yourself
Open Claude 5 Opus in Multichat with the evaluation prompt prefilled.
What the response revealed
The response contains 103 whitespace-delimited words, three over the required count.
It outputs only the story and avoids aliens, time travel, hallucinations, and simulations.
Clock drift on the comms array explains the future timestamp, but the draft composed at 04:24 arriving in a log at 04:12 is left for the reader to reconcile, and the breach being 'sealed off from them by the door now standing open' is physically unclear.
The logs running minutes ahead of the wall panel and Kepler being the one who logged the message both quietly prepare the resolution.
The last two lines convert the message from an external warning into Kepler's own confession, changing who sent it and why.
Clipped, concrete sentences carry the tension, and the sensory beat inside the greenhouse lands before the alarm.
The model receives the prompt without web access or external tools. Claude 5 Opus is run through Claude CLI at medium reasoning effort. We preserve the response as generated apart from display rendering.
Performance
Claude 5 Opus benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches. Several launch figures are Anthropic's own internal evaluations and had no independent replication at the time of writing.
Artificial Analysis Intelligence Index
61
aggregate intelligence · Artificial Analysis
ARC-AGI 3
30.2%
novel problem solving · Anthropic
Organic chemistry evaluation
10.2
life sciences · Anthropic
Protein prediction tasks
7.7
life sciences · Anthropic
Automated behavioral audit, misaligned behavior
2.3
alignment · Anthropic
Output speed
45.7
latency · Artificial Analysis
Family position
Claude 5 positioning
Opus 5 is Anthropic's recommended default for complex agentic coding and enterprise work. Fable 5 remains the highest-capability option and the choice for the most demanding long-running agents, while Sonnet 5 stays the cheaper route for routine work. The pitch for Opus 5 is near-Fable capability at Opus pricing.
Claude Fable 5
Highest general capability
$10 input
$50 output
Claude Opus 5
Complex agentic coding and enterprise
$5 input
$25 output
Claude Opus 4.8
Previous-generation Opus
$5 input
$25 output
Claude Sonnet 5
Balanced speed and intelligence
$3 input
$15 output
API pricing
Per million text tokens
Input
$5
Cached input
$0.50
Cache write
$6.25
Output
$25
Where Opus 5 stands out
Frontier capability at Opus pricing
Independent benchmarking placed it first on the Artificial Analysis Intelligence Index at launch while it bills $5 per million input tokens and $25 per million output tokens, the same rates as Opus 4.8 and half of Fable 5.
Novel-problem reasoning
The reported 30.2% on ARC-AGI 3, a benchmark built around unfamiliar interactive environments, is several times the best previously published result and is the clearest evidence of a genuine capability jump rather than benchmark saturation.
Efficient at lower effort
Anthropic reports matching Opus 4.8's maximum-effort output quality while generating about 26% fewer tokens on average, which makes the low and medium effort levels worth testing before defaulting to the top of the ladder.
Limitations to know
Verbose and slow by default
Independent measurement shows below-median output speed, a long time to first token, and above-average output volume, and early hands-on reviews single out verbosity as the model's main irritation. Expect to tune prompts for length.
Launch benchmarks are mostly self-reported
Most headline results come from Anthropic's own evaluations, including internally named benchmarks. Independent replication was limited at launch, so treat the numbers as directional.
Safety classifiers can decline requests
The model ships with elevated cybersecurity safeguards, and flagged requests can be refused or routed to a fallback model. Benign security and life-sciences work occasionally trips them, so applications need to handle refusals explicitly.
Capabilities and specifications
Knowledge cutoff
May 1, 2026
Inputs
text, image, PDF
Reasoning
Supported
Default effort
high