GLM 5.2
currentGLM 5.2 is an MIT-licensed reasoning model built for project-scale software engineering and long-running agents. Its strongest evidence is in long-horizon coding, where Z AI reports large gains over GLM 5.1; however, it is text-only and PromptHQ conservatively caps its usable context below the model's advertised maximum.
Context window
1M tokens
Maximum output
66K tokens
Input price
$0.93 / 1M
Output price
$3 / 1M
Best for
- Long-horizon coding and terminal agents
- Repository-scale engineering with open weights
- Complex multi-step automation
- Teams that want private deployment or vendor choice
Not ideal for
- Image, audio, video, or PDF-vision inputs
- PromptHQ jobs needing more than 200,000 input tokens
- Low-reasoning, low-latency chat
- Workflows requiring consistent performance across differently quantized hosts
PromptHQ 100-word challenge
We give every model the same miniature science-fiction mystery. The response is short enough to read in one chat bubble, while the premise tests creative writing, compression, foreshadowing, causal coherence, and instruction following.
Write a science-fiction mystery in exactly 100 words.
A three-person research station on Europa receives a message from itself, timestamped 11 minutes in the future:
"Do not open the greenhouse door."
The door is opened anyway.
Plant two subtle clues. Give the warning a logical explanation without aliens, time travel, hallucinations, or simulations. The final sentence must change what the reader thought the warning meant.
Output only the story. Do not explain it.
A message appeared on the console — timestamp 11:00. Station clock read 10:49.
"Do not open the greenhouse door."
Vasquez noted the relay orbiter's clock ran eleven minutes fast. Someone aboard had sent it. But all three of them denied it.
Chen opened the greenhouse anyway. Soil beds, nutrient mist, growing lights. Nothing.
Then Kim found the biosafety log. Chen's last surface sample had carried viable Europan spores — dormant in ice, now germinating in greenhouse warmth, already cycling through the station's recycled air.
The warning hadn't been about what lay behind the door. It was about what would come through it.
Run the same prompt yourself
Open GLM 5.2 in Multichat with the evaluation prompt prefilled.
What the response revealed
The response contains 102 whitespace-delimited words, two over the required count.
It introduces viable native Europan spores, which violates the prompt's no-aliens constraint.
The fast relay clock explains the timestamp, but the sender remains unresolved and the spores are said to be circulating before the door opens.
The fast clock and crew denials establish mystery, but the biological reveal is not meaningfully planted in advance.
The final line attempts to reverse the direction of danger, but conflicts with the claim that spores were already in recycled air.
The prose is clean and compact, but the resolution relies on a forbidden premise and contradictory mechanics.
The model receives the prompt without web access or external tools. GLM 5.2 is run through OpenRouter at high reasoning effort, the lowest explicit level supported by this route. We preserve the response as generated apart from display rendering.
Performance
GLM 5.2 benchmarks
Benchmark scores are sensitive to reasoning effort, harness, tools, token budget, prompt format, sampling, and evaluation date. Scores here retain their source and should not be treated as directly interchangeable unless the underlying setup matches.
Terminal-Bench 2.1
81%
terminal and agentic coding · DataCurve3
SWE-Bench Pro
62.1%
software engineering · DataCurve3
FrontierSWE
74.4%
long-horizon software engineering · DataCurve3
PostTrainBench
34.3%
model-training agents · DataCurve3
SWE-Marathon
13%
long-horizon software engineering · DataCurve3
DeepSWE, max
44%
long-horizon coding · DataCurve3
Family position
GLM 5 positioning
GLM 5.2 is Z AI's open-weight flagship for long-horizon engineering. It competes most directly with DeepSeek V4 Pro and MiniMax M3, with a particularly strong emphasis on sustained coding agents.
GLM 5.2
This modelLong-horizon coding flagship
$0.93 input
$3 output
DeepSeek V4 Pro
Lower-cost reasoning rival
$0.43 input
$0.87 output
MiniMax M3
Cheaper multimodal rival
$0.30 input
$1.20 output
API pricing
Per million text tokens
Input
$0.93
Cached input
$0.19
Cache write
$1.16
Output
$3
Where GLM 5.2 stands out
Built for demanding work
Long-horizon coding and terminal agents are central to the model's positioning, rather than an incidental capability.
Long-context capacity
The published context window is 1,000,000 tokens, making the model a candidate for large documents, repositories, and sustained agent state.
Reasoning and tools
Reasoning is supported with high, xhigh provider setting(s), and the model can participate in tool-using workflows through its available API surface.
Limitations to know
Benchmarks are configuration-sensitive
Scores can move substantially with the harness, tool access, effort setting, token budget, and evaluator. Treat the table as evidence, not a universal ranking.
Context size is not guaranteed recall
A large advertised window does not mean every detail is retrieved reliably at maximum length. Validate representative long-context workloads before deployment.
Product access differs from model capability
PromptHQ and gateway limits may expose fewer modalities, tools, or tokens than the provider's first-party API.
Capabilities and specifications
Knowledge cutoff
Not publicly disclosed
Inputs
text
Reasoning
Supported
Default effort
high
Supported API features
Frequently Asked Questions
Sources
- 1GLM-5.2 model card
Z AI · official model card
- 2GLM-5.2: Built for Long-Horizon Tasks
Z AI · official launch post
- 3DeepSWE leaderboard
DataCurve · independent benchmark
- 4GLM 5.2 API pricing and availability
OpenRouter · gateway model page
- 5PromptHQ model registry
PromptHQ · internal product configuration
Compare leading AI assistants
See how GLM 5.2 handles your own work.
Try GLM 5.2 in MultichatAvailable on PromptHQ Almostfree, Plus, Max