Introduction
GLM-4.7 is a 358B-parameter open-weight model built around coding, tool use and long-running agent workflows. It supports a 200K context window and up to 128K output tokens, while Z.ai currently prices the API at $0.60 per 1M input tokens, $0.11 for cached input and $2.20 per 1M output tokens.
The benchmark results are still competitive for its generation: GLM-4.7 scored 73.8% on SWE-bench Verified, 66.7% on SWE-bench Multilingual, 41.0% on Terminal Bench 2.0 and 42.8% on Humanity’s Last Exam when tools were enabled. Local deployment is possible, but this is a very large model: popular Q4 GGUF builds are roughly 205–219 GB, so VRAM and system-memory requirements are far more demanding than a typical consumer LLM.
This GLM-4.7 review covers those benchmark results, current API pricing, local and GGUF deployment, VRAM requirements, agentic coding, and using GLM-4.7 for SillyTavern and roleplay.
Here’s the fastest way to orient yourself.
GLM-4.7 Snapshot: Pricing, Specs & Local Requirements
The key numbers for developers: current API pricing, context length, benchmark performance, agentic coding features and local deployment.
| Category | GLM-4.7 Details |
|---|---|
| Coding plan | Current: from $10/month Launch offer: $3/month The $3 price was the launch-era entry offer; current Z.ai Coding Plans start at $10 per month. |
| GLM-4.7 API pricing |
Pay-as-you-go
Input
$0.60
Cached input
$0.11
Output
$2.20
Prices are per 1 million tokens.
|
| Context & output | 200K context window with up to 128K output tokens. |
| Model focus | Agentic coding, long-running tool workflows, improved function calling, front-end generation and Preserved Thinking across agent turns. |
| Headline benchmark |
42.8% on
Humanity’s Last Exam with tools enabled.
The non-tool HLE result is 24.8%, so the 42.8% score specifically measures the tool-enabled setup.
|
| Local deployment | Open weights Approximately 358B parameters. Current Q4 GGUF builds are roughly 205–219 GB. Dual 24 GB GPUs cannot hold the full Q4 model entirely in VRAM; lower-memory systems require substantial RAM/CPU offloading. |
Table of Contents
1. The GLM-4.7 Hype: “Benchmaxxing” Or Real Breakthrough?
The headline matters because it changes who gets to test serious models. When a system shows up inside popular agent shells and costs less than a streaming subscription, it stops being a weekend experiment. It becomes a daily-driver candidate.
That also explains the skepticism. r/singularity users have seen enough “big jump” releases to ask the obvious question, did capability really move, or did the eval harness get friendlier?
My read is blunt. The interesting story is not one heroic number. It is a set of engineering choices aimed at making agentic workflows less fragile. The marketing talks about intelligence. The docs talk about stability and control.
1.1 My Expert Take
If you are hunting the best LLM for coding 2025, watch failure modes, not highlights. Great models keep their bearings across hours of back-and-forth, adapt when tools return messy output, and stay consistent when you change requirements mid-task. GLM-4.7 is clearly optimized for that style of work.
2. Spec Sheet: What Makes GLM-4.7 Different??

Specs are boring until they stop your workflow from breaking. A 200K context window is not a flex, it is permission to keep your design notes, logs, code, and constraints in one place without playing token Tetris. With GLM-4.7, that context ceiling is high enough to feel practical, not theoretical.
The headline specs are straightforward:
- 200K context length.
- Maximum output up to 128K tokens.
- Open weights release, with a permissive license in the model card.
That last bullet is the quiet one. It means you can take the model out of the hosted environment and into infrastructure you control. For privacy, for compliance, or for the simple joy of not being rate-limited mid-sprint, that matters.
2.1 The Capability Menu That Actually Matters
The feature list hits the modern essentials: thinking modes, streaming, function calling, context caching, and structured outputs like JSON. The differentiator is how much of that is tuned for agent loops. Z.ai GLM is selling “model plus agent ergonomics,” not a raw autocomplete engine.
3. The “Preserved Thinking” Feature Explained

This is the feature that made me stop scrolling. Preserved Thinking means the model retains its internal reasoning blocks across turns in coding agent scenarios, instead of re-deriving its plan from scratch each time.
In human terms, it reduces the goldfish problem. Many agents do something impressive, run a tool, then come back and narrate a slightly different universe. That inconsistency compounds. You end up debugging the agent instead of your code.
GLM-4.7 pairs Preserved Thinking with Interleaved Thinking, meaning it thinks before responses and tool calls. It also supports turn-level control so you can disable deep reasoning when you just want formatting, not philosophy.
3.1 Why This Changes Agent Stability
In agent work, the enemy is drift. Plans slowly mutate because the system forgot a constraint or lost a conclusion. Preserved Thinking is a direct countermeasure. It is a stability feature dressed up as a reasoning feature, and it matches the way real coding sessions unfold.
4. GLM-4.7 Vs Claude 4.5 Sonnet And GPT-5.2: The Coding Face-Off
If you are comparing top models for coding, you care about three outcomes:
- It fixes real bugs and writes real code.
- It survives long agent loops without babysitting.
- It generates UI you would not be embarrassed to ship.
GLM-4.7 is positioned directly against Claude Sonnet 4.5 in the agent-coder lane. The published benchmark table includes SWE-bench Verified, multilingual SWE-bench, Terminal Bench, and tool benchmarks like τ²-Bench and BrowseComp. The margins are close enough that behavior will matter more than rank.
4.1 Vibe Coding, The Unsexy Metric
“Vibe coding” sounds like a meme until you have to ship front-end. UI generation has a brutal threshold. Either the output is clean enough to keep, or you toss it and do it yourself.
The release notes claim a real jump in front-end aesthetics, cleaner pages, better slide layouts, and more accurate sizing. If that holds in your stack, it translates into fewer edits, fewer layout bugs, and faster iteration.
4.2 My Expert Take
Claude often feels like a careful senior engineer. GPT-5.2 tends to feel like a fast generalist. GLM-4.7 is trying to feel like a persistent agent teammate that keeps state. If that “statefulness” sticks, it is a competitive advantage you will notice on day two, not day one.
5. GLM-4.7 Benchmark Results: HLE, SWE-bench and Terminal Bench
Humanity’s Last Exam is the benchmark keyword everyone is repeating because it is hard and because it sounds like a movie trailer. The number attached to this launch is 42.8%.
GLM-4.7 Benchmark Results
Coding, software engineering, terminal-agent and reasoning performance across six published evaluations.
| Benchmark | GLM-4.7 Score |
|---|---|
SWE-bench Verified | 73.8% |
SWE-bench Multilingual | 66.7% |
LiveCodeBench v6 | 84.9% |
Terminal Bench 2.0 | 41.0% |
Humanity’s Last Exam | 24.8% |
Humanity’s Last Exam
with tools | 42.8% |
That score is for the tool-enabled setting. The same table shows a much lower score without tools. The delta is the story. Tool use gives a big lift compared to the previous generation’s tool setting.
That changes how you should interpret it. Tool-assisted evaluation is not a memory quiz. It is a workflow test. The model has to decide when to call tools, how to form calls, and how to integrate results without losing the thread.
5.1 Tool Use Is Not Cheating, It Is The Job
Real work looks like this: read context, form a hypothesis, run a command, adjust. Benchmarks that force tool use are closer to that loop. If GLM-4.7 is genuinely better at multi-step tool use, the payoff will show up in agent frameworks, not in a single one-shot answer.
6. GLM-4.7 API Pricing: Coding Plan vs Pay-As-You-Go Costs
At launch, one of GLM-4.7’s biggest attractions was Z.ai’s $3 entry-level coding-plan offer. That figure should now be treated as launch-era pricing rather than the current subscription price. Z.ai’s current documentation advertises Coding Plan access starting at $10 per month.
The GLM-4.7 API itself remains available separately. Z.ai currently charges $0.60 per 1M input tokens, $0.11 per 1M cached input tokens and $2.20 per 1M output tokens. Cached-input storage is currently listed as free for a limited time. GLM-4.7-Flash is currently listed as free by Z.ai, while GLM-4.7-FlashX costs $0.07 input, $0.01 cached input and $0.40 output per 1M tokens; these are separate models from the full GLM-4.7.
6.1 The Second Order Cost People Miss
Agent costs spike on output. A model that rambles is an expensive model. If you are doing GLM-4.7 vs Claude 4.5 Sonnet LLM API pricing math, track output length and tool chatter, not just input rates. Turn-level thinking control matters because you can be strict about when it “thinks hard” and when it just returns an answer.
GLM-4.7 Scenario Fit Table
| Scenario | Best Fit | Why |
|---|---|---|
| You want cheap daily coding help inside an agent | Z.ai coding plan | Low entry price, designed for agent shells |
| You are building product features | GLM API | Predictable billing, full control over prompts and tools |
| You need maximum privacy or fewer content filters | run GLM-4.7 locally | You control data flow, moderation, and logging |
7. GLM-4.7 SillyTavern & Roleplay: Sampling and Local Control
Coding is the headline, but creative communities care about different failure modes. Users want character consistency, lore tracking, and prose that does not collapse into repetitive “AI voice” filler.
The reported vibe is positive. People describe less “slop,” fewer stock phrases, better character continuity, and stronger handling of established universes. RWBY lore pops up as a common stress test because small continuity errors expose weak models quickly.
For SillyTavern, GLM-4.7 can be used through a compatible hosted API or a local backend serving the model or a quantized GGUF. Z.ai’s official default sampling settings are temperature 1.0 and top_p 0.95, which are sensible starting points before tuning for your roleplay style. SillyTavern itself is only the frontend, so whether a setup feels “uncensored” depends on the model or quantization, backend and moderation layer. Running the open weights locally gives you more control over that stack; it does not mean the official GLM-4.7 base model is specifically marketed as an uncensored model.
8. Run GLM-4.7 Locally: VRAM Requirements, GGUF and Hardware

GLM-4.7 can run locally, but its size changes the hardware equation. The official model is approximately 358B parameters, and current Q4 GGUF builds are roughly 205–219 GB before allowing additional memory for context, runtime overhead and the KV cache.
That means a pair of 24 GB RTX 3090s or 4090s cannot hold a full Q4 GLM-4.7 quant entirely in VRAM. Systems with less GPU memory can use system-RAM or CPU offloading, but performance will depend heavily on memory bandwidth, context length and how much of the model remains GPU-resident.
For GGUF users, quantized builds are available for llama.cpp-compatible applications, Ollama and LM Studio. For server-style deployment, the official GLM-4.7 model supports vLLM and SGLang, including OpenAI-compatible serving and tool calling.
8.1 A Practical Local Checklist
- GGUF / llama.cpp: best when you need quantization and CPU, RAM or mixed CPU/GPU offloading.
- vLLM: suited to GPU server deployment and OpenAI-compatible APIs.
- SGLang: particularly relevant if you want GLM-4.7’s preserved-thinking controls for agentic workloads.
- SillyTavern: use it as the frontend and connect it to whichever local or hosted backend is actually serving GLM-4.7.
Local is not the cheapest route. It is the most controllable route.
9. How To Migrate From GLM-4.6 To 4.7
Migrate is refreshingly boring. Update the model identifier, keep your prompts crisp, and decide how you want thinking and streaming handled.
Two details matter:
- Default sampling: temperature 1.0 and top_p 0.95, tune one at a time.
- Streaming tool calls: enable tool_stream=true to receive arguments as they form.
Minimal OpenAI-style SDK example:
from openai import OpenAI
client = OpenAI(
base_url="https://api.z.ai/api/paas/v4",
api_key="YOUR_ZAI_API_KEY",
)
resp = client.chat.completions.create(
model="glm-4.7",
messages=[{"role": "user", "content": "Briefly describe the advantages of GLM-4.7 for coding agents."}],
thinking={"type": "enabled"},
temperature=1.0,
)
print(resp.choices[0].message.content)Endpoints, depending on product surface:
Streaming tool-call pattern, trimmed to the essence:
response = client.chat.completions.create(
model="glm-4.7",
messages=[{"role": "user", "content": "Get weather for Beijing, then summarize."}],
tools=[...],
stream=True,
tool_stream=True,
)
final_args = {}
for chunk in response:
delta = chunk.choices[0].delta
if delta.tool_calls:
for tool_call in delta.tool_calls:
idx = tool_call.index
final_args[idx] = final_args.get(idx, "") + tool_call.function.argumentstool_call.index to rebuild the full JSON.
10. Safety And Privacy: The Open Source Advantage
Hosted platforms enforce policies. Local models enforce your policies. That split is the real “open” advantage.
If your work touches private repos, sensitive documents, or regulated data, running locally is less about edgy content and more about control. You decide what gets logged. You decide what leaves your machine. You decide how strict your environment should be.
Hosted still wins on convenience and updates. For many teams, the GLM API route is the cleanest compromise.
11. Pros And Cons Summary
Pros
- Cheap entry via the Z.ai coding plan, easy to plug into agent tools.
- Strong tool use focus, with competitive tool-assisted results.
- Preserved Thinking for multi-turn agent stability.
- Open weights option for local control and privacy.
Cons
- Local runs demand serious VRAM, often multi-GPU territory for good latency.
- Hosted filtering can clash with some creative scenarios.
- Long outputs can get expensive fast if you do not control verbosity.
12. Final Verdict: Who Is GLM-4.7 For?
For developers, the plan is the obvious starting point. Drop it into your agent shell and give it real work, a failing test, a messy log, and a UI request. Demos are easy. Long sessions are where models earn trust.
For product builders, the GLM API path offers predictable costs and clean integration, especially if you rely on function calling, streaming, and structured output.
For hobbyists and privacy maximalists, run GLM-4.7 locally. You get freedom and control, plus the delightful experience of debugging drivers when you would rather be writing code.
My closing take is simple. GLM-4.7 is not interesting because it claims to be smart. It is interesting because it is engineered to stay on task inside an agent. If you have been waiting for a model that feels less like a chatty assistant and more like a persistent coworker, test it on a real repo this week.
Make it earn its keep. If it saves you time and keeps you confident enough to ship, that is the only benchmark that matters.
Is the Z.ai Coding Plan really just $3?
Yes. The Z.ai GLM Coding Plan starts at $3/month and is a subscription designed for coding tools like Claude Code, Cline, OpenCode, and Roo Code. It uses a quota system (resetting every 5 hours) and does not convert into pay-as-you-go token billing. If you want direct API usage beyond those supported tools, you use the standard GLM API with per-token pricing (e.g., GLM-4.7 is billed per 1M tokens).
Is GLM-4.7 better than Claude 4.5 Sonnet?
It depends on what you mean by “better.” GLM-4.7 looks strongest when you run agentic, tool-using workflows and long multi-step tasks. Claude 4.5 Sonnet can still feel cleaner for some zero-shot coding prompts and “vibe” polish, especially for minimal-instruction tasks.
Can I run GLM-4.7 locally on my GPU?
Yes, but plan for real hardware. GLM-4.7 is big, so most people rely on 4-bit quantized builds. Practical setups include high-VRAM GPUs (often multi-GPU) or Apple Silicon configurations that can sustain large memory footprints.
What is “Humanity’s Last Exam” (HLE) in AI?
HLE is a reasoning-heavy benchmark designed to stress test real problem solving. The headline number people cite is GLM-4.7’s tool-assisted score, which signals the model can combine reasoning with external tool calls instead of only “answering from memory.”
Does Z.ai train on my code and private data?
For API usage, Z.ai’s API Data Processing Addendum says the company does not store the content you or your users provide or generate via API calls, it’s processed in real time. If you need maximum privacy and control, running open-weights locally keeps everything on your own machine.
