Fundamentals Published Aug 14, 2026 · 48 min read · 10,649 words
GLM-5.3 Review: Coding, Security, Benchmarks (2026)
An engineer's review of GLM-5.3 (Zhipu AI, 2026): post-training wins, Terminal-Bench 28.3, ExploitBench 54.4, the open-weights safety debate, cost, self-hosting.
Every six weeks or so, Zhipu AI ships another GLM. GLM-5 on February 12, GLM-5.1 in early April, GLM-5.2 on June 16, and now GLM-5.3 on August 14, 2026. Four flagship releases in roughly six months is not a cadence most labs can sustain, and it is the most important thing to understand about this release before you look at a single benchmark: Zhipu has decided that model quality is a shipping cadence problem, not a base-model problem. GLM-5.3 is the same 743B-parameter Mixture-of-Experts base as GLM-5.2, and every improvement in it comes from post-training.
That is the headline, and it is also the story that makes this review worth reading. GLM-5.3 is not a bigger model. It is a better-trained model: the terminal-use and software-engineering benchmarks moved more in one release than most labs move in a year, the security benchmark numbers look like a typo until you check them twice, and the open-weights release — delayed behind a safety review for the first time in the 5.x line — turned a routine model launch into a genuine controversy in the open-source community. As engineers who actually run these models in production, we have opinions about all of it.
This review is the honest version. We will walk through what GLM-5.3 actually is, how Zhipu builds models at this cadence, what the coding benchmarks really measure, what the cyber-defense jump means and what it does not, whether the safety-staging of the weights is good governance or a downgrade, how it compares to the field, what it costs to run (API and self-host), and exactly how you would use it tomorrow. Where the vendor claims are thin, we will say so. Where the numbers are genuinely impressive, we will say that too.
One note on method before we go further, because it determines how much of this you should trust. We are reviewing the model, not the company and not the marketing: we read the published benchmark methodology where it exists, we built the cost math from hardware economics rather than press releases, and we treat any vendor claim we cannot verify as a claim rather than a fact. Every number in this post that is not labeled as an estimate comes from Zhipu's public release materials for the 5.x line. Every number labeled as an estimate is our arithmetic, with the assumptions stated openly so you can redo it yourself. That distinction is the whole discipline of this review, and it is the difference between useful coverage and a press release with a byline.
Key takeaways
- GLM-5.3 is a post-training release on an unchanged base. Same 743B MoE as GLM-5.2; every gain comes from training, which is a deliberate and working strategy at Zhipu's six-week shipping cadence.
- The coding story is real. Terminal-Bench 3.0: 4.6 to 28.3 — a 6x jump and first among open-source models. DeepSWE v1.1: 46.2 to 66.9. This is the strongest single-revision jump in the category this year.
- The cyber-defense numbers are the controversy and the story. ExploitBench 24.4% to 54.4%, ExploitGym 2-hour tasks 29 to 105. A model that is this good at exploitation is why the weights are staged.
- Open weights are delayed, not canceled. Zhipu shipped GLM-5.2's weights day-one under MIT; 5.3's weights sit behind a safety review, expected in about two weeks. GLM-5.2 remains the fully open self-host fallback.
- Text-only, 1M context, 128K output. No vision — the top community request, polled by founder Jie Tang on June 29, did not make the cut.
- Cost is in the same band as 5.2. API pricing wasn't fully public at launch; the reference is $1.40 in / $4.40 out per 1M tokens, and self-host math lands around $0.14-2.78 per 1M output tokens on 8x H100 depending on utilization.
- Treat benchmark-vs-benchmark comparisons with care. Terminal-Bench 2.1 (5.2: ≈81) and Terminal-Bench 3.0 (5.3: 28.3) are different exams; the 6x claim is within-version and valid, cross-version leaderboard comparisons are not.
What GLM-5.3 actually is
GLM-5.3 is Zhipu AI's flagship text model, released August 14, 2026, under the tagline "Built to Code. Ready for Cyber Defense." It is a 743B-parameter Mixture-of-Experts architecture with roughly 40B active parameters per token — the same base used by GLM-5.2, which itself was a 753B/40B refinement of the 5.x line. The model keeps the 1M-token lossless context window that GLM-5.2 introduced and extends output to 128K tokens. It is text-only. There is no vision, no audio, no tool-specific modality: this is a language model aimed squarely at code, agents, and — new for this release — cyber-defense workloads.
The positioning matters. GLM-5.2 was the open-weights flagship: MIT license, day-one weights, top of the open-source Artificial Analysis Index at 51. GLM-5.3 is a different kind of release. The weights are staged behind a safety review that Zhipu expected to take about two weeks, the API launched with all GLM Coding Plan and ZCode quotas reset, and the security language is not marketing filler — the benchmark suite Zhipu published includes exploit-writing and cyber-gym evaluations that most labs would not have published at all. Whether that is confidence or recklessness is the question at the heart of this release, and we will come back to it.
To place it in the family history: GLM-4.5 (2025) was Zhipu's first natively agentic model at 128K context. GLM-4.6 (late 2025) was the flagship coding release at 200K. GLM-4.7 (2026) consolidated the line at 200K. Then the 5.x series changed the game: GLM-5 (February 12, 2026) introduced 200K context, a 744B-total / 40B-active architecture, DeepSeek Sparse Attention, async reinforcement learning, and the thesis Zhipu published as "GLM-5: From Vibe Coding to Agentic Engineering." GLM-5-Turbo followed in March for high-throughput agent workloads, GLM-5.1 in early April pushed the line to 754B/40B with support for autonomous runs up to eight hours, and GLM-5V-Turbo in April added multimodality. GLM-5.2 in June brought the 1M-token context, the "IndexShare" sparse-attention indexer that cut FLOPs 2.9x at 1M context, MIT weights, and the #1 open-weights Artificial Analysis Index spot.
GLM-5.3 sits on top of all of that, and the striking thing is how little of it is architecture. 743B total versus 5.2's 753B is noise — the active 40B is unchanged, the context window is inherited, the sparse attention is inherited. Everything that moved between June and August moved in the training pipeline. That is not a criticism. It is the most interesting engineering fact about Zhipu in 2026: they have turned model improvement into a fast-iteration post-training loop, and GLM-5.3 is the proof that the loop produces results.
The text-only constraint deserves its own paragraph, because it is a real decision with real consequences and not just a missing feature. The 1M-token context and the 128K output ceiling are both text budgets: you can hand the model an entire repository of source files, a multi-hour agent transcript, or a stack of documents, and it will hold them losslessly. But you cannot point it at a screenshot, a diagram, or a video frame, and any workflow that needs vision must either route those inputs to GLM-5V-Turbo or keep a second model in the loop. In our own agent stacks, that split is a genuine cost — every multimodal hop adds latency, a second vendor or model, and a new failure surface. Zhipu knows the demand is there; the June 29 poll put vision at the top of the community list. They chose to ship the coding story first. For a team whose work is code and text, that ordering is defensible. For a team whose work includes anything visual, it is the single most important reason to keep GLM-5V-Turbo or a multimodal competitor in the rotation.
The post-training story: why Zhipu ships every six weeks
Most labs release a model when they have a new base. Zhipu appears to have inverted that: the base is the stable platform, and the model you actually use is the product of a continuous post-training pipeline that they keep feeding. The cadence — four flagship releases in about six months, with turbo and multimodal variants in between — only makes sense if post-training is a pipeline you can run on a schedule, not a research project you hope converges.
The pipeline has three visible stages, and GLM-5.3 is what they look like when they work. First is supervised fine-tuning on code and agentic traces. The training mixture is the secret sauce, but the shape is public in the way the results look: the model is taught not just to produce correct code but to behave correctly inside long, tool-using loops — plan, act, observe, recover. Second is reinforcement learning at scale, the async-RL machinery Zhipu pioneered with GLM-5. The rewards are not next-token likelihoods; they are outcomes. Does the terminal command succeed? Does the patch pass the test suite? Does the agent finish the task within the budget? That shift — training against terminal and test outcomes rather than against human-preferred text — is the single biggest reason the terminal and software-engineering benchmarks moved the way they did. Third is the safety staging, which is new and which we will treat separately, because it is a policy change as much as a training step.
The result is that GLM-5.3 inherits the 1M-token context and the sparse-attention economics of 5.2 but behaves differently in almost every agentic evaluation. Zhipu's internal claim is that GLM-5.3 is about 50% better than GLM-5.2 on coding evals. Public numbers broadly agree, though as always the vendor's internal mix is not disclosed and you should discount the exact percentage.
The data side of that pipeline is where the real moat sits, and it is worth being explicit about what we do and do not know. Zhipu does not publish its post-training mixtures, and no one outside the company can reconstruct them. What the benchmark shape tells us is that the mixture is heavy on terminal transcripts, tool-calling trajectories, and code-review pairs — the model behaves as if it has seen a lot of successful and unsuccessful agent runs and learned to tell the difference. The other thing the release cadence tells us is that Zhipu has built evaluation into the loop: you cannot ship a flagship every six weeks unless you have a benchmark pipeline fast enough to grade candidate checkpoints on that schedule, which means the eval harness itself is first-class infrastructure rather than an afterthought bolted on before a launch. That shows in the benchmark hygiene: within-version numbers are reported cleanly, and the methodology notes for the 5.3 release are specific enough to reproduce.
The deeper lesson for anyone who builds with these models is operational. A six-week cadence means the model under your workload changes constantly, and that has real consequences: benchmark churn, prompt-compatibility drift, and a decision about whether you pin a version or ride the release train. GLM-5.2 was a stable MIT-weights target you could lock down. GLM-5.3, with staged weights and a reset quota, is a moving target for the first few weeks. If you are running agents in production, decide now whether you are a "ship every six weeks" team or a "pinned release" team, because the release cadence forces the question either way.
The coding story: Terminal-Bench and DeepSWE
The coding numbers are the reason this release matters. On Terminal-Bench 3.0, GLM-5.3 scores 28.3 against GLM-5.2's 4.6 — a 6x jump and the best score among open-source models. On DeepSWE v1.1, it scores 66.9 against 46.2. On Agents' Last Exam, 28.5 against 23.8, first among open-source models in the CLI category. Before you file these away as marketing, it is worth understanding what each benchmark actually measures, because the story they tell is more specific and more interesting than "coding got better."
Terminal-Bench is a terminal-use benchmark. The model is dropped into a shell with a task — debug a failing service, write a script, refactor a repo, operate infrastructure — and scored on whether it can drive a real terminal toward the correct outcome through tool calls, error recovery, and environment manipulation. A score of 4.6 on GLM-5.2 means it could almost not do this at all: a terminal is a hostile environment where one wrong command corrupts state and there is no autocomplete safety net. A score of 28.3 means it can actually work in one for a meaningful fraction of tasks. The jump is not "better at writing functions." It is "better at operating computers." That is a categorical difference, and it is the difference Zhipu's post-training was explicitly built to produce, because terminal competence is the substrate of agentic work.
DeepSWE is a software-engineering benchmark in the SWE-bench family: real GitHub issues, real repositories, and a requirement to produce a patch that passes the project's actual test suite. GLM-5.2 was already strong at this (46.2); GLM-5.3's 66.9 puts it in a different tier. To put that in context, GLM-5.2 scored 62.1 on SWE-bench Pro — a related but different exam — so 5.3's DeepSWE number is consistent with a model that has genuinely gotten better at end-to-end issue resolution, not just at pattern-matching common fixes. The important methodological point is that DeepSWE requires the model to navigate a large unfamiliar codebase, identify the root cause, and produce a patch that passes tests it has never seen. That is close to what a working engineer actually does.
Two honest caveats about the coding story. First, the vendor's "50% better than GLM-5.2 on internal evals" claim is unverifiable — internal evals are whatever they say they are, and you should weight the public benchmarks far more heavily than that sentence. Second, benchmark-vs-benchmark comparisons across versions are treacherous. GLM-5.2 scored about 81 on Terminal-Bench 2.1, and GLM-5.3 scores 28.3 on Terminal-Bench 3.0. Anyone reading those two numbers naively would conclude 5.3 got dramatically worse. It did not: 3.0 is a harder revision of the exam, and the 6x claim is only valid within the same version (4.6 to 28.3 on 3.0). This is exactly the kind of cross-version comparison you will see misquoted all over social media for the next month, and now you can win that argument.
What the coding story means in practice: if you run agents that touch terminals, run builds, or resolve real issues, GLM-5.3 is the first open-family model where the agentic coding loop is not the weak link. The API is the fastest way to test that claim against your own repos, and the DeepSWE jump is the strongest public signal that the improvement is not benchmark gaming.
Let us be concrete about what a Terminal-Bench task feels like, because the abstraction hides the difficulty. You are not given a clean problem statement and a test file. You are given an environment — a misconfigured server, a half-finished script, a repository with a failing test — and the model must discover what is wrong by operating the machine: reading logs, inspecting processes, querying state, editing files, running commands, and interpreting the results of each action. There is no oracle. A wrong guess can destroy state the model still needed. GLM-5.2's 4.6 on version 3.0 is the score of a model that occasionally succeeds on the easy tasks and fails hard on anything that requires multi-step diagnosis. GLM-5.3's 28.3 is the score of a model that can chain diagnosis, action, and verification. The difference is not vocabulary. It is the learned ability to recover from its own mistakes mid-task — the most important capability for terminal work and the hardest thing to teach with supervised data alone, which is why it shows up in a release that leaned on reinforcement learning.
The agentic story: from benchmarks to eight-hour runs
The agentic story is where the coding numbers stop being about code and start being about autonomous work. Zhipu's stated direction since "GLM-5: From Vibe Coding to Agentic Engineering" has been to make models that can run long, unsupervised, tool-using loops. GLM-5.1 supported autonomous runs up to eight hours. GLM-5.3 inherits the 1M-token context that makes those runs possible, because an eight-hour agent session is mostly context: tool outputs, file diffs, error logs, intermediate plans. A 200K window chokes on that. A 1M window, with the sparse-attention IndexShare machinery that cut FLOPs 2.9x at 1M context, is what lets the loop breathe.
Agents' Last Exam is the benchmark to watch here. GLM-5.3 scores 28.5 against GLM-5.2's 23.8, first among open-source models in the CLI category. ALE is a hard, human-crafted exam of agentic tasks — not a coding quiz, but a test of whether the model can plan, use tools, and persist toward a goal in open-ended scenarios. The +4.7 is smaller than the coding jumps, and that is honest information: agentic competence is harder to move with post-training than terminal competence, because it requires the model to hold a plan together over long horizons under uncertainty. The improvement is real, but it is incremental, and anyone claiming 5.3 is suddenly a reliable autonomous worker for hours at a time is over-reading a benchmark that itself shows only incremental movement.
The practical reality of long agent runs is still what it has always been: the model is the least fragile part of the system. The 8-hour run is bounded by your tooling, your state management, your retry logic, and your budget far more than by the model's context window. We have run multi-hour agent sessions on GLM-5.2 and seen the model do impressive mid-run recovery — and we have seen the same sessions fail for reasons that had nothing to do with the model: a tool returned malformed JSON, a subprocess hung, a rate limit fired at exactly the wrong moment. GLM-5.3's improved terminal competence directly attacks the most common failure mode, which is the agent painting itself into a corner it cannot recover from. That is the real agentic story of this release: not "agents are now autonomous," but "the model is now less likely to be the thing that kills the run."
If you build agent loops, the 1M context is not a feature you will always need — it is a safety margin you will occasionally need very badly. A long run that exceeds 200K tokens and dies is not a rare event; it is a Tuesday. GLM-5.3 gives you headroom and 128K output tokens to write back large artifacts. Treat that as the operational argument for upgrading, independent of the benchmark story.
There is a second, subtler agentic gain in this release that the benchmarks do not capture directly: planning behavior. In our testing of models in this family, the practical failure mode of a long run is not that the model stops producing sensible text; it is that the model commits to a plan too early and then cannot revisit it, or spends the run oscillating between two plans without committing to either. Post-trained models in the 5.3 style show markedly more stable plan commitment — they pick a course, execute it, and change course only when evidence warrants, rather than when the token stream wanders. That behavior is hard to score in a benchmark and easy to observe in a real run, and it compounds with the terminal gains: a model that both plans stably and recovers from errors is qualitatively closer to a junior engineer than any single benchmark number captures.
The cyber-defense side: reading the security benchmarks honestly
Here is the part of this release that makes it different from every other model launch this year. Zhipu shipped GLM-5.3 with a tagline about cyber defense and published a security benchmark suite: CyberGym 77.2% to 84.5%, ExploitBench 24.4% to 54.4%, and ExploitGym moving from 29/39 completed tasks at the 2-hour/6-hour marks to 105/130. Read those numbers plainly, because the marketing framing does not survive contact with them: ExploitBench and ExploitGym are exploit-writing benchmarks. The model got dramatically better at writing working exploits. That is not a defensive capability. It is an offensive capability with defensive applications.
Let us be precise about what each benchmark measures, because the difference between them is where the honest analysis lives. CyberGym is the defense-leaning one: it tests the model's performance in cyber-defense and security-operations scenarios — detecting, triaging, and responding to threats inside a simulated environment. An 84.5% score there is a legitimate defensive story: a model that can triage alerts, analyze indicators of compromise, and draft responses is a real SOC multiplier. The 77.2 to 84.5 movement is solid but not explosive. ExploitBench is different. It is a benchmark of exploit generation — taking a vulnerability and producing a working exploit. The 24.4% to 54.4% jump is more than double, and it is the number that should make anyone pause. ExploitGym is the same story in gym form: a model placed in a gym environment is scored on how many exploit tasks it completes in 2-hour and 6-hour budgets. Going from 29 to 105 (2-hour) and 39 to 130 (6-hour) is a roughly 3.5x jump in offensive capability.
The honest reading has three parts. First, the defensive positioning is not a lie — a model this good at CyberGym genuinely helps defenders, and Zhipu is right that the same skills (protocol knowledge, system internals, careful adversarial reasoning) serve both sides of the field. Second, the offensive capability is real, large, and the direct cause of the staged weights. You do not hold back weights for a model that is merely good at defense. The weights are staged because the model is good at exploitation, and any company that shipped 54.4%-on-ExploitBench weights day-one under MIT would be — correctly — criticized. Third, the benchmarks are gym environments, which means the real-world difficulty distribution is unknown. Exploiting a contrived gym vulnerability is not the same as chaining zero-days in production infrastructure. What the jump demonstrates is capability headroom, not immediate weaponization. Both of those statements are true, and anyone who tells you only one of them is selling you something.
What this means operationally for a working engineer is straightforward. If you are building defensive tooling — alert triage, log analysis, incident response drafting — GLM-5.3 is the most interesting open-family model this year, and the API path is available today. If you are deploying it self-hosted, treat the safety staging as the signal it is: run your own red-team before you put it in front of anything sensitive, because a model that writes exploits is a model whose outputs you cannot trust by default. The dual-use reality is not a reason to avoid the model. It is a reason to deploy it the way you would deploy any powerful tool: with review, monitoring, and the assumption that it can be misused.
The other honest point about the security story is the framing asymmetry. "Ready for Cyber Defense" reads as if the model were built for defenders, and the CyberGym score supports that reading. But the same release contains a tripling of exploit-generation capability, and the tagline does not mention that half. A more honest tagline would be "Built to Code. Capable of Cyber Operations." That is not a reason to refuse the model — dual-use capability is the normal state of powerful tools, and the open-source ecosystem has always shipped such tools with appropriate caveats. But it is a reason to name the asymmetry in your own risk assessment, especially if you are a security vendor: the model that drafts your incident-response playbooks is also the model a sophisticated adversary can point at the same playbooks from the other side. The mitigation is not less capability; it is the discipline you already apply to every tool in a security stack — least privilege, logging, and the assumption that whatever you deploy can be used against you.
The safety-staging controversy: good governance or a downgrade?
GLM-5.2 shipped its weights on day one under MIT. Anyone could download them, self-host, fine-tune, and sell the result. GLM-5.3 does not. The weights are staged behind a safety review — Zhipu says roughly two weeks — and the API launched first, with all GLM Coding Plan and ZCode quotas reset at launch. This is the first time in the 5.x line that the open-weights pattern broke, and the community reaction split exactly along the fault line you would expect.
One camp calls it responsible engineering. The model's exploit-writing capability is not hypothetical; the numbers are public. Shipping a 54.4%-on-ExploitBench model under an MIT license with zero review would invite exactly the "irresponsible open weights" criticism that has dogged every lab that released a capable dual-use model without guardrails. Staging the release, running a safety review, and publishing the timeline is the same pattern Anthropic and OpenAI use for their frontier releases — Zhipu is doing the open-weights equivalent of a staged rollout. Under this reading, the two-week delay is the cost of maturity, and GLM-5.2 staying MIT-licensed means nobody is left stranded.
The other camp calls it a downgrade dressed as safety. The skeptic's version goes like this: the safety review is a soft landing for what is really a commercial decision. Zhipu now has a flagship model whose API is live and whose weights are not — for the first time, the only legal way to run the newest model is through Zhipu's API or its coding-plan products, at Zhipu's prices. The timing is convenient: launch day resets quotas, the API is the only path, and every team that wants 5.3 today must become a paying API customer. Under this reading, the safety review is a policy that will quietly become precedent — every future GLM flagship will have "a safety review," and open-weights day-one will be a thing of the past.
Both readings are partly right, and the honest position is the uncomfortable middle. The safety rationale is credible — the benchmarks genuinely are the kind of capability that responsible labs gate — and Zhipu published the timeline and named the process, which is more transparency than the purely-commercial reading would predict. But the commercial incentive is also real and undisclosed. The way to tell which motive dominates is to watch what happens after the weights drop: if GLM-5.4 and 5.5 ship with the same staged pattern, the safety story is doing a lot of commercial work. If future releases return to day-one weights when the capability profile allows it, the safety story was honest. For now, the correct stance is: treat the review as legitimate, plan for the two-week wait, and keep GLM-5.2 pinned as the self-host fallback. The controversy is real, but it is a two-week controversy, and the model under it is the same either way.
There is one more angle that the loudest voices on both sides miss: the safety review itself is a black box. Zhipu has not published what it evaluates, what thresholds gate a release, or what would happen if the review failed. For a company asking the open-source community to accept a delay "for safety," the absence of a published review protocol is a genuine gap — the more legitimate the safety rationale is, the more the community deserves to see the evaluation. That is the specific critique that neither camp is making well, and it is the one that will age the best.
There is also a practical licensing question hiding under the controversy that no one is answering yet, because it depends on what license the staged weights eventually ship under. GLM-5.2's MIT license meant commercial use, fine-tuning, and redistribution with essentially no friction — that license is a large part of why it became the de facto open-family workhorse for self-hosted production. If GLM-5.3's weights clear review under a similarly permissive license, the two-week delay is a footnote in the history of the line. If they ship under a more restrictive license — source-available, or with use restrictions attached to the cyber capabilities — then the staging was never really about timing, and the open-weights community has a materially different relationship with Zhipu going forward. The review window is short, but the license decision is the durable one, and it is worth waiting the two weeks to see it.
Model comparisons: GLM-5.3 vs the field
Where does GLM-5.3 actually sit? Against GLM-5.2, the answer is unambiguous: same base, materially better post-training. On Terminal-Bench 3.0, 4.6 to 28.3. On DeepSWE v1.1, 46.2 to 66.9. On Agents' Last Exam, 23.8 to 28.5. On the cyber suite, roughly a 1.1x to 3.6x movement depending on the benchmark. If you already run GLM-5.2 in production, the upgrade decision is about whether the coding and agentic gains justify moving off a stable, MIT-licensed, fully-known model to a new one with staged weights — for many teams, that is a "wait two weeks for the weights" decision, not a "never" decision.
Against the rest of the open field, the positioning is sharp. GLM-5.2 was already the #1 open-weights model on the Artificial Analysis Index (51), with SWE-bench Pro 62.1, AIME 2026 99.2, and GPQA-Diamond 91.2. GLM-5.3 takes that base and pushes the agentic-coding dimension hard, which is exactly where the open-source competition — DeepSeek and Qwen most notably — has been weakest. DeepSeek's strength has been raw reasoning and training efficiency; Qwen's has been breadth and multimodal coverage. Neither has published a terminal-use and exploit-capability story like this. The honest summary: GLM-5.3 is not simply "better than everything open" — it is specifically better at the agentic-coding-and-terminal slice that DeepSeek and Qwen have under-invested in. If your workload is long-horizon tool use, this is the open-family model to beat. If your workload is math, codegen-in-a-file, or multimodality, the gap is smaller and the competition (including GLM-5.2) remains strong.
Against the closed frontier, the claim from Zhipu is that GLM-5.3 is approaching Claude Fable 5 on coding and agentic benchmarks. We cannot verify that with public cross-model numbers, and you should treat "approaching" as vendor-speak that means "not there yet on everything." What the public numbers support is narrower and more defensible: first among open-source models on Terminal-Bench 3.0 and on Agents' Last Exam (CLI category), and a DeepSWE score that would be competitive in the closed tier. The interesting structural point is that the closed frontier's advantage used to be the training itself — the data, the reward design, the RL machinery. Zhipu's post-training pipeline has demonstrated it can close that gap with the same base architecture it already controls, which is a different and more durable kind of competition than buying more GPUs for a bigger base.
The comparison that matters most for working engineers is not GLM-5.3 versus the leaderboard. It is GLM-5.3 versus your specific workload, and the only way to run that comparison is with your own terminal tasks and your own repos, on the API, which is available today. Benchmark tables tell you where the model was trained to win. Your workload tells you whether it wins where you actually work. If the two do not agree, trust your workload.
One more comparison deserves a mention because it is the one that usually gets skipped: GLM-5.3 versus the older GLM line that many teams are still running. GLM-4.6 and GLM-4.7 are capable 200K-context models that were flagships a year ago, and for simple, well-scoped workloads — a summarization pipeline, a classification job, a single-file code assistant — the upgrade to 5.3 buys you little and costs you a migration. The jump to the 5.x line is worth it when your workload is agentic: tool use, multi-turn loops, long context, terminal operation. The 4.6/4.7 generation was trained in a different era of the reward function, and no amount of prompting closes the gap that outcome-based reinforcement learning opened up. Our rule of thumb is embarrassingly simple: if your workload is one call and done, stay where you are; if your workload is a loop, move. The cadence of this release line means there will be a 5.4 to reconsider soon anyway.
Cost and access: API vs self-host
Let us do the money math, because it is the part everyone gets wrong. At launch, GLM-5.3's API pricing was not fully public — it shipped in the same staged roll-out as the safety review — so the reference point is GLM-5.2's published pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens. The honest expectation is that 5.3 lands in the same band; do not buy anyone's invented 5.3 price until Zhipu publishes it. The access picture at launch is: the API is live, the GLM Coding Plan and ZCode quotas all reset, and the open weights are expected in about two weeks. For the first two weeks of this model's life, the API is the only legal way to run it.
The self-host math is where the interesting economics live, and it starts with the architecture. GLM-5.3 is 743B total with about 40B active parameters per token. A 40B-active MoE is very servable: it is roughly the active-parameter class of a 40B dense model, but with the expert weights of a much larger model. On an 8x H100-class box (the reasonable minimum), the numbers work out like this. Eight H100s peak around 8 petaFLOPs of BF16 compute; at a realistic 40% MFU with a big batch, you get roughly 40,000 output tokens per second for a 40B-active model. That puts one million output tokens at about 25 seconds of 8x H100 time. At a typical rental of about $2.50 per hour per H100 — $20 per hour for eight — that is roughly $0.14 per 1M output tokens at full batch utilization. Compare that to the API's $4.40 per 1M output tokens, and self-hosting at full utilization looks about 30x cheaper on raw compute.
The catch, and it is the entire catch, is utilization. That 40,000 tokens-per-second number assumes a saturated batch. A dev team with an interactive workload, small batches, and latency-sensitive requests runs at a fraction of that. At 25% utilization, the per-token cost rises to roughly $0.56 per 1M. At 5% utilization — which is where many "we host our own model" setups actually live — it is about $2.78 per 1M, already above half the API price. At 2% utilization, self-hosting costs more per token than the API. Add in the hardware itself: an 8x H100 box is roughly $20,000 per month to rent or a seven-figure capital cost to buy, plus engineering time to serve, monitor, and update it. The crossover is not at "do I have GPUs." The crossover is at "do I have enough sustained throughput to keep those GPUs busy."
So the honest decision framework is short. Use the API if you are evaluating, if your usage is spiky, if your latency budget is strict, or if the weights are not out yet — which, for the first two weeks, is everyone. Self-host once the weights drop if you have sustained high-volume throughput, hard data-residency requirements, or a price per token that must approach the raw-compute floor. And remember the middle path that most teams miss: GLM-5.2 is MIT-licensed, self-hostable today, and close enough to 5.3 for many workloads that the two-week wait for 5.3 weights costs nothing but patience.
The cost asymmetry between input and output tokens is worth a separate line, because it shapes how you pay for agent workloads in particular. GLM-5.2's pricing — and by extension the 5.3 band — charges roughly three times more per output token than per input token, which is the industry norm and which punishes exactly the kind of work this model is good at: a long agent run reads a lot of context (cheap) and writes a lot of tokens back into the world (expensive). The 128K output ceiling compounds this: if you use it, you pay for it. Two mitigations apply to any agent budget. First, cache aggressively — if your provider offers prompt caching, a long shared context is the ideal cache key, and it turns the input side from a cost into nearly a free resource. Second, be honest about output discipline: many agent loops write verbose intermediate logs that the model then re-reads, and trimming what the model writes back into its own context is the single highest-leverage cost optimization in a token-priced world.
Hands-on: running GLM-5.3
Enough theory. Here is how you actually use this model tomorrow, on both paths — the API that is live today, and the self-host path you will take once the weights clear review.
The API path (available now). GLM-5.3's API is OpenAI-compatible, which is the single most pragmatic thing about the release. If your code already talks to an OpenAI-compatible endpoint, you point base_url at Zhipu's endpoint, set the model to glm-5.3, and you are done. A minimal call in Python looks like this:
from openai import OpenAI
client = OpenAI(
base_url="https://api.z.ai/api/paas/v4",
api_key="ZAI_API_KEY", # from your Zhipu account
)
resp = client.chat.completions.create(
model="glm-5.3",
messages=[
{
"role": "system",
"content": (
"You are a senior engineer driving a Linux terminal. "
"Plan first, then execute one command at a time, "
"and recover from errors by reading the output."
),
},
{
"role": "user",
"content": (
"A service on this box died overnight. Find why it "
"stopped, fix the root cause, and restart it so it "
"survives reboots."
),
},
],
max_tokens=8192,
)
print(resp.choices[0].message.content)
Two practical notes for the terminal-style workloads this model targets. First, the system prompt matters more than it does on less capable models: explicitly asking for plan-then-execute behavior measurably improves terminal-task success, which is consistent with how the model was trained. Second, for real agent loops you want the response to include structured tool calls, so use the tool-calling interface rather than asking for commands in prose — the model was trained on exactly that pattern.
On access, the GLM Coding Plan and ZCode products both reset user quotas at launch, which matters more than it sounds: it means your existing subscription tier immediately gets a fresh allocation of GLM-5.3 usage, so the cheapest way to evaluate the model is often the plan you already pay for rather than a new API spend. The reset is a deliberate launch-day move — Zhipu wants the new model under as many fingertips as possible in the first two weeks, both to collect usage data and to convert curiosity into habit before the weights are public. Take advantage of it: run your actual terminal tasks and your real repositories through the API before you spend a dollar on infrastructure, because the evaluation result is the one number that should drive your deployment decision.
The self-host path (in about two weeks). Once the weights drop, the fastest route is a serving stack that already handles MoE models well. The pattern looks like this:
pip install vllm
vllm serve zai-ai/glm-5.3 \
--tensor-parallel-size 8 \
--max-model-len 1048576 \
--trust-remote-code
That command assumes eight GPUs with enough VRAM for a 743B MoE at the target precision — the exact repo path and quantization options will be published with the weights, so treat the model name as illustrative until then. If you do not have 8x H100-class hardware, your options are quantization (4-bit drops the footprint dramatically while MoE keeps the active-parameter quality largely intact) or renting, and the honest math from the cost section applies: rent until your utilization justifies owning.
The GLM-5.2 bridge. There is a third path that costs nothing: GLM-5.2 is MIT-licensed and self-hostable right now, with the same 40B-active footprint, the same 1M context, and close-enough behavior for most workloads. For a team that wants the open path today, pinning GLM-5.2 and swapping to 5.3 when the weights clear is the lowest-risk upgrade in the entire release. The benchmarks say 5.3 is meaningfully better at terminal use; they do not say 5.2 is bad. The two-week staging window is precisely the length of time it should take you to decide which workloads justify the jump.
Honest caveats
Every review of a model this hyped needs a cold-water section, and this one is no different. Six caveats, in no particular order.
First, the vendor's headline claims are unverifiable. "About 50% better than GLM-5.2 on internal evals" and "approaching Claude Fable 5" are not benchmark results; they are positioning. The public numbers are the only numbers you should weight, and even they are single-run vendor-published figures until third parties reproduce them.
Second, benchmark revisions break naive comparisons. Terminal-Bench 2.1 and 3.0 are different exams; the 6x jump is within-version only. Expect a wave of misleading cross-version charts in the next month and treat every one with suspicion, including — especially — the ones that flatter the model.
Third, the cyber-defense numbers are gym numbers. ExploitGym and ExploitBench measure capability in sandboxed environments, not weaponization in the wild. The jump is real capability headroom; it is not evidence of real-world exploitation competence, and it is not a reason to panic — or to be complacent.
Fourth, the safety review is an opaque process. Zhipu has not published what the review evaluates or what would fail it. A legitimate safety rationale deserves a published protocol, and its absence is a fair criticism regardless of which camp you are in.
Fifth, long-agent-run reliability is still the hard part. The benchmarks moved, but an eight-hour autonomous run is bounded by your tooling, your state management, and your error handling more than by the model. GLM-5.3 raises the ceiling; it does not remove the floor.
Sixth, the pricing is not fully public. Anyone quoting you a specific GLM-5.3 API price before Zhipu publishes one is guessing. The reference band — $1.40 in / $4.40 out, in the style of 5.2 — is the honest expectation, and the self-host numbers in this review are estimates built on stated assumptions, not vendor figures.
None of these caveats changes the verdict. They change how confidently you can act on it, and that distinction is the entire job of a review.
FAQ
Is GLM-5.3 open source? Not on day one. Zhipu staged the open weights behind a safety review — a deliberate departure from GLM-5.2's day-one MIT release — and expects them in about two weeks. GLM-5.2 stays fully MIT-licensed and is the self-host fallback. The API, GLM Coding Plan, and ZCode access were live at launch with all quotas reset.
How much does GLM-5.3 cost? GLM-5.3 API pricing was not fully public at launch because it shipped in the staged roll-out. The reference point is GLM-5.2's published pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens, with 5.3 expected in the same band. Self-host math runs roughly $0.14 to $2.78 per 1M output tokens on an 8x H100-class box depending on utilization, before engineering and hardware costs.
Is GLM-5.3 safe? It is genuinely dual-use, and the published exploit benchmarks are why. CyberGym 84.5% is a real defensive capability; ExploitBench 54.4% and ExploitGym 105/130 are real offensive capability. The staged weights are the acknowledgment of that. If you self-host, run your own red-team and misuse evaluation.
Does GLM-5.3 have vision? No. It is text-only. Vision was the top community request — confirmed in founder Jie Tang's June 29 poll — but was not included. Zhipu's multimodal option remains GLM-5V-Turbo.
Can I self-host GLM-5.3? Once the weights clear review, expected in about two weeks. It is a 743B MoE with 40B active, so plan for 8x H100-class hardware or a quantized deployment. Until then, GLM-5.2 (MIT, 753B/40B, 1M context) is the self-host path.
How good is GLM-5.3 at coding? The strongest part of the release. Terminal-Bench 3.0: 4.6 to 28.3, a 6x jump and first among open-source models. DeepSWE v1.1: 46.2 to 66.9. Zhipu claims about 50% better than 5.2 internally and approaching Claude Fable 5; discount the claims, trust the public benchmarks.
What is GLM-5.3's context window? 1M-token lossless context, inherited from GLM-5.2, with up to 128K output tokens. Text-only. That is enough for a mid-size codebase or a long agent run in a single pass.
What is new in GLM-5.3 versus GLM-5.2? The same 743B MoE base — all gains come from post-training. Terminal, software-engineering, agentic, and cyber-defense benchmarks moved sharply; open weights moved to a staged model; and the cyber-defense positioning is new. API and coding-plan quotas all reset at launch.
Further reading
- Web scraping APIs in 2026 — the buy-versus-build decision, applied to hosted fetching and rendering.
- Web scraping tools — the tooling landscape, from HTTP clients to browser automation.
- Web scraping with Python — the language most model-assisted scraping loops are written in.
- Web scraping with Node.js — the async-first alternative for high-concurrency collectors.
- Scraping APIs for AI agents — how agent loops like the ones GLM-5.3 powers consume web content at scale.
Frequently Asked Questions
Is GLM-5.3 open source?
Not on day one. Zhipu staged the open weights behind a safety review — a deliberate departure from GLM-5.2's day-one MIT release. Weights are expected within about two weeks of the August 14, 2026 launch. GLM-5.2 stays fully MIT-licensed and is the self-host fallback. The API and GLM Coding Plan / ZCode access were live at launch, with all user quotas reset.
How much does GLM-5.3 cost?
GLM-5.3 API pricing was not fully public at launch because it shipped in the same staged roll-out as the safety review. The reference point is GLM-5.2's published pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens. Expect 5.3 to land in the same band. Self-hosting costs roughly $0.14 to $2.78 per 1M output tokens on an 8x H100-class box depending on utilization, before engineering and maintenance.
Is GLM-5.3 safe?
It is a genuinely dual-use model and Zhipu treats it that way. The cyber-defense numbers are public and real — ExploitBench 54.4%, CyberGym 84.5%, ExploitGym 105/130 tasks — which is exactly why the open weights are staged behind a safety review instead of shipping day-one. If you self-host, run your own red-team and misuse evaluation. The capability is the point of the release, and the risk review is the honest acknowledgment of that.
Does GLM-5.3 have vision?
No. GLM-5.3 is text-only. Vision was the top community request — founder Jie Tang ran a June 29 poll and it won — but it was not included in this release. Zhipu's multimodal option remains GLM-5V-Turbo, from April 2026. If you need image understanding, 5.3 is not the model.
Can I self-host GLM-5.3?
Once the weights clear the safety review, yes — expected within about two weeks of launch. It is a 743B-parameter Mixture-of-Experts with 40B active parameters, so plan for 8x H100-class (or equivalent) hardware for respectable throughput, or a quantized deployment. Until the weights drop, GLM-5.2 (MIT license, 753B/40B, 1M context) is the self-host fallback and a very close approximation.
How good is GLM-5.3 at coding?
It is the strongest part of the release. On Terminal-Bench 3.0 GLM-5.3 scores 28.3 versus GLM-5.2's 4.6 — a 6x jump and the top score among open-source models. DeepSWE v1.1 goes 46.2 to 66.9. Zhipu claims 5.3 is about 50% better than 5.2 on internal evals and approaching Claude Fable 5. Treat the vendor claims with skepticism, but the public benchmark movement is real and directionally consistent.
What is GLM-5.3's context window?
GLM-5.3 keeps GLM-5.2's 1M-token lossless context window and adds up to 128K output tokens. That is enough to hold an entire mid-size codebase, a long agent run, or several books in a single pass. The model is text-only, so that context is for text.
What is new in GLM-5.3 compared to GLM-5.2?
The base model is the same 743B MoE, so every benchmark gain comes from post-training, not a bigger architecture. The visible changes: large jumps in terminal and software-engineering benchmarks, a major cyber-defense capability jump (ExploitBench 24.4% to 54.4%), and the safety-staging of open weights — a change from 5.2's day-one MIT release. API, GLM Coding Plan, and ZCode quotas all reset at launch.
Keep reading
The Best Web Scraping Tools in 2026: Full Breakdown
The honest 2026 ranking of web scraping tools — libraries, GUI scrapers, APIs and platforms — with real pricing, honest pros and cons, and a decision guide.
What Actually Gets You Blocked When Web Scraping (2026)
We run a web scraping API and see millions of requests a day. Here's what actually gets you blocked in 2026 — and what doesn't, signal by signal.
Web Scraping for AI Training Data in 2026
How AI companies build training datasets: Common Crawl, web corpora, domain scraping, the crawl-to-JSONL pipeline, what makes good data, the 2026 legal landscape, and how a small team builds its own.
Found this useful? Cite it as: webscraping.space. “GLM-5.3 Review: Coding, Security, Benchmarks (2026).” https://webscraping.space/blog/glm-5-3-model-review. Published 2026-08-14.