Model or dataset
zai-org/GLM-5 avatar
zai-org/GLM-5

GLM-5: Z.ai's open-weights model family for long-horizon agentic coding

GLM-5: From Vibe Coding to Agentic Engineering

7,196 stars948 forksUnknownApache-2.0

At a glance

What is it?
The zai-org/GLM-5 repository collects four generations of Z.ai's GLM models, from the 744B-parameter GLM-5 base to the post-trained GLM-5.3. It is a weights-and-API release, not a runnable application, and the README is explicit about which gains come from training and which from architecture.
Who is it for?
Adopt GLM-5 if you are building an agent that runs for hundreds of tool calls and you want open weights behind it, or if you need a coding model reachable through an OpenAI-style endpoint. Do not adopt it if you expect a pip install that starts a chat UI: this repository publishes model cards, weights and a requirements.txt, and the README does not document local serving, quantisation or rollback between model versions.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What GLM-5 is, and the four models sharing one repository

The repository is a model release, not a framework. It documents four generations under one README: GLM-5, GLM-5.1, GLM-5.2 and GLM-5.3 (plus GLM-5.3-Flash). Each is aimed at the same buyer, engineers building agents that must keep working after the first plausible patch fails.

The generational story is worth reading carefully because the README is unusually specific about where the gains come from. GLM-5 scales from GLM-4.5's 355B parameters (32B active) to 744B parameters (40B active) and grows pre-training data from 23T to 28.5T tokens. GLM-5.1 is described as the agentic-engineering generation, with the README claiming it stays productive over longer sessions and sustains optimisation over hundreds of rounds and thousands of tool calls. GLM-5.2 adds a stated 1M-token context and IndexShare, which reuses one indexer across every four sparse attention layers and is said to cut per-token FLOPs by 2.9x at 1M context. GLM-5.3, by contrast, uses the same base model as GLM-5.2; the README says every gain comes from post-training.

That last sentence is the most useful line in the document for anyone choosing a model. It tells you GLM-5.3 is not a new architecture you need to re-benchmark for memory or throughput. It is the same weights plus a different post-training recipe, so the delta shows up in task success, not in serving cost.

The mechanism: sparse attention, IndexShare and asynchronous RL

Two architectural threads run through the README. The first is attention efficiency. GLM-5 integrates DeepSeek Sparse Attention to reduce deployment cost while preserving long-context capacity. GLM-5.2 layers IndexShare on top, sharing one indexer across four sparse attention layers. GLM-5.3-Flash goes further and introduces a hybrid of sparse and linear attention, which the README says sharply reduces long-context serving costs, plus Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency. If you serve these models yourself, this is the part that determines your KV-cache budget and your per-token cost at long context.

The second thread is training infrastructure. The README credits slime, an asynchronous RL infrastructure, with improving training throughput and enabling finer-grained post-training iterations. That matters because the GLM-5.3 claim (same base model, all gains from post-training) only makes sense if post-training is cheap enough to iterate on repeatedly. The repository itself does not ship slime; it is a separate project at github.com/THUDM/slime.

What the repository does ship is thin: a requirements.txt pinning transformers, pre-commit and accelerate, an example directory containing a single ascend.md file, and a skills directory. There is no inference server, no Dockerfile, no quantisation script. The README points to the Z.ai API Platform for hosted access and to z.ai for the product. Treat this as a weights-and-documentation repository and plan your serving layer separately.

Installing GLM-5: what the repository actually gives you

There is no install command for the model itself in the README. The repository provides a requirements.txt, and the download table in the README lists model sizes and precisions alongside download links. The requirements file pins the Python-side tooling:

bash
pip install -r requirements.txt

That installs transformers>=5.15.0, accelerate>=1.14.0 and pre-commit>=4.6.2. Note the transformers floor: 5.15.0 is well above the 4.x line most teams still have pinned, so expect to resolve that in a separate environment rather than in an existing training venv.

The README's own recommended path for most users is the hosted API rather than local weights. It links to the Z.ai API Platform at docs.z.ai/guides/llm/glm-5.3 for GLM-5.3 and GLM-5.3-Flash. The README does not publish the endpoint URL, the model identifier string or an authentication example, so read those from the platform documentation rather than guessing. What the README does give you is the shape of the decision: weights if you need them, API if you want to start today.

One concrete detail for anyone running on non-NVIDIA hardware: the example directory contains a single file, example/ascend.md, which suggests Ascend deployment is documented there. It is the only example in the repository.

Where GLM-5 is the wrong tool

The README is a launch document, and it reads like one. Several things you would need before a production decision are simply absent.

The download table is truncated in the repository text, so model sizes and precisions for individual checkpoints are not fully readable from the README alone. There are no quantised variants described, no GGUF or similar format mentioned, and no local serving instructions. If your plan is to run GLM-5 on a workstation, the README does not tell you how, and the requirements.txt does not help.

The benchmark numbers are almost entirely in-house or vendor-reported. Z.ai Code Bench is described as in-house, and the 50% improvement of GLM-5.3 over GLM-5.2 is measured there. The public benchmarks cited (Terminal Bench 3.0, Agents' Last Exam, SWE-bench Pro, Terminal-Bench 2.1, CyberGym) are real names, but the comparisons to Claude Opus 4.8 and Gemini 3.1 Pro are the vendor's own framing. Independent reproduction is the thing to look for, and the README does not cite any.

Finally, the cyber capability section is a genuine consideration, not a marketing line. The README states that cyber capability developed faster than expected as post-training scaled, that GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and that its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks. If your threat model includes a capable open-weights model being used offensively, that is a fact about this release, and it is stated by the authors. It is also a reason some organisations will route around open weights entirely.

GLM-5 versus a general-purpose chat model

The obvious alternative is a frontier closed model accessed through the same kind of API, and the README makes the comparison itself: on Terminal-Bench 2.1, GLM-5.2 scores 81.0 against Claude Opus 4.8 at 85.0, while staying ahead of Gemini 3.1 Pro. That is the vendor's number, but it frames the real trade-off well. You give up a few points of peak capability and you get weights you can host, fine-tune and inspect.

The difference in approach is not just price. A closed chat model is optimised for a single strong answer to a well-posed question. GLM-5.1's stated design goal is the opposite: the README says previous models exhaust their repertoire early and plateau, and that GLM-5.1 instead stays effective over hundreds of rounds and thousands of tool calls. The Vending Bench 2 example makes this concrete: the model runs a simulated vending machine business over a one-year horizon and finishes with a final account balance of $4,432, which the README says approaches Claude Opus 4.5. That is a different evaluation target from a chat benchmark, and it is the one that matters if you are building an agent that has to keep making decisions after the first hundred fail.

If your workload is a single question and a single answer, the long-horizon machinery is overhead you are paying for and not using. A smaller model will be cheaper and just as good.

Licence, maintenance and the cost of upgrading between generations

The repository is licensed Apache-2.0. That is a permissive licence, and it is the reason many teams will look at this release at all, but the licence file is the authority on what it covers and this article is not legal advice. Check whether the licence text in the repository extends to the model weights or only to the code and documentation, because the two are not automatically the same thing.

The last push to the repository was on 2026-09-01, which is recent enough that the project is not dormant. The repository is not archived. There are no retrieved releases in the repository, so there is no versioned changelog to read and no tagged upgrade path. In practice that means tracking changes means watching the README, which currently documents four model generations in one file. When GLM-5.4 arrives, the GLM-5.3 section will likely move down the page rather than into a release note.

The upgrade cost is unusual in one specific way. Because GLM-5.3 uses the same base model as GLM-5.2, moving from 5.2 to 5.3 should not require re-validating your memory footprint or throughput assumptions. Moving from GLM-5 or GLM-5.1 to either of those does, since the architecture changed. The README does not document rollback between model versions, so if you pin a model string in production, keep the previous one reachable.

Editorial conclusion

Adopt GLM-5 if you are building an agent that runs for hundreds of tool calls and you want open weights behind it, or if you need a coding model reachable through an OpenAI-style endpoint. Do not adopt it if you expect a pip install that starts a chat UI: this repository publishes model cards, weights and a requirements.txt, and the README does not document local serving, quantisation or rollback between model versions. Before committing, verify three things: which of GLM-5, 5.1, 5.2 or 5.3 your target benchmark actually refers to, whether the 1M-token context in GLM-5.2 is available at the effort level you plan to run, and what the Z.ai API Platform charges for the specific model string you intend to call.

Frequently asked questions

What is the GLM-5 model?

GLM-5 is Z.ai's open-weights model family for complex systems engineering and long-horizon agentic tasks. The repository documents four generations: GLM-5, GLM-5.1, GLM-5.2 and GLM-5.3, plus GLM-5.3-Flash. GLM-5 scales to 744B parameters with 40B active and was pre-trained on 28.5T tokens.

How much does GLM-5 cost?

The README does not list prices. It points to the Z.ai API Platform for GLM-5.3 and GLM-5.3-Flash API services, and to z.ai for the product, so pricing has to come from those pages. The weights themselves are published under Apache-2.0, which is what you would self-host.

How do I install GLM-5.2 locally?

The repository does not document local installation of the model. It ships a requirements.txt pinning transformers>=5.15.0, accelerate>=1.14.0 and pre-commit>=4.6.2, and a download table with links, model sizes and precision. Serving instructions are not in the README.

How do I use GLM-5.2 in Claude Code?

The README does not cover Claude Code integration. It directs API users to the Z.ai API Platform at docs.z.ai/guides/llm/glm-5.3, which is where the endpoint and model identifier would be documented. Any client-side wiring is outside this repository.

Is GLM a Chinese model?

The README does not state a country of origin. It does link a WeChat community group alongside Discord, and the model is served through the Z.ai API Platform at z.ai. The technical report is hosted on arXiv.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. zai-org/GLM-5 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zai-org-glm-5.svg)](https://hysenlabs.com/projects/zai-org-glm-5)