ChatGLM-6B: transformers pinned to 4.27.1, no official app, and a form before commercial use
ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型
At a glance
- What is it?
- ChatGLM-6B is an Apache-2.0 bilingual Chinese-English dialogue model of 6.2 billion parameters that can run in 6GB of VRAM at INT4. The repository stopped moving in June 2024 and its own front page now advertises a successor, which tells you most of what you need to know about adopting it.
- Who is it for?
- ChatGLM-6B fits an academic or evaluation context, a Chinese-language dialogue experiment, or a small local deployment on a consumer GPU where 6GB of VRAM is the ceiling. It does not fit a long-document product, given the 2K context of this generation, and it does not fit anyone who wants a supported application, since the project states plainly that it has built none.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 27 months ago, on June 27, 2024.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The front page now advertises GLM-4, so this repository documents the previous generation
Read the top of the README before anything else. The first substantive section is not about ChatGLM-6B at all: it announces the newer GLM-4 models and points to a separate repository for the open-sourced GLM-4-9B series, to a hosted chat product for the latest model including GLMs and All tools, and to an API platform listing GLM-4-0520, GLM-4-air, GLM-4-airx, GLM-4-flash, GLM-4, GLM-3-Turbo, CharacterGLM-3 and CogView-3, with System Prompt, Function Call, Retrieval and Web_Search supported on GLM-4 and GLM-3-Turbo. The model documented on the rest of the page is therefore the earlier one, and the work has moved elsewhere: GLM-4 lives in its own repository and the commercial path is an API platform. The main branch here received its last push on 2024-06-27 and the repository publishes no GitHub releases. Treat this as a frozen artifact kept available, not as the place where the project is being developed.
transformers is pinned to 4.27.1 while everything else is a bare name
The dependency file is eight lines and the asymmetry matters:
protobuf
transformers==4.27.1
cpm_kernels
torch>=1.10
gradio
mdtex2html
sentencepiece
accelerateExactly one package is pinned, and it is pinned to an exact version with no floor and no ceiling: transformers==4.27.1. Everything else is either a floor, as with torch>=1.10, or completely unversioned. protobuf, cpm_kernels, gradio, mdtex2html, sentencepiece and accelerate can all resolve to whatever is current when you install, which is a decade of drift away from the transformers release the code was written against. The consequence is predictable. The hard pin will block or downgrade a modern environment, and in an environment where you relax the pin to get something else working, the unversioned web stack can break the demos on its own. Install this into a dedicated virtual environment, keep the pin intact, and treat the model code as something that needs the era it was written for.
The project states outright that it has never shipped an application
There is a bolded sentence in the Chinese README that is worth translating precisely, because it is unusual to see it stated this plainly. It says the project team has not developed any application based on ChatGLM-6B, and it names the categories: no web front end, no Android app, no iOS app, no Windows app. The page then points you to a hosted platform to try larger ChatGLM models instead. The practical consequence is direct. Any mobile or desktop build carrying this model's name was made by somebody else, with their own licence, their own build of the weights and their own idea of what the disclaimers mean. The disclaimers themselves are strict: because the model is small in scale and subject to probabilistic randomness, output accuracy is not guaranteed and the model is easy to mislead, and the project accepts no responsibility for data security, reputational risk, or misuse. Those terms travel with the weights in this repository, not with a repackaged application.
Commercial use is free, but only after you register through a form
The licensing paragraph has two halves and they are often quoted as one. The ChatGLM-6B weights are fully open for academic research, and free commercial use is also permitted after you register by filling in a questionnaire on the vendor's open platform. So the term is not open source in the usual sense, it is open with a registration step, and a company that ships commercially without completing that form is outside what the project has stated. There is also a separate MODEL_LICENSE file distinct from the Apache-2.0 LICENSE that covers the code, and the README asks developers to follow it, specifically asking that the model, the code and derivatives not be used for anything that could harm the country and society, and not for any service that has not passed a security assessment and filing. For an overseas deployment, that is a legal question about a Chinese regulatory regime, and it is the one thing here that no amount of engineering resolves.
This generation has a 2K context, and the successor already admits a long-document limit
Context length is the constraint that decides whether this model can do your job. ChatGLM-6B was trained with a 2K context length. The update notes describe how the successor changed this: ChatGLM2-6B extended the base model's context from ChatGLM-6B's 2K to 32K using FlashAttention, and used an 8K context during the dialogue stage. The same paragraph then concedes a limit on the successor, noting that the current ChatGLM2-6B has limited ability to understand very long single-turn documents and that the team will work on it in later iterations. So the older model in this repository offers a quarter of the smaller of those two numbers. The consequence for planning is that document question answering, whole-file summarisation and long transcript processing are out of scope on ChatGLM-6B itself. CodeGeeX2, the code model built on ChatGLM2-6B, raises sequence length to 8192, which is a separate model in a separate repository rather than a setting you can turn on here.
Three web demos sit at the root and nothing says which one is current
The entry points at the top level are api.py, cli_demo.py, utils.py, and then web_demo.py, web_demo2.py and web_demo_old.py, plus cli_demo_vision.py and web_demo_vision.py. A file named web_demo_old.py is unambiguous about its status, and the presence of a plain web_demo.py alongside a web_demo2.py leaves the actual choice to you. That matters because the demos are the interface to the model and they are where an incompatible dependency will show up first, and none of the three are pinned to a Gradio version. The vision demos have a larger footprint still: running VisualGLM-6B through cli_demo_vision.py and web_demo_vision.py requires installing SwissArmyTransformer and torchvision, neither of which appears in the eight-line requirements file. If you are starting fresh, read the README's usage section for the canonical script rather than picking by filename, and expect to resolve the Gradio and transformers versions yourself before any of them will import cleanly.
6GB is an inference floor at INT4, and fine-tuning costs one gigabyte more
The hardware table is small and it is the most operationally useful thing in the README. It gives minimum GPU memory for two tasks across three quantisation levels. For inference, FP16 with no quantisation needs 13GB, INT8 needs 8GB and INT4 needs 6GB. For efficient parameter fine-tuning, the same three levels need 14GB, 9GB and 7GB respectively. The headline figure people quote, 6GB, is therefore the best case for running the model and is not enough to train it, and the fine-tuning path is the P-Tuning v2 implementation with its guide under ptuning. Read the column headers carefully, because these are GPU memory figures. Nothing in this table describes CPU-only execution, Apple silicon, or a system without a discrete GPU, so anyone on a laptop iGPU or a CPU-only server has to look at the third-party projects instead. The model is small enough that they exist, and they make their own claims about what is achievable.
The examples folder is screenshots, and the low-resource runners are all other people's projects
The examples directory contains image files, not runnable code: ad-writing, blog-outline, comments-writing, two email-writing samples, information-extraction, role-play, self-introduction, sport and tour-guide, all as PNGs. They are output samples, useful for judging output style and nothing else, and a reader looking for a starting script will find nothing there. The interesting part of the page is the friend-links section, which is where the reach of this model actually lives. Four projects claim to make it faster or lighter: lyraChatGLM for inference acceleration with a claim of up to 9000 or more tokens per second, ChatGLM-MNN as an MNN-based C++ implementation that assigns work between GPU and CPU based on available memory, JittorLLMs claiming FP16 operation in 3GB of memory or with no GPU at all on Linux, Windows and Mac, and InferLLM as a lightweight C++ implementation for x86 and Arm with 4GB of running memory, including on phones. None of those figures are measurements from this repository, and all four are maintained elsewhere.
Editorial conclusion
ChatGLM-6B fits an academic or evaluation context, a Chinese-language dialogue experiment, or a small local deployment on a consumer GPU where 6GB of VRAM is the ceiling. It does not fit a long-document product, given the 2K context of this generation, and it does not fit anyone who wants a supported application, since the project states plainly that it has built none. Before you build on it, confirm the exact transformers version your environment can hold, register through the stated form if the use is commercial, read MODEL_LICENSE separately from the Apache-2.0 code licence, and treat any APK or desktop build you find as third-party work with its own terms.
Frequently asked questions
What is ChatGLM used for?
ChatGLM-6B is an open bilingual Chinese and English dialogue model of 6.2 billion parameters based on the GLM architecture, trained on about 1T bilingual tokens with supervised fine-tuning, feedback self-training and human feedback reinforcement learning, and optimised for Chinese question answering. With quantisation it can run locally on a consumer GPU, needing 6GB of memory at INT4.
Is ChatGLM open source?
The weights are fully open for academic research, and free commercial use is permitted after registering by filling in a questionnaire on the vendor's open platform. The code is under Apache-2.0, with a separate MODEL_LICENSE that asks users not to employ the model or derivatives for purposes harmful to the country and society, or in services that have not passed a security assessment and filing.
Are GLM models Chinese?
They are bilingual rather than Chinese only. Training used about 1T Chinese and English tokens, and a v1.1 checkpoint added English instruction fine-tuning data to balance the two, which fixed English answers that had Chinese words mixed into them.
What is the GLM AI model?
GLM is the underlying architecture, and ChatGLM-6B is a 6.2 billion parameter dialogue model built on it. The README now points to the newer GLM-4 line, including an open-sourced GLM-4-9B series in a separate repository and a commercial API platform offering GLM-4-0520, GLM-4-air, GLM-4-flash and others.