GLM-V: the archived GLM-4.6V, 4.5V and 4.1V weights, and what to check before you pull them
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
At a glance
- What is it?
- The zai-org/GLM-V repository holds the open GLM-4.6V, GLM-4.5V and GLM-4.1V vision-language models plus the VLM Reward System used to train the 4.1V-Thinking line. The repository itself stopped being maintained on 2026-09-02, so the useful question is not whether to adopt it but which of its parts you still want.
- Who is it for?
- Use GLM-V if you need open weights for grounding, image-to-text or video understanding and you are content to pin a revision, because the repository stopped being maintained on 2026-09-02 and no further fixes are coming. Do not expect it to be a service you can file issues against, and do not treat it as the home of GLM-5.3-Flash, which the README points at the separate GLM-5 repository.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What GLM-V actually ships, and who it is for
GLM-V is the open-weight side of Zhipu's vision-language work. The README states plainly that the repository contains the GLM-4.6V, GLM-4.5V and GLM-4.1V series models, and the download table lists five weights you can pull today: GLM-4.6V, GLM-4.6V-FP8, GLM-4.6V-Flash, GLM-4.5V, GLM-4.5V-FP8, plus GLM-4.1V-9B-Thinking and GLM-4.1V-9B-Base. Each is mirrored on Hugging Face and ModelScope, and the README notes that GGUF builds are published through a separate ggml-org collection, which is the route for people running local inference rather than a hosted endpoint. The audience is narrow and specific. You are building grounding, image-to-text or video-understanding features, you want the weights rather than an API key, and you can read a model card and a conversation template without hand-holding. The topics on the repository say the same thing: image2text, reasoning, video-understanding, vlm. If you want a managed endpoint instead, the README points at the z.ai API platform for GLM-5.3-Flash, which is a different model line. That distinction matters more than it looks, because the two get conflated constantly.
The split between this repository and the transformers implementation
The most easily missed design fact is that GLM-V does not contain the model algorithm. The README says the GLM-4.5V and GLM-4.6V implementation lives in transformers under src/transformers/models/glm4v_moe, and that GLM-4.1V-9B-Thinking lives under src/transformers/models/glm4v. So the data flow is: weights and preprocessing come from here, the architecture and the forward pass come from the transformers release you install. The README also warns that both model families share identical multimodal preprocessing but use different conversation templates, and asks you to distinguish carefully. That is a real failure mode rather than a footnote. If you copy a prompt template from a GLM-4.1V example and point it at a GLM-4.6V checkpoint, the preprocessing will accept your input and the template will be wrong, which produces bad answers rather than an error message. The repository layout reflects the same split: inference/, examples/, glmv_reward/ and skills/ are here, while the model code is not. Treat GLM-V as a distribution and tooling repository, not as a self-contained library.
Installing the dependencies and running a first grounding call
The repository pins its dependencies in requirements.txt. The notable entries are torch>=2.10.0, torchvision>=0.25.0, transformers>=5.13.0, accelerate>=1.14.0, plus av and torchcodec for video, and PyMuPDF for document input. Install them before anything else.
pip install -r requirements.txtThat pulls a recent torch and a transformers release new enough to contain the glm4v modules. If your environment pins an older transformers, the model classes will simply not exist, and you will get an import error rather than a graceful fallback.
The grounding example is the clearest first use. The README gives this prompt shape, where <expr> is the description of the target object:
Help me to locate <expr> in the image and give me its bounding boxes.The output is a quadruple [x1,y1,x2,y2] for the top-left and bottom-right corners, with each x normalized by image width and each y by image height, then scaled by 1000. So a box reported as [250,100,750,900] means the object starts a quarter of the way across the image and ends three quarters of the way across. The README also documents a second prompt form for requesting a specific output format, and mentions special tokens in the response, though the README text available here is truncated at that point. For a desktop debugging loop, the README points at the GLM-4.5V demo app, downloadable as an installer from a Hugging Face Space or buildable from source under examples/vllm-chat-helper/README.md. That app captures your screen via screenshots or recordings and sends them to GLM-4.5V.
The VLM Reward System is the part most people ignore
Buried in the project updates is the piece with the longest shelf life. On 2025-07-16 the project open-sourced the VLM Reward System used to train GLM-4.1V-Thinking, with code under glmv_reward/ and a runnable demo. The README gives the command:
python examples/reward_system_demo.pyA reward model for vision-language outputs is a narrower artifact than a chat model, and it is the kind of thing that keeps working after the base weights are superseded, because you are using it to score outputs rather than to generate them. The same pattern shows up in the downstream releases the README lists: Glyph, a framework for scaling context length through visual-text compression, and UI2Code^N, a UI coding model with UI-to-code, UI-polish and UI-edit capabilities. Both are described as trained on GLM-4.1V-Base. That tells you the 4.1V base was treated as a starting point for further training rather than as a finished product, and it is the strongest argument for cloning this repository even now. If you are doing reinforcement learning on multimodal tasks, the reward code is the reason to be here. If you are not, you will scroll past it.
Hardware paths, GGUF, and where this is the wrong tool
The repository carries examples for non-NVIDIA hardware, with directories for AMD_GPU and Ascend_NPU, and separate example trees for gui-agent, midscene-ts-demo and midscene-yaml-demo. That is broader hardware coverage than most VLM repositories bother with, and it is worth checking those directories before you assume your accelerator is unsupported. The GGUF route is the other escape hatch, published through the ggml-org collection on Hugging Face rather than hosted here.
Where GLM-V is the wrong tool is easier to state. It is not a serving stack. There is no documented batching server, no documented quantization pipeline beyond the FP8 and GGUF weights that already exist, and the README does not document rollback or version-pinning guidance for the weights. If you need a stable endpoint with an SLA, the README itself directs you to the API platform instead. It is also not a general chat model repository: the whole point is multimodal input, and if your task is text-only you are paying for vision preprocessing you will never use. Finally, the repository stopped being maintained on 2026-09-02, and the README says so directly, pointing GLM-5.3-Flash questions at the separate GLM-5 repository. Anything you build here sits on a frozen base.
GLM-V against a general-purpose VLM stack
The obvious alternative for someone who wants vision-language inference without model-specific plumbing is a general VLM serving stack such as vLLM, which the repository itself references through examples/vllm-chat-helper. The difference in approach is real. GLM-V gives you the weights, the preprocessing contract, the conversation templates and the reward code, and leaves serving to you or to a tool you bring. A serving framework gives you the runtime, the batching and the OpenAI-compatible surface, and expects you to supply weights it already knows how to load. If your goal is to train, fine-tune or score, GLM-V is the side with the artifacts. If your goal is to put a vision endpoint behind a load balancer this afternoon, the serving stack is the side with the artifacts, and GLM-V becomes a source of weights for it. The two are not competitors so much as different halves of the same pipeline, and the repository's own examples directory acknowledges that by shipping a vLLM helper. Choosing wrongly here usually means writing a serving layer that already exists.
Licence and the cost of adopting a frozen repository
The repository is Apache-2.0, which is permissive and permits commercial use, modification and redistribution provided you keep the notices. That covers the code in this repository. It does not automatically tell you the terms attached to each weight file on Hugging Face or ModelScope, which are separate downloads with their own model cards, and it does not cover the GGUF conversions published by a third party. Check the licence on the specific artifact you pull, not on the repository that links to it. This is not legal advice; read the model card yourself.
The upgrade cost is the more practical concern. The README states that the repository will no longer be maintained as of 2026-09-02, so there is no upstream to pull fixes from and no issue tracker that will produce patches. The requirements.txt pins are a snapshot of what worked at that date: torch>=2.10.0, transformers>=5.13.0 and the rest. As those lower bounds move forward, nothing here will be updated to follow. The safe posture is to pin exact versions in your own environment rather than tracking the minimums, and to vendor a copy of the weights you depend on rather than assuming a Hub revision stays where it is.
Editorial conclusion
Use GLM-V if you need open weights for grounding, image-to-text or video understanding and you are content to pin a revision, because the repository stopped being maintained on 2026-09-02 and no further fixes are coming. Do not expect it to be a service you can file issues against, and do not treat it as the home of GLM-5.3-Flash, which the README points at the separate GLM-5 repository. Before you commit, verify which of GLM-4.6V, GLM-4.6V-FP8, GLM-4.6V-Flash, GLM-4.5V or GLM-4.1V-9B-Thinking you actually need, that its conversation template matches the code path you are calling, and that your transformers version satisfies the transformers>=5.13.0 pin in requirements.txt.
Frequently asked questions
What is GLM-V?
It is the open-source repository for the GLM-4.6V, GLM-4.5V and GLM-4.1V series vision-language models, containing model downloads, inference and example code, skills, and the VLM Reward System used to train GLM-4.1V-Thinking.
What is a GLM vs an LLM?
The README frames GLM-V as a vision-language model family, meaning it takes images and video as well as text and is aimed at multimodal reasoning, grounding and video understanding, rather than being a text-only language model.
What is GLM used for?
According to the README, the GLM-V models are used for multimodal perception and reasoning tasks including object grounding with bounding boxes, image-to-text, video understanding and multimodal agents, with a desktop assistant app provided for debugging.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zai-org-glm-v)