PaddleFormers: a PaddlePaddle training stack for DeepSeek-V4, GLM-4.5 and Qwen3
PaddleFormers is an easy-to-use library of pre-trained large language model zoo based on PaddlePaddle.
At a glance
- What is it?
- PaddleFormers is PaddlePaddle's answer to Hugging Face Transformers, aimed at teams that train LLMs and VLMs rather than serve them. The README claims training performance above Megatron-LM on DeepSeek-V3 and GLM-4.5-Air, and the v1.2 release adds DeepSeek-V4 support.
- Who is it for?
- Adopt PaddleFormers if your training stack is already PaddlePaddle, or if you need DeepSeek-V4, GLM-4.5, ERNIE-4.5 or Qwen3 post-training on domestic accelerators such as Kunlun P800, Tianshu Tiangai 150 or Metax C550. Do not adopt it if you only need inference, or if your serving path is built on PyTorch checkpoints and you have no reason to move.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PaddleFormers is for, and who it is not for
The README describes PaddleFormers as a Transformers library built on PaddlePaddle, positioned as the PaddlePaddle ecosystem's counterpart to Hugging Face Transformers, with training support for large language models and vision language models. The scope is deliberate: this is a training library. The stated capabilities run from pretraining through post-training, with CPT, SFT, SFT-LoRA, DPO and DPO-LoRA named as the supported post-training methods. There is no serving runtime in the description, and no inference benchmark. If you want to run a model, the README points elsewhere: it says checkpoints are written in Safetensors format matching the weights hosted on Hugging Face, and names FastDeploy, vLLM and SGLang as tools that can consume them. That framing tells you the intended user. It is an engineer who already has a PaddlePaddle environment, or who has been told to train on domestic accelerators, and who needs the distributed machinery (tensor, pipeline and expert parallelism) without writing it. The model list backs this up: DeepSeek-V3, DeepSeek-V3.2, DeepSeek-V4-Pro and DeepSeek-V4-Flash, GLM-4.5-Air and GLM-4.5, the Qwen2 and Qwen3 families up to Qwen3-235B-A22B, gemma-3 from 270m to 27b, Llama-3 through 3.3, OLMo2, phi-4, gpt-oss-20b and gpt-oss-120b, Granite 3.2, plus Baidu's own ERNIE-4.5 line and PaddleOCR-VL. That is a pretraining and fine-tuning audience, not an application developer audience.
How the training stack is put together
The architecture visible from the repository is a Python package, paddleformers/, sitting on top of PaddlePaddle, with examples/ holding runnable configurations and best practices, tests/ covering data, datasets, generation, mergekit, nn, peft, trainer, transformers and utils, and third_party/ for vendored code. The parallelism story is the part the README spends words on. Tensor parallelism, pipeline parallelism and expert parallelism are all listed as supported, alongside automatic mixed precision. For DeepSeek-V4 specifically, the v1.2 notes describe a stack of additions: Document Mask and Packing data flow, context parallelism (CP), a FlashMLA attention operator, a fused FP8 MoE module, DeepEP and HybridEP expert-parallel communication libraries, TileLang fused operators, and MoE Auto-Subbatch for dynamic memory load balancing. Read that list as a statement about where the engineering effort went. It is not a general-purpose trainer with a plugin for MoE; it is a trainer whose optimizations assume Mixture-of-Experts models with long contexts. The v1.1 notes point the same way, adding single-step and multi-step MTP training for GLM-4.5 and a switch to freeze the backbone network when training the MTP module. The output side is simpler than the training side: models are saved in Safetensors, which the README calls full support and describes as format-compatible with Hugging Face-hosted weights.
Installation and a first run
The README's install section is a Makefile target, not a plain pip line, and the target begins by checking for CUDA. If nvcc is not on PATH it prints an error and exits, so a CPU-only machine cannot use this path at all. It then reads the CUDA version from nvcc --version and picks a PaddlePaddle package index accordingly, with explicit branches for 12.6 and 12.9 visible in the Makefile excerpt.
make installRun that from the repository root. The target echoes the detected CUDA version before selecting a source, so the first thing you should see is a line like Detected CUDA version: 12.6, or an ERROR line telling you nvcc was not found. The README also states Python 3.10+ and Linux or Windows as the supported platforms, and requirements.txt pins transformers to >=5.0.0, <5.4.0 alongside paddlecodec>=0.2, <0.3 and fast_dataindex>=0.1.1, both of which are marked Linux-only in that file. That pin on transformers is worth noticing: this library depends on the Transformers API surface while reimplementing it, so the two move together.
For a first real use, the repository ships example configurations rather than a single documented entry point. The examples directory contains config/, experiments/, tools/ and best_practices/, and the README links to examples/best_practices/DeepSeek-V4/ for the v1.2 training recipe. The practical first step is to open a config under examples/config/, set the model path to one of the identifiers in the model list (for example Qwen/Qwen3-8B or deepseek-ai/DeepSeek-V4-Flash), and check that the chat template column in the model table matches, since the table lists templates such as deepseek3, qwen, glm4_moe and ernie. A mismatch there is a data-formatting problem, not a model problem.
To run the test suite the way the project does, the Makefile defines a unit-test target that sets DOWNLOAD_SOURCE=aistudio and runs pytest with coverage over paddleformers/.
make testExpect tests to pull data from AiStudio, given that environment variable, so the suite is not fully offline.
The limits the README does not talk about
The performance claim in the README is directional and unsourced. It says training performance on DeepSeek-V3, GLM-4.5-Air and DeepSeek-V4 is clearly above Megatron-LM, and for v1.1 it gives one number: Qwen3-VL 30B-A3B improved 48 percent over the previous version and leads Megatron-LM by 6 percent. No hardware configuration, batch size, sequence length or measurement method accompanies those figures. Treat them as the project's own positioning, not as a result you can plan capacity around, and reproduce them on your own cluster before you make a decision that depends on them. The second limit is platform. The Makefile install target requires nvcc, and the dependency markers in requirements.txt restrict paddlecodec and fast_dataindex to Linux. Windows appears in the badge, but the install target as shown is CUDA-first and shell-based, so the Windows path is not demonstrated in the repository files. Third, the domestic-accelerator support that the README highlights (Kunlun P800, Tianshu Tiangai 150, Metax C550, with a claim of DeepSeek V3 SFT on 128 Kunlun P800 cards) is described in the feature list and release notes, but those are not the same as a documented, reproducible setup for each chip. Fourth, this is a young library with a fast release cadence: v1.0.0 in January 2026, 1.1.1 in April, v1.2.0 in June, and the default branch is develop, not a stable branch. The last push to the repository was on 2026-09-21, so work is ongoing, but a develop default branch means the code you clone is the code under construction.
Where PaddleFormers is the wrong choice
Pick something else if your goal is inference. Nothing in the README describes a serving path, an OpenAI-compatible endpoint, or a quantized runtime for deployment; it explicitly hands that job to FastDeploy, vLLM and SGLang by writing compatible Safetensors. Pick something else if your team's existing training code, checkpoints and tooling are PyTorch-based and there is no external reason to change. The value here is the PaddlePaddle runtime and the domestic-chip adaptations, and moving a working pipeline to a different framework to gain a performance number you have not measured is a poor trade. Pick something else if you need a small, single-GPU fine-tuning tool for a laptop. The install path wants CUDA and nvcc, the dependency set includes librosa, visualdl and modelscope, and the feature list is organised around distributed strategies. It will run, but you are paying for machinery you are not using. Finally, if your model is not in the table, the chat template column gives you no mapping, and the training support for that architecture is unverified.
How it differs from Hugging Face Transformers and Megatron-LM
The README frames the project against two references, and the differences are concrete. Against Hugging Face Transformers, the stated goal is an equivalent model interface and functional experience, but implemented on PaddlePaddle rather than PyTorch, with the distributed training strategies (tensor, pipeline, expert parallelism, automatic mixed precision) treated as built-in rather than assembled from separate libraries. The dependency on transformers>=5.0.0, <5.4.0 in requirements.txt shows the relationship is not purely parallel: the package is installed, and the reimplementation sits alongside it. Against Megatron-LM, the difference is the target hardware and the model set. Megatron-LM is a PyTorch training framework; PaddleFormers is a PaddlePaddle one, and its distinguishing additions are the domestic accelerator adaptations (Kunlun P800, Tianshu Tiangai 150, Metax C550) and first-party support for Baidu's ERNIE-4.5 and PaddleOCR-VL models, which Megatron-LM does not cover. If neither PaddlePaddle nor domestic silicon is in your picture, the comparison reduces to a performance claim you would have to verify yourself, and there is little reason to prefer this over the framework your team already knows.
Licence, releases and what upgrades cost
PaddleFormers is Apache-2.0, the same licence family as many of the frameworks it competes with, and the LICENSE file sits at the repository root with the Apache 2 badge in the README. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you keep the licence and notice files. That is the general shape of the terms, and it is not legal advice: check the LICENSE text and your own obligations, particularly if you redistribute a modified paddleformers/ package or ship a product built on it. On upgrade cost, the release history is short and dense: v1.0.0 on 2026-01-21, 1.1.1 on 2026-04-02, v1.2.0 on 2026-06-19, with the develop branch receiving pushes as recently as 2026-09-21. Each release has added model families rather than only fixing bugs, and v1.2 added a substantial set of new operators and communication libraries for DeepSeek-V4. The practical consequence is that the version you pin determines which models you can train at all, so pinning is not just a stability measure here; it is a capability decision. The transformers>=5.0.0, <5.4.0 constraint in requirements.txt is the other thing to watch, because it caps a dependency that the wider ecosystem moves quickly, and a bump there is the kind of change that reaches into model code.
Editorial conclusion
Adopt PaddleFormers if your training stack is already PaddlePaddle, or if you need DeepSeek-V4, GLM-4.5, ERNIE-4.5 or Qwen3 post-training on domestic accelerators such as Kunlun P800, Tianshu Tiangai 150 or Metax C550. Do not adopt it if you only need inference, or if your serving path is built on PyTorch checkpoints and you have no reason to move. Before committing, verify that the model you intend to train appears in the model list with a matching chat template, and check that your CUDA version is one the Makefile install target recognises, because the target aborts when nvcc is missing.
Frequently asked questions
What is PaddleFormers and which models does it support?
It is a PaddlePaddle-based Transformers library for training large language models and vision language models, described in the README as the PaddlePaddle ecosystem's counterpart to Hugging Face Transformers. The model list covers more than 100 models, including DeepSeek-V3, DeepSeek-V3.2, DeepSeek-V4-Pro and DeepSeek-V4-Flash, GLM-4.5 and GLM-4.5-Air, the Qwen2 and Qwen3 families, Llama-3, gemma-3, OLMo2, phi-4, gpt-oss, Granite 3.2, ERNIE-4.5 and PaddleOCR-VL.
How do I install PaddleFormers?
The README's install path is the Makefile install target run from the repository root, which checks for nvcc and selects a PaddlePaddle package index based on the detected CUDA version. It requires CUDA to be present and reports Python 3.10+ with Linux or Windows as supported platforms.
Can I use a PaddleFormers checkpoint with vLLM or SGLang?
The README states that models are saved in Safetensors format matching the weights hosted on Hugging Face, and names FastDeploy, vLLM and SGLang as frameworks that can use that format. PaddleFormers itself is described as a training library, with no serving runtime documented.
Does PaddleFormers support domestic AI accelerators?
The feature list names Kunlun P800, Tianshu Tiangai 150 and Metax C550 as supported domestic compute platforms, and the v1.0 notes mention PaddleOCR-VL adaptation on Kunlun P800 and Tianshu Tiangai 150. The README also claims DeepSeek V3 SFT on 128 Kunlun P800 cards as a post-training configuration.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paddlepaddle-paddleformers)