Model or dataset
PaddlePaddle/PaddleNLP avatar
PaddlePaddle/PaddleNLP

PaddleNLP: LLM Training, Compression, and Inference Toolkit on PaddlePaddle

Easy-to-use and powerful LLM and SLM library with awesome model zoo.

12,977 stars3,026 forksPythonApache-2.0

At a glance

What is it?
PaddleNLP is an Apache 2.0 toolkit for large language model development on the PaddlePaddle framework, covering 4D parallel training, FP8 and 4-bit quantization, and multi-hardware inference across NVIDIA GPUs and several Chinese AI chips.
Who is it for?
PaddleNLP is the right choice for teams already inside the PaddlePaddle ecosystem who need a full training-to-deployment pipeline with explicit support for Kunlun XPU, Ascend NPU, and other Chinese AI chips alongside NVIDIA GPUs. Teams working entirely within the Hugging Face ecosystem will find the integration overhead significant: PaddleNLP uses its own Trainer, tokenizer, and checkpoint format.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 129 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What PaddleNLP Is and Who It Targets

PaddleNLP is a large language model development toolkit built on the PaddlePaddle deep learning framework from Baidu. The README describes its purpose as assisting developers with high-efficiency industrial-scale LLM deployment. It covers the full pipeline: pretraining, supervised fine-tuning, alignment, quantization, and serving.

The primary audience is engineering teams that need to train or serve large models on hardware configurations that include non-NVIDIA chips. PaddleNLP explicitly supports NVIDIA GPU, Kunlun XPU, Ascend NPU, Suiyuan GCU, and Hygon DCU. This multi-hardware support is a distinguishing design goal. The API for switching between these hardware backends is described in the README as requiring minimal code changes.

The model zoo covers LLaMA series, Baichuan series, Bloom series, ChatGLM series, Gemma series, Mistral series, OPT series, and Qwen series. As of the 2025-04-29 update noted in the README, Qwen3 series models including Qwen3-235B-A22B and Qwen3-30B-A3B are also supported. DeepSeek V3/R1/R1-Distill and QwQ-32B were added in the v3.0 Beta4 release in March 2025.

4D Parallel Training Architecture

The README describes a 4D parallel training configuration that combines four strategies simultaneously: data parallelism, sharded data parallelism (grouping parameter slices), tensor model parallelism, and pipeline model parallelism. The Trainer API accepts distributed strategy configuration, which the README describes as reducing the complexity of combining these four modes.

Unified Checkpoint is a custom storage format for distributed training. The README states it accelerates checkpoint saves by 95% and reduces storage space by 78.5% compared to standard approaches. It also supports dynamic resource scaling: a checkpoint saved on one machine configuration can be resumed on a different number of nodes or GPUs. Asynchronous saving is part of the mechanism.

FlashMask is a proprietary attention mask representation that the README describes as reducing memory consumption during training by using a columnar sparse attention mask format. It is listed as one of the techniques that enabled DeepSeek-R1 training on PaddlePaddle with reduced GPU memory usage.

Installing PaddleNLP and Running a First Training Job

The Makefile in the repository shows the standard install sequence:

bash
pip install --pre paddlepaddle -i https://www.paddlepaddle.org.cn/packages/nightly/cpu/
pip install -r requirements-dev.txt
pip install -r requirements.txt

The requirements.txt includes datasets 3.6.0, pyarrow 20.0.0, huggingface_hub 0.19.2 or later, safetensors, tokenizers, fastapi, uvicorn, and several other packages including jieba for Chinese text processing. A pre-commit hook is installed as part of the dev setup.

The llm/ directory in the repository contains the large model training and inference code. The README points to https://paddlenlp.readthedocs.io for full documentation. For inference deployment, the README mentions a one-click deployment command described in the inference documentation at https://paddlenlp.readthedocs.io/zh/latest/llm/server/docs/general_model_inference.html, which covers popular models including DeepSeek.

The v3.0 Beta4 release notes in the README state that single-machine FP8 inference outputs over 1000 tokens per second and 4-bit inference outputs over 2100 tokens per second. These numbers appear in the release notes without specifying the hardware configuration.

Quantization and High-Performance Inference

PaddleNLP's inference module supports FP8, INT8, and 4-bit quantization for supported models. The README notes that DeepSeek V3/R1 full models support FP8, INT8, and 4-bit quantization with MTP speculative decoding. Quantization is integrated into the serving workflow rather than requiring a separate conversion step.

The PP-UIE model (a proprietary information extraction model from Baidu) supports up to 8192 tokens for document-level information extraction. According to the README, its training efficiency is 1.8 times faster than LLaMA-Factory on comparable tasks. Zero-shot and few-shot learning are described as supported for cold-start scenarios.

MergeKit is a model merging utility bundled with PaddleNLP. The README describes it as addressing alignment tax, the performance degradation that can occur after alignment fine-tuning. The RsLoRA+ algorithm, introduced in v2.8, is described in the README as significantly improving PEFT training convergence speed.

The PaddlePaddle Dependency as a Hard Constraint

PaddleNLP is not framework-agnostic. Every component requires PaddlePaddle. Teams using PyTorch natively cannot use PaddleNLP without switching frameworks or running a compatibility layer. The requirements.txt lists paddle as a known third-party package in the isort configuration, but there is no PyTorch path in the codebase.

The documentation is primarily in Chinese. The README includes an English version linked as README_en.md, but the release notes and community communication channels listed in the README are in Chinese. Teams without Chinese-language engineering support may find the documentation gap significant when troubleshooting production issues.

The repository has a docs/ directory and a readthedocs configuration, but the English documentation at paddlenlp.readthedocs.io is described as separate from the Chinese documentation. Gaps between the two exist, and the release notes in the README are not translated.

PaddleNLP vs Hugging Face Transformers

Hugging Face Transformers is the dominant open-source LLM library outside China, built on PyTorch and JAX with a large community, thousands of pretrained model weights, and extensive documentation in English. It is hardware-agnostic within the PyTorch/CUDA ecosystem and has broad support for export to ONNX, TensorRT, and other inference formats.

PaddleNLP's differentiation is hardware breadth: it supports Kunlun XPU, Ascend NPU, and other Chinese AI chips natively, which Hugging Face Transformers does not. The Unified Checkpoint format and 4D parallel configuration are also specific features not present in the core Transformers library. For teams deploying on hardware environments that mix NVIDIA and non-NVIDIA chips, PaddleNLP is designed to handle that configuration within a single toolkit.

The model compatibility list overlaps significantly with Hugging Face, since PaddleNLP supports the same model families. Migration between the two requires converting checkpoint formats and adapting training scripts to the PaddlePaddle Trainer API.

Editorial conclusion

PaddleNLP is the right choice for teams already inside the PaddlePaddle ecosystem who need a full training-to-deployment pipeline with explicit support for Kunlun XPU, Ascend NPU, and other Chinese AI chips alongside NVIDIA GPUs. Teams working entirely within the Hugging Face ecosystem will find the integration overhead significant: PaddleNLP uses its own Trainer, tokenizer, and checkpoint format. Before adopting it, verify that the models you need are in the supported list. The last push to the repository was on 2026-05-23, and the most recent release is rl-v1.0.0 from 2025-05-21.

Frequently asked questions

What hardware does PaddleNLP support?

The README lists NVIDIA GPU, Kunlun XPU, Ascend NPU, Suiyuan GCU, and Hygon DCU as supported hardware for both training and inference. The API is described as allowing fast switching between hardware backends.

What models does PaddleNLP support?

The README lists LLaMA, Baichuan, Bloom, ChatGLM, Gemma, Mistral, OPT, and Qwen series models. DeepSeek V3/R1/R1-Distill, QwQ-32B, and the Qwen3 series including Qwen3-235B-A22B were added in 2025.

How does PaddleNLP's Unified Checkpoint differ from standard checkpoint saving?

The README states that Unified Checkpoint speeds up model saves by 95% and reduces storage space by 78.5% compared to the standard approach. It also supports restoring training with a different number of nodes or GPUs than were used when the checkpoint was saved.

Official sources

  1. License: Apache-2.0
  2. PaddlePaddle/PaddleNLP on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/paddlepaddle-paddlenlp.svg)](https://hysenlabs.com/projects/paddlepaddle-paddlenlp)
Community notes

Community notes