Open-source project
OFA-Sys/Chinese-CLIP avatar
OFA-Sys/Chinese-CLIP

Chinese-CLIP: cross-modal retrieval and embeddings for Chinese text and images

Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

6,011 stars551 forksJupyter NotebookMIT

At a glance

What is it?
OFA-Sys/Chinese-CLIP packages five pretrained Chinese image-text models, an embedding API, and finetuning scripts on top of open_clip. It is a good fit when your captions, queries and labels are Chinese; it is not a general-purpose vision library.
Who is it for?
Adopt Chinese-CLIP if your queries, captions or class names are Chinese and you need zero-shot retrieval or classification without training a model from scratch. Skip it if your content is English (original CLIP or open_clip is the natural choice) or if you only need object detection rather than image-text matching.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Chinese-CLIP fills: CLIP's text encoder does not read Chinese

Original CLIP was trained on English image-text pairs, and its text encoder is a byte-pair tokenizer built for English. Feed it Chinese captions and the tokenization is poor, so retrieval quality drops. Chinese-CLIP retrains the contrastive objective on large-scale Chinese data, described in the README as roughly 200 million image-text pairs, and swaps the text backbone for Chinese-aware encoders: RBT3 for the smallest model and RoBERTa-wwm variants for the rest. The vision side stays conventional (ResNet50 or ViT at 224 or 336 resolution).

The target user is an engineer building Chinese-language search, tagging or classification: a product search box where users type Chinese, a moderation queue that needs to match Chinese captions to images, or a zero-shot classifier where the labels are Chinese words. If your data is English, this project adds nothing over the models it was derived from.

Five checkpoints, and the memory cost each one implies

The README lists five open models. chinese-clip-rn50 is 77M parameters (ResNet50 vision, RBT3 text) at 224 resolution. chinese-clip-vit-base-patch16 is 188M. chinese-clip-vit-large-patch14 is 406M, and the 336px variant is 407M at 336 resolution. chinese-clip-vit-huge-patch14 is 958M, with a 632M ViT-H/14 vision tower and a 326M RoBERTa-wwm-Large text tower.

That spread matters more than the retrieval numbers. The huge model is more than twelve times the parameter count of the RN50 model, and the 336px variant processes roughly 2.25 times the pixels of the 224px one at the same depth. The README does not publish per-model latency or memory figures, so the parameter counts and input resolution are the only sizing signals available before you run anything. Weights are hosted on Hugging Face Hub and ModelScope, and the README links both for every checkpoint.

Installing cn_clip and computing a first image-text similarity

The repository ships a setup.py that installs the package as cn_clip, version 1.5.1, pulling dependencies from requirements.txt (numpy, tqdm, six, timm, lmdb==1.3.0, torch>=1.7.1, torchvision). The README describes a quick API for computing image and text features and their similarity in a few lines.

Install from the repository root:

bash
pip install -r requirements.txt
python setup.py install

The README also notes that from 2022-12-01 the model code and feature extraction API were merged into the Hugging Face transformers library, so you can load a checkpoint through transformers instead of the local package. The README does not spell out the full transformers snippet, so check the model card on the Hub for the exact loading call before relying on that path.

After installation, the intended first use is the embedding API: encode an image and a set of Chinese candidate captions, then compare the vectors. The repository's examples directory contains pokemon.jpeg plus three image_retrieval_result images, which are the worked example of image-to-text retrieval. The README gives the API as a short code path but does not reproduce it inline in the section that was captured, so the safest first step is to open the API section of README_En.md and run its example verbatim before adapting it.

Where the design shows strain: batch size, tokenization and deployment

Contrastive training depends on batch size, because the loss is computed over in-batch negatives. The README added gradient accumulation on 2023-03-20 specifically to simulate larger batches when memory is limited. That is a workaround, not a fix: accumulated gradients approximate a large batch but do not give you the same negative pool, and the README does not claim they do.

Tokenization is the second constraint. The text towers are Chinese-specific, so the models are not a drop-in replacement for multilingual retrieval. If your corpus mixes Chinese and English queries, the README gives no guidance on how the tokenizer handles the English half.

Third, deployment. ONNX and TensorRT export arrived on 2023-01-15, with pretrained TensorRT models offered, and a PyTorch-to-CoreML conversion script was added on 2023-11-30. Those are separate documents (deployment.md, and the script under cn_clip/deploy). The README does not document rollback or version pinning for exported artifacts, so treat conversion as a one-way step you should validate against the PyTorch output yourself.

The alternative worth comparing: open_clip and the original CLIP

Chinese-CLIP is built on open_clip, and the README states that directly. The difference in approach is the training data and the text encoder. open_clip gives you the architecture and a training framework with English-centric pretrained weights; Chinese-CLIP gives you weights already trained on Chinese pairs plus Chinese text backbones. If you want to train your own Chinese model, open_clip is the lower-level base and Chinese-CLIP's training scripts are the concrete example of how to do it. If you just want Chinese retrieval to work today, Chinese-CLIP saves you the training run.

The README's own MUGE text-to-image comparison puts CN-CLIP at 63.0 R@1 zero-shot against Wukong at 42.7 and R2D2 at 49.5, and at 68.9 R@1 after finetuning. Those are the project's reported numbers on the official validation set, not an independent evaluation, but they are the clearest statement of what the Chinese training data buys you.

Licence and the cost of keeping up

The repository is MIT licensed, with MIT-LICENSE.txt at the top level. MIT covers the code in the repository. The model weights are hosted on Hugging Face Hub and ModelScope and are downloaded separately, so the terms attached to a given checkpoint are a separate question from the code licence. The README does not restate weight terms, so check the model card for the checkpoint you actually deploy. This is a description of what the repository contains, not legal advice.

The maintenance picture is mixed in a specific way. The last push to the repository was on 2026-03-31, which is recent. But the README's news section stops at 2023-11-30, and there have been no tagged releases retrieved. The feature work described in the news (FlashAttention, FLIP, distillation, ONNX/TensorRT, CoreML) all landed between 2022 and 2023. So the code sees activity while the documented feature set has been stable for a long time. Upgrading is cheap in the sense that the API surface has not churned; it is uncertain in the sense that you cannot tell from the README what changed after 2023.

Finetuning and zero-shot classification, and when to skip both

Two downstream tasks are documented with code. Cross-modal retrieval has training and evaluation scripts under run_scripts, and the README reports finetune results on MUGE, Flickr30K-CN and COCO-CN. Zero-shot image classification has its own section and supports the ELEVATER benchmark; the Chinese version of the ELEVATER classification datasets was released on 2022-12-03 and is described in zeroshot_dataset.md.

Zero-shot classification is the cheapest thing to try: write your class names in Chinese, encode them as text, encode the images, and take the nearest class. No training data, no labels. The failure mode is that class names are short and out of distribution compared to captions, and the README offers no prompt-engineering guidance for Chinese label wording. If accuracy on your label set is poor, the documented next step is finetuning, which brings you back to the batch-size and GPU-memory constraints above.

Skip the whole project if your task is detection, segmentation or OCR. Chinese-CLIP produces image and text embeddings; it does not localize objects in an image, and nothing in the repository suggests it does.

Editorial conclusion

Adopt Chinese-CLIP if your queries, captions or class names are Chinese and you need zero-shot retrieval or classification without training a model from scratch. Skip it if your content is English (original CLIP or open_clip is the natural choice) or if you only need object detection rather than image-text matching. Before committing, verify the model size against your GPU memory using the parameter counts in the model table, confirm that the tokenizer handles your text domain, and check that the licence terms of the specific checkpoint you download match your intended use, since the repository code is MIT but the weights are downloaded separately.

Frequently asked questions

What is Chinese-CLIP?

It is the Chinese version of the CLIP model, trained on large-scale Chinese data (the README says roughly 200 million image-text pairs) to support Chinese image-text feature and similarity computation, cross-modal retrieval, and zero-shot image classification. The code is built on the open_clip project and optimized for Chinese data.

What is meant by CLIP, and what does a CLIP model do?

CLIP is a contrastive image-text model: it encodes images and text into a shared embedding space so that similarity between them can be computed directly. Chinese-CLIP applies that mechanism to Chinese text, which is what makes retrieval and zero-shot classification with Chinese captions and labels possible.

How do I use Chinese-CLIP?

Install the cn_clip package from the repository (pip install -r requirements.txt, then python setup.py install), then use the feature extraction API described in the README to compute image and text features and their similarity. The README also notes the model code and API were merged into Hugging Face transformers on 2022-12-01, and weights are hosted on Hugging Face Hub and ModelScope.

Official sources

  1. Issues
  2. License: MIT
  3. OFA-Sys/Chinese-CLIP on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ofa-sys-chinese-clip.svg)](https://hysenlabs.com/projects/ofa-sys-chinese-clip)