Open-source project
openai/CLIP avatar
openai/CLIP

OpenAI CLIP: zero-shot image classification without task-specific training

CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image

34,376 stars4,047 forksJupyter NotebookMIT

At a glance

What is it?
OpenAI's CLIP repository ships the code and pretrained weights for a model that scores text against an image directly. This review covers how the contrastive mechanism works, how to install it, where it fails, and how it differs from OpenCLIP.
Who is it for?
Adopt openai/CLIP when you need a working zero-shot classifier with a small amount of code and you are comfortable pinning PyTorch and torchvision yourself. Do not adopt it if you need large ViT-L/14 or ViT-G/14 checkpoints, a supported training script, or a maintained release process, because the repository has no published releases and its README documents inference only.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What openai/CLIP solves, and who actually needs it

The problem CLIP addresses is the cost of labels. A conventional image classifier needs a fixed set of classes and enough labeled examples per class to train a head. CLIP instead learns a shared embedding space from (image, text) pairs, so at inference time you supply the candidate labels as plain strings and read off a probability. The README states that CLIP matches the performance of the original ResNet50 on ImageNet zero-shot without using any of the 1.28M labeled examples. That claim is the whole pitch: no fine-tuning run, no annotation budget, no fixed label set baked into the weights.

The audience is therefore narrow but real. Researchers who want a frozen vision backbone for probing, engineers prototyping a search or routing feature over an image collection, and anyone who needs a baseline before deciding whether to invest in a trained model. It is not a product. There is no server, no endpoint, no CLI. The repository is a Python package plus notebooks, and the primary language listed is Jupyter Notebook, which tells you where the intended use lives.

The contrastive mechanism behind clip.load and the similarity scores

CLIP has two encoders. The vision tower turns an image into a feature vector; the language tower turns a token sequence into a feature vector. The README describes the joint call precisely: model(image, text) returns logit scores that are cosine similarities between the corresponding image and text features, multiplied by 100. Classification is then just a softmax over the text axis.

That multiplication by 100 matters more than it looks. Cosine similarity lives in [-1, 1], and a softmax over such a narrow range produces nearly uniform probabilities. Scaling by 100 sharpens the distribution, which is why the README example prints a label probability of about 0.9927937 for the top class rather than something close to 0.33. If you reimplement the scoring yourself and forget the scale factor, your confidence numbers will be wrong even though your ranking is right.

The two encoders are also usable separately, and this is the part most people miss. encode_image and encode_text return the raw features. Those vectors are what you cache when you want to classify thousands of images against a fixed label set, or build a nearest-neighbour search. Calling the joint model on every pair is simpler but recomputes text features on every batch.

Installing openai/CLIP and running a first zero-shot prediction

The README's install path assumes conda and a CUDA machine. It pins PyTorch 1.7.1 (or later) and torchvision, then installs the small dependencies and the package itself from GitHub. Note that the package is not distributed under a different name on PyPI in these instructions; the final line installs directly from the repository URL.

bash
$ conda install --yes -c pytorch pytorch=1.7.1 torchvision cudatoolkit=11.0
$ pip install ftfy regex tqdm
$ pip install git+https://github.com/openai/CLIP.git

Replace cudatoolkit=11.0 with the appropriate CUDA version on your machine, or with cpuonly when installing on a machine without a GPU, as the README instructs. After that, the first real use is loading a checkpoint and scoring one image against a few candidate strings. The README's opening example does exactly this with ViT-B/32 and three labels.

python
import torch
import clip
from PIL import Image

device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device)

image = preprocess(Image.open("CLIP.png")).unsqueeze(0).to(device)
text = clip.tokenize(["a diagram", "a dog", "a cat"]).to(device)

with torch.no_grad():
    image_features = model.encode_image(image)
    text_features = model.encode_text(text)

    logits_per_image, logits_per_text = model(image, text)
    probs = logits_per_image.softmax(dim=-1).cpu().numpy()

print("Label probs:", probs)  # prints: [[0.9927937  0.004

The first run downloads the checkpoint, so expect a pause and make sure your network allows it. The printed array has one probability per label in the order you passed them, and the values sum to one across the row.

Where CLIP breaks: context length, prompt sensitivity and the missing training path

The most concrete limitation is in the tokenizer signature. clip.tokenize takes context_length=77, and the README gives no truncation or chunking strategy. A long caption, a paragraph of product copy, or a multi-sentence query will be cut to fit. There is no documented warning when this happens, so a silently truncated prompt can degrade a score without any error surfacing.

Prompt sensitivity is the second trap. The README's own CIFAR-100 example wraps each class name as f"a photo of a {c}" rather than passing the bare label. That is not decoration. The model was trained on captions, and a bare noun is a different distribution from a caption fragment. Teams that benchmark CLIP with bare labels and then wonder why accuracy lags published numbers have usually skipped this step.

The third limitation is scope. The repository documents inference. There is no training script in the README, no data pipeline, no evaluation harness beyond the two notebook-style examples. If your goal is to train a CLIP-style model on your own image-text pairs, this repository is the wrong tool and you should look at the training-focused forks instead. The README also points to OpenCLIP for larger models, which is an admission that this codebase does not cover that ground.

OpenCLIP and the Hugging Face port: same idea, different coverage

The README's See Also section names two alternatives, and the difference is not cosmetic. OpenCLIP, from mlfoundations, includes larger and independently trained CLIP models up to ViT-G/14. The openai/CLIP package loads the models returned by clip.available_models(), and the README's examples use ViT-B/32. If your accuracy target needs a larger backbone, the official repository does not offer it.

The Hugging Face implementation of CLIP takes a different route: it targets easier integration with the HF ecosystem. That means the transformers model classes, the Hub for weight distribution, and the surrounding tooling for tokenizers and pipelines. The trade-off is that you inherit the transformers abstractions and their versioning rather than calling clip.load directly.

A practical way to think about it: openai/CLIP is the reference implementation with the smallest surface area, OpenCLIP is the one that scales up, and the Hugging Face port is the one that fits an existing transformers codebase. They are not competing on the same axis.

Maintenance, licence and what upgrading costs you

The repository is not archived, but no last push date is published alongside it, so there is no basis for calling it actively maintained. Treat the version string in setup.py, which is fixed at 1.0, as the honest signal: this is a research release, not a package with a release cadence. No releases are listed for the repository.

The licence is MIT, which is permissive and, unlike some model releases, does not add a separate usage policy in the repository root. That is a genuine advantage for commercial embedding, but the licence covers the code in this repository. The model card file is a separate document, and the README links to it, so read it before you ship anything. This is not legal advice; the point is that the code licence and the model terms are two different files here.

The upgrade cost is dominated by the PyTorch pin. The README specifies 1.7.1 or later and pairs it with a matching cudatoolkit version. Moving to a newer PyTorch means re-checking the CUDA build, and the package itself has no versioned releases to migrate between. Your real upgrade surface is torch and torchvision, not clip.

Editorial conclusion

Adopt openai/CLIP when you need a working zero-shot classifier with a small amount of code and you are comfortable pinning PyTorch and torchvision yourself. Do not adopt it if you need large ViT-L/14 or ViT-G/14 checkpoints, a supported training script, or a maintained release process, because the repository has no published releases and its README documents inference only. Before you commit, run clip.available_models() on your target machine, confirm the checkpoint download works behind your proxy, and check whether your labels need prompt engineering such as the "a photo of a {c}" format the README uses, since that string is the difference between a usable score and a misleading one.

Frequently asked questions

What is OpenAI's CLIP?

CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet given an image, without directly optimizing for the task.

What is CLIP pretraining?

The name expands to Contrastive Language-Image Pre-Training. The model learns a shared space between image and text features during pretraining, which is what lets it score an image against arbitrary text labels at inference time.

How do I install openai/CLIP?

The README installs PyTorch 1.7.1 or later with torchvision via conda, then pip install ftfy regex tqdm, then pip install git+https://github.com/openai/CLIP.git. On a machine without a GPU, replace the cudatoolkit version with cpuonly.

Is openai/CLIP free and open source?

The repository is licensed under MIT, which is a permissive open source licence. The README also links to a separate model card document, so the code licence and the model terms are not the same file.

What is openai/CLIP used for?

The README demonstrates zero-shot prediction over a label set and linear-probe evaluation, where logistic regression is fit on top of the encoded image features. Both use a frozen model; the repository documents inference rather than training.

What is openai/CLIP ViT-B/32?

It is one of the model names accepted by clip.load, and it is the checkpoint used in every example in the README. The full set of names is returned at runtime by clip.available_models().

Official sources

  1. Official README
  2. Project repository