OpenCLIP: Open-Source CLIP Training and Inference for ML Researchers
An open source implementation of CLIP.
At a glance
- What is it?
- OpenCLIP (open_clip_torch) is an open-source reimplementation of CLIP that ships a full training pipeline, multiple architecture extensions, and HuggingFace Hub weight loading. It targets ML researchers who need to train or fine-tune image-text contrastive models on custom data rather than relying on fixed pretrained weights.
- Who is it for?
- ML researchers who need a full image-text contrastive training pipeline should start with the v3 branch or the latest 3.x release on PyPI if they have existing training jobs, then test the main branch separately because the default training precision changed from amp to amp_bf16 and Horovod support was removed.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OpenCLIP Solves and Who Uses It
OpenAI's original CLIP model was released as pretrained weights alongside a paper describing contrastive language-image pretraining. It provided no public training pipeline. OpenCLIP fills that gap: a full, reproducible training loop that any team can run on their own data, combined with a growing registry of community-trained pretrained checkpoints loaded directly from HuggingFace Hub. The pyproject.toml description says it is an "Open reproduction of contrastive language-image pretraining (CLIP) and related."
The primary audience is ML researchers and engineers who need to move past fixed model cards. That includes teams training on domain-specific datasets where the standard web-crawled checkpoints perform poorly, researchers reproducing published results on DataComp or LAION training sets, and developers building audio-text or generative multimodal models by extending the CLIP architecture. The library is not a production inference SDK with a simplified API surface; it is a research platform that expects its users to engage with config files, training flags, and PyTorch distributed training concepts.
How OpenCLIP Encodes Images and Text
The core architecture follows the original CLIP design. A separate image tower and a text transformer run in parallel, both producing fixed-length embeddings. Contrastive loss pulls together the embeddings for matching image-text pairs and pushes apart non-matching ones. OpenCLIP packages this as a model factory that accepts an architecture name and a pretrained tag, downloading the matching checkpoint from HuggingFace Hub on first use.
The image tower in the standard families uses ViT variants (ViT-B/32, ViT-L/14, and others). The text tower in the standard families is a transformer with a fixed context length. Both towers are exposed as separate `encode_image` and `encode_text` methods so callers can embed one modality without computing the other.
The main branch has extended this base with a modern text tower configuration. Setting `text_cfg.text_arch="modern"` switches to an encoder with RoPE positional embeddings, SwiGLU or ReLU-squared activations, RMSNorm, and masked pooling strategies including eos, mean, and map. Optional additions include qk-norm, gated attention, register tokens, sandwich norm, value residuals, and zero-init residuals. The `moderntext-*` config families use this tower. This closes the gap between the original CLIP text encoder and more recent language model design patterns without changing the contrastive training objective.
Installing open_clip_torch for Local Development
The Makefile provides the standard local development install sequence. First, upgrade pip and install the package in editable mode:
python -m pip install -U pip
python -m pip install -e .For training workloads that require WebDataset pipelines, Pandas, and the Hugging Face Transformers sentencepiece integration, the training extras are installed separately:
python -m pip install -r requirements-training.txtThe test suite can be run after installing the test dependencies:
python -m pytest -x -s -v testsThe core runtime dependencies include torch 2.6 or later, torchvision, regex, ftfy, tqdm, huggingface_hub, safetensors, and timm 1.0.29 or later. Running on Python 3.8 or below is not supported; the package classifiers list Python 3.9 through 3.12. The package name on PyPI is `open_clip_torch`, not `open_clip`, so install commands targeting the wrong name will fail silently by installing a different package.
After install, the `hf-hub:` prefix in the pretrained string lets callers load from any HuggingFace Hub repo. Research groups hosting their own fine-tuned checkpoints can use the same loading path as the main weight registry.
New Model Families Introduced on the Main Branch
The main branch diverged significantly from the v3 training API and introduced several new architecture families. The README warns that training scripts and downstream integrations should review the changes before upgrading.
NaFlex CLIP replaces the fixed-resolution image tower with a variable-resolution ViT from the timm naflexvit family. It uses token-budget batching via the `--use-naflex` flag and `--naflex-*` config options. This removes the fixed-resolution preprocessing constraint that standard ViT pipelines carry: images at different aspect ratios no longer require square-padding before encoding.
NaFlex CLAP extends the same variable-length pipeline to audio. Audio-text contrastive training accepts variable-duration audio inputs via the naflexclap_* config families, with `--audio-*` and `--audio-zeroshot-*` flags controlling preprocessing and evaluation.
The generative models, NaFlex GenLIP and NaFlex GenLAP, add prefix-LM attention for image and audio captioning within the same training script. They use packed `[media ; text]` row format with tiktoken text processing.
MaMMUT is a single text decoder used in two passes per the linked paper. Bi-directional without cross-attention for contrastive learning, causal with cross-attention for captioning. The `mammut_*` config family reproduces the original LAION-fork numerics exactly and loads the released LAION openMaMMUT-ViT-L-14 weights. The `mammut2_*` family uses corrected defaults including masked-mean text pooling and pad masking.
CoCa v2 configs (coca2_*) add paper-faithful attentional pooling and a corrected CLS and pad attention mask. Existing `coca_*` configs and released weights are unchanged.
Breaking Changes from v3 to the Current Main Branch
The README is explicit about the risk of upgrading training code from v3 to the main branch. Several changes are silent: they produce no error, but they change the output.
The default training precision changed from `amp` (float16 AMP) to `amp_bf16` (bfloat16 AMP). Runs that relied on fp16 without passing `--precision amp` explicitly now use bfloat16. On hardware where bfloat16 gives different numerical behavior than float16, this changes training dynamics without any warning.
Horovod support was removed. The only distributed training modes are DDP and FSDP2. Any training infrastructure that set `--horovod` will fail at flag parsing rather than silently. The `--torchscript` and `--trace` flags are also gone following upstream PyTorch deprecation of TorchScript support.
SigLIP's `--loss-dist-impl` now defaults to `gather` rather than the bidirectional ring exchange. Gather stores all ranks' text features on each rank, changing memory usage and communication patterns versus the previous default. Pass `--loss-dist-impl bidir` to restore the old behavior.
The `--naflex-max-tokens-per-batch` flag no longer defaults to 16384. It is unset by default, with the local token budget inferred from `--batch-size` and `--naflex-seq-lens`. Older runs that relied on the previous 16384 default need an explicit token budget to reproduce.
The safest path for existing training jobs is to pin to the v3 branch or a specific 3.x release on PyPI and test the main branch separately.
Where OpenCLIP Is the Wrong Tool
OpenCLIP is a research-grade library, not a production inference SDK. It ships no ONNX export utilities or model quantization tools in the main repository. Researchers who only need CLIP inference without any training machinery may find that the HuggingFace Transformers CLIP implementation, which uses a simpler API and a broader ecosystem of deployment tools, is easier to maintain.
The hard dependency on PyTorch 2.6 or later, stated in requirements.txt, blocks use on any infrastructure that has not upgraded beyond that version. This is not a soft constraint: the package will not install on PyTorch 2.5.
The new model families on the main branch require careful config-file alignment. NaFlex pipelines use token-budget batching parameters that interact with distributed training in non-obvious ways, and the README notes that the scope of the main branch has grown well beyond the original refactor. Teams that need a stable, documented training API should pin to the v3 branch.
Finally, the licensing situation in pyproject.toml is listed as MIT, but the top-level repository metadata shows `NOASSERTION` for the license field. Teams with strict compliance requirements should review the actual LICENSE file before incorporating the repository.
OpenCLIP vs OpenAI's CLIP Repository
OpenAI's original CLIP repository provides pretrained weights and basic inference code but no public training pipeline. OpenCLIP provides both: a training loop tested at scale on datasets including DataComp and LAION, and a growing set of pretrained checkpoints from community training runs. The functional difference is that OpenAI's CLIP releases are fixed model cards, while OpenCLIP lets a team extend the architecture (NaFlex, CLAP, a modern text tower) and run the full training job under the MIT license as stated in pyproject.toml.
The trade-off is operational complexity. The main branch is under active development, with the last push on 2026-09-25, and the training API changes with each major cycle. OpenAI's CLIP, while not publicly trainable, has a stable, well-documented inference interface that downstream projects have integrated widely. OpenCLIP's inference interface for pretrained image-text models is intended to remain backward-compatible, but the training CLI is not.
Editorial conclusion
ML researchers who need a full image-text contrastive training pipeline should start with the v3 branch or the latest 3.x release on PyPI if they have existing training jobs, then test the main branch separately because the default training precision changed from amp to amp_bf16 and Horovod support was removed. Teams that only need inference should verify whether HuggingFace Transformers' built-in CLIP classes meet their needs before adding the PyTorch 2.6 or later requirement. Anyone upgrading from v3 should audit the SigLIP loss distribution change and the FSDP2-only distributed backend before rerunning existing training scripts. The package name on PyPI is open_clip_torch.
Frequently asked questions
What does OpenCLIP do?
OpenCLIP is an open-source reimplementation of CLIP that provides both a full training pipeline and pretrained model weights. It supports image-text contrastive learning with multiple architecture variants including standard ViT-based models, variable-resolution NaFlex models, and audio-text CLAP models.
How do I install open_clip?
Install the package in editable mode for local development with `python -m pip install -e .` after upgrading pip. The package name on PyPI is open_clip_torch, not open_clip. It requires Python 3.9 or later and PyTorch 2.6 or later.
What is open_clip_torch?
open_clip_torch is the PyPI package name for the OpenCLIP library from the mlfoundations organization. It is distinct from any package that might be found under the name open_clip. The package provides image-text contrastive model training and inference.
What is the difference between open clip and CLIP?
OpenAI's CLIP repository provides pretrained weights and inference code but no public training pipeline. OpenCLIP adds a full training loop, community-trained checkpoints from datasets like DataComp and LAION, and additional model families including NaFlex variable-resolution and CLAP audio-text models.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlfoundations-open-clip)