facebookresearch/perception_models: Perception Encoder and PLM for image, video and audio
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
At a glance
- What is it?
- Meta FAIR's Perception family ships four kinds of encoder checkpoints plus a decoder line, under two separate licences. Here is what each checkpoint is for, how to install the package, and where the repository stays silent.
- Who is it for?
- Adopt it if you need a single codebase for zero-shot image and video classification, retrieval, dense prediction or audio-visual embedding, and you can live with the FAIR Noncommercial Research License that setup.py declares for the package. Do not adopt it if your product is commercial, or if you want a supported release cadence: the repository lists no releases, and the last push was on 2026-04-13.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 169 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Perception Encoder and Perception Language Model actually cover
This repository is the distribution point for two model lines from Meta AI Research, FAIR. Perception Encoder (PE) is the encoding side: it turns images, video and audio into embeddings. Perception Language Model (PLM) is the decoding side: it consumes an aligned PE and produces text. The README frames the split directly, describing PE as the encoder for image, video and audio and PLM as the decoder.
PE is not one model. The README lists four checkpoint types, each aimed at a different job. PE core is a CLIP-style model for zero-shot image and video classification and video retrieval. PE lang is aligned to a language model and is what powers PLM. PE spatial is tuned for vision-centric work such as detection, depth estimation and tracking. PE audio-visual embeds audio, video, audio-video and text into one joint space. The December 2025 update added PE-AV and PE-A-Frame to that audio-visual line.
The intended user is someone who already knows which of those four jobs they have. If you need a general-purpose vision backbone, PE spatial is the wrong pick. If you need dense prediction, PE core is the wrong pick. The repository is organised around that distinction rather than around a single default model, and the per-app READMEs under apps/pe and apps/plm are where the detail lives.
How the checkpoints and training recipe fit together
The README states that all PE variants follow the same contrastive pretraining recipe, and that the difference between them comes from how the encoder is tuned afterwards. PE core is the base contrastive model. PE spatial is the spatially tuned variant. PE lang is tuned to align with an LLM so that PLM can be built on top of it. PE audio-visual extends the same idea to a joint audio, video and text embedding space.
That shared recipe matters for evaluation, because the repository reports every family against different benchmarks. PE core is scored on IN-1k, IN-v2, IN-A, ObjectNet, COCO-T2I, Kinetics-400 and VTT-T2V. PE lang is scored on Doc VQA, InfoQA, TextVQA, MVBench, PerceptionTest and EgoSchema. A model that wins on one table is not automatically the right encoder for the other.
The PLM tables make the coupling explicit. The controlled setting pairs PE-Lang-L14-448 and PE-Lang-G14-448 with the decoder. The SotA setting uses the tiling-aligned checkpoints PE-Lang-L14-448-Tiling and PE-Lang-G14-448-Tiling, and the README carries a footnote saying those were aligned with tiling and should be used when the LLM decoder runs above 448 resolution with tiling. Choosing a lang checkpoint without checking that footnote is an easy way to get weaker numbers than the table suggests.
Installing perception_models and running a first retrieval check
The repository root carries setup.py and requirements.txt, so installation is from a clone. setup.py declares python_requires of 3.11 or newer and pulls its dependency list from requirements.txt, which pins versions such as numpy==2.1.2, timm==1.0.15, torchdata==0.11.0 and transformers>=4.48.0. Note that requirements.txt is pinned tightly; if your environment already holds a different numpy or timm, resolve that conflict before installing rather than after.
git clone https://github.com/facebookresearch/perception_models.git
cd perception_models
python -m pip install -r requirements.txt
python -m pip install -e .The editable install registers the perception_models package. The package_data entry in setup.py ships core/vision_encoder/bpe_simple_vocab_16e6.txt.gz, which is the tokenizer vocabulary the vision encoder needs, so do not strip package data from the install.
The README points at a Colab notebook for a working PE session, apps/pe/docs/pe_demo.ipynb, and at apps/pe/README.md for checkpoint-level detail. That notebook is the fastest way to see the expected call shape before you write your own loader. The README also notes that PE was integrated into timm, so if you only need the encoder as a backbone you can load it through timm instead of this repository. The README does not document a CLI entry point, so the intended surface is the Python API used in the demo notebook.
The licence split is the first thing to check, not the last
The repository root contains two licence files, LICENSE.PE and LICENSE.PLM, and a general Apache-2.0 code licence badge in the README. Those are not the same thing, and setup.py makes the distinction concrete: it sets license to "FAIR Noncommercial Research License" and classifies the package as "License :: Other/Proprietary License", even though the README's code badge says Apache 2.0.
So the code in the repository and the weights you download are governed by different documents. The README attaches an Apache-2.0 model licence badge to the PE section, and the PLM section has its own file. Anyone planning to ship a product needs to read LICENSE.PE and LICENSE.PLM themselves; the setup.py metadata alone should stop a commercial team from assuming the whole thing is Apache-2.0. This is not a legal opinion, it is a reading of what the files say.
The practical consequence is that the package metadata and the README badges disagree in emphasis. If your build pipeline checks licences automatically, it will likely flag the package as non-commercial because that is what setup.py declares.
Where this repository leaves you on your own
There are no retrieved releases. The README's Updates list is a chronological log of model and integration announcements, not a versioned changelog, and the last push was on 2026-04-13. Anyone who needs a stable tag to pin against will not find one here.
Dependency pinning is the second friction point. requirements.txt pins numpy==2.1.2, timm==1.0.15, tokenizers==0.21.1 and transformers>=4.48.0, among others, and lists scipy and sentencepiece twice. Duplicate lines are harmless to pip, but they signal a list maintained by hand. If you are adding PE to an existing training stack, expect to reconcile pins.
Third, the README does not document rollback, checkpoint deprecation, or what happens to older checkpoints when a new family lands. The July 2025 update added eight new PE checkpoints, and the December 2025 update added PE-AV and PE-A-Frame; the README does not say whether earlier checkpoints remain supported or will be removed from Hugging Face. Pin the exact checkpoint identifier you use.
Finally, this is a research release. The README describes benchmark leadership against SigLIP2, InternVideo2, QwenVL2.5, InternVL3 and DINOv2, which tells you where it is intended to compete, not that it comes with an SLA.
When open_clip or timm is the better fit
If your task is zero-shot image classification with a CLIP-family model and nothing else, open_clip is the narrower dependency. It exists to load and run CLIP-style models, and its API is built around that single job. perception_models carries a wider dependency list because it also has to support video decoding, audio, tiling and the PLM decoder path. For a text-to-image or image-to-text pipeline with no video and no audio, that extra surface buys you nothing.
The counter-argument is real, though: the README states that PE core outperforms SigLIP2 on image and video benchmarks, and that PE was integrated into timm. If you are already using timm as your backbone library, loading PE through timm gives you the encoder without adopting this repository's dependency set at all. That is the middle path, and for many teams it is the right one.
The choice is clearer at the other end. If you need a multimodal LLM that answers questions about documents or video, PLM is the point of the repository, and open_clip has no equivalent. The README notes PLM was added to lmms-eval, so you can reproduce the reported VideoBench results through that harness rather than writing your own evaluation loop.
Questions people ask before adopting perception_models
The questions below cover the points that come up when a team first evaluates this repository: what the model families are, where the weights live, and how the licence applies. Answers stay within what the README, setup.py and requirements.txt state.
Editorial conclusion
Adopt it if you need a single codebase for zero-shot image and video classification, retrieval, dense prediction or audio-visual embedding, and you can live with the FAIR Noncommercial Research License that setup.py declares for the package. Do not adopt it if your product is commercial, or if you want a supported release cadence: the repository lists no releases, and the last push was on 2026-04-13. Before writing any integration code, read LICENSE.PE and LICENSE.PLM, confirm which of the four checkpoint families you actually need, and check that the pinned requirements.txt resolves against your existing torch and transformers versions rather than assuming it will.
Frequently asked questions
What are the four checkpoint types in perception_models?
The README lists PE core for vision-language tasks such as zero-shot image and video classification and video retrieval, PE lang for powering PLM, PE spatial for vision-centric tasks such as detection, depth estimation and tracking, and PE audio-visual for embeddings over audio, video, audio-video and text.
Where do I download the Perception Encoder checkpoints?
The README links each checkpoint to its own Hugging Face model page, for example PE-Core-T16-384 through PE-Core-G14-448 and PE-Lang-L14-448, and points to a Hugging Face collection for Perception Encoder.
What Python version does perception_models require?
setup.py sets python_requires to >=3.11, so the package will not install on older interpreters.
Is perception_models available under Apache-2.0?
The README shows an Apache-2.0 code licence badge and an Apache-2.0 model licence badge for PE, but setup.py declares the package licence as "FAIR Noncommercial Research License" and classifies it as Other/Proprietary. LICENSE.PE and LICENSE.PLM are the files to read for the weights.
Can I use Perception Encoder with timm?
The README states that Perception Encoder was integrated into timm in May 2025, so the encoder can be loaded through timm rather than through this repository.
Which PE lang checkpoint should I use with a tiling decoder?
The README's SotA table uses PE-Lang-L14-448-Tiling and PE-Lang-G14-448-Tiling, with a footnote saying these were aligned with tiling and should be used when you run higher than 448 resolution with tiling in the LLM decoder.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-perception-models)