TRIBE v2: a multimodal model that predicts what your brain does to a video
This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
At a glance
- What is it?
- Meta releases the training code and pretrained weights for a model that maps video, audio and text onto the cortical surface, with predictions on a 20,000 vertex mesh.
- Who is it for?
- TRIBE v2 is one of the more usable brain encoding models, because the inference path is three lines and the weights are on HuggingFace rather than behind a request form. What it does not give you is an individual brain.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 105 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 23, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Inference is three lines and a HuggingFace download
The quick start block is the whole public interface, and it is worth reading closely because several important details live in it.
from tribev2 import TribeModel
model = TribeModel.from_pretrained("facebook/tribev2", cache_folder="./cache")
df = model.get_events_dataframe(video_path="path/to/video.mp4")
preds, segments = model.predict(events=df)
print(preds.shape) # (n_timesteps, n_vertices)`get_events_dataframe` takes a video path and returns something called an event dataframe. The word events is doing specific work: the model consumes a stream of annotated stimulus intervals rather than a flat array of frames, and `predict` returns both predictions and segments, so temporal boundaries come back alongside the signal.
The print statement is more informative than it looks. Output shape is timesteps by vertices, meaning this is not a per-image classifier output but a dense cortical time course. The README states predictions are for the average subject, following the paper, and live on the fsaverage5 cortical mesh at roughly 20,000 vertices.
Two further details follow in prose. Predictions are offset by five seconds in the past, to compensate for the hemodynamic lag, which is the kind of thing that quietly invalidates a comparison if you do not know about it. And you can pass `text_path` or `audio_path` to `get_events_dataframe` instead of a video, in which case text is automatically converted to speech and transcribed to obtain word-level timings.
Three extras split inference, plotting and training
Installation is an editable install with optional extras, which is a clear sign of where the authors expect most readers to land.
pip install -e .That base install is labelled inference only. Brain visualization comes as a second extra, and training dependencies as a third:
pip install -e ".[plotting]"pip install -e ".[training]"Reading `pyproject.toml` tells you what each extra actually contains. The `plotting` group is nibabel, matplotlib, seaborn, colorcet, nilearn, scipy, pyvista and scikit-image, which is two visualization stacks: nilearn for statistical brain plotting and PyVista for mesh rendering. The `training` group is nibabel, torchmetrics, wandb and lightning, which is PyTorch Lightning plus experiment tracking. A third extra, `test`, holds only pytest.
The base dependency list is where the design shows. `neuralset==0.0.2` and `neuraltrain==0.0.2` are pinned Meta packages that sit at version 0.0.2, which tells you these are new code paths rather than established interfaces. Then `torch>=2.5.1,<2.7`, `numpy==2.2.6` exactly, `torchvision>=0.20,<0.22`, and `x_transformers==1.27.20`, which is the attention library the Transformer encoder is built on. `exca==0.5.20` handles the experiment config. Python 3.11 is the floor.
Training expects multi-study data and a cluster
The training instructions are written for people who already have fMRI studies on disk. Step one is two environment variables:
export DATAPATH="/path/to/studies"
export SAVEPATH="/path/to/output"The README adds that you configure the Slurm partition in the same place, or edit `tribev2/grids/defaults.py` directly. That single sentence reveals the intended environment: Meta's internal HPC cluster.
Step two has a local path and a cluster path. For a quick check:
python -m tribev2.grids.test_runFor the real thing there are two grid entry points, one cortical and one subcortical:
python -m tribev2.grids.run_cortical
python -m tribev2.grids.run_subcorticalThe module paths are the documentation here. `tribev2.grids` holds both `defaults.py`, described as the full default experiment configuration, and `test_run.py` as the quick local entry point. The fact that cortical and subcortical are separate runs rather than one configurable run is a design choice worth noticing, since the two involve different vertex counts and different surface projection steps.
`utils.py` is described as multi-study loading, splitting and subject weighting, which confirms the training path is built around combining several datasets and balancing subjects across them rather than fitting one study at a time.
What each module in the package is for
The README includes a project structure block that is more informative than the prose above it. `main.py` is the experiment pipeline, described as Data and TribeExperiment. `model.py` holds FmriEncoder, the Transformer-based multimodal to fMRI model. `pl_module.py` is the PyTorch Lightning training module. `demo_utils.py` holds TribeModel and the helpers for inference from text, audio or video, which tells you the public class lives in a file named for demo purposes.
`eventstransforms.py` is described as custom event transforms, with word extraction and chunking named as examples. That connects back to the transcribe-then-time behaviour in the quick start: the text path produces word-level timings through these transforms.
`utils_fmri.py` handles surface projection from MNI to fsaverage plus ROI analysis, which is the bridge between the fMRI volume data people have and the surface mesh the model predicts on.
The `plotting/` directory is described as brain visualization with PyVista and Nilearn backends, and `studies/` as dataset definitions with Algonauts2025 and Lahner2024 named. Those two study names are the concrete answer to what data the code expects, and they are the first place to look if you want to train on something the authors have already wired up.
Where the README points you away from itself
There is a paper, a demo and a Colab, and the README treats each as the place where a different question gets answered. The paper is arXiv 2605.04326, titled a foundation model of vision, audition, and language for in-silico neuroscience, authored by d'Ascoli, Rapin, Benchetrit, Brooks, Begany, Raugel, Banville and King. The demo is hosted at aidemos.atmeta.com. Weights are on HuggingFace at facebook/tribev2.
The Colab notebook is called `tribe_demo.ipynb` and sits in the repository root, and the README describes it as a full walkthrough with brain visualizations. That is the fastest way to see what the output looks like in anatomical terms, since the base install gives you an array of numbers on a mesh and the notebook shows what those numbers mean when placed on a brain surface.
The average-subject detail, which the README calls out twice, is the thing the paper is needed for. Nothing in the quick start suggests per-subject prediction is offered, and the phrase points to details in the paper rather than a configuration flag. Anyone expecting individual brain decoding from these weights is going to be disappointed, and the README does at least warn them.
Two supporting files exist without being described: `CONTRIBUTING.md` and `CODE_OF_CONDUCT.md`, and the README links the former.
Licence, versions and the shape of the release
The licence is CC-BY-NC-4.0, stated both in the badge and in a dedicated section, and the tree contains a `LICENSE` file. `pyproject.toml` reads it as `license = {file = "LICENSE"}`. This is a non-commercial licence, which is normal for research output and means the weights and code are not available for commercial deployment without a separate agreement. The README also has an open science section asking users to share results through a citation, naming an arXiv journal reference.
Version 0.1.0 in `pyproject.toml` and no tagged releases at all. That is consistent with how this kind of research code usually ships: one code drop accompanying a paper, with the version number set for the initial state rather than tracking a release train.
The repository is small and notebook-first. The tree holds `.gitignore`, `CODE_OF_CONDUCT.md`, `CONTRIBUTING.md`, `LICENSE`, `README.md`, `pyproject.toml`, `tribe_demo.ipynb` and the `tribev2/` package. The declared language of the repository is Jupyter Notebook, which reflects that the primary artefact a reader is expected to open is the demo notebook rather than a test suite.
There are 3,231 stars, 697 forks and 48 open issues, and the last push was on 2026-06-23. The repository is not archived.
Editorial conclusion
TRIBE v2 is one of the more usable brain encoding models, because the inference path is three lines and the weights are on HuggingFace rather than behind a request form. What it does not give you is an individual brain. The published weights predict a response for the average subject on the fsaverage5 surface, which is the right granularity for comparing stimuli and the wrong granularity for clinical or personal questions. The base install is enough to run inference, the `plotting` extra adds brain visualization, and the `training` extra is where you go if you have your own fMRI studies to fit. The CC-BY-NC-4.0 licence means academic and research use, so check that before anything else.
Frequently asked questions
What is tribev2?
TRIBE v2 is Meta's multimodal brain encoding model. It combines text, audio and video models in a single Transformer architecture and predicts fMRI responses to naturalistic stimuli, mapping representations onto the cortical surface on the fsaverage5 mesh of about 20,000 vertices.
How do I run TRIBE v2 on my own video?
Load the weights with `TribeModel.from_pretrained("facebook/tribev2")`, call `model.get_events_dataframe(video_path=...)` to turn the video into an event dataframe, then `model.predict(events=df)`. That returns predictions shaped timesteps by vertices along with the segments. You can pass `text_path` or `audio_path` in place of a video instead.
Does TRIBE v2 predict individual subjects or just an average brain?
The released weights predict the average subject, and the README points to the paper for details rather than offering a per-subject option. The output lands on the fsaverage5 surface at roughly 20,000 vertices, which makes it a group-level model in its current form.
Can I use TRIBE v2 commercially?
Not under the terms that ship with it. The project is licensed CC-BY-NC-4.0, which permits non-commercial use with attribution. Commercial use of the code or the weights would need a separate agreement with Meta.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-tribev2)