# Scenic: A JAX Research Codebase for Attention-Based Computer Vision Models

> Scenic is a JAX and Flax library from Google Research that provides shared training utilities and a collection of fully implemented vision model projects. It targets researchers running large-scale multi-device experiments on images, video, audio, and multimodal data.

**google-research/scenic** — Scenic: A Jax Library for Computer Vision Research and Beyond

- Repository: https://github.com/google-research/scenic
- Stars: 3,839 · Forks: 482
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-research-scenic

## What Scenic Provides and Who It Is For

Scenic is a research codebase, not a general-purpose library. The README describes it as two things: a set of shared lightweight libraries for common tasks in large-scale vision training, and a collection of projects that contain fully implemented training and evaluation loops for specific problems.

The shared libraries cover experiment launching, summary writing, logging, profiling, optimized training and evaluation loops, losses, metrics, bipartite matchers for detection, input pipelines for popular vision datasets, and a set of strong non-attentional baseline models.

The projects layer implements complete pipelines for classification, segmentation, detection, video understanding, audio modeling, and multimodal tasks. Each project is a self-contained directory with its own training loop, configuration, and evaluation logic.

The intended users are machine learning researchers who want to experiment with new attention-based architectures at multi-device or multi-host scale in JAX, without writing all the boilerplate for distributed training from scratch. Practitioners who need a stable inference API or a deployment-ready model should look elsewhere.

## The Fork-and-Copy Philosophy and What It Means for Code Organization

Scenic's stated philosophy is to prefer forking and copy-pasting over adding complexity or increasing abstraction. Functionality is promoted to shared libraries only when it proves useful across many models and tasks.

This means each project in Scenic is relatively self-contained. A researcher who wants to experiment with a model from the projects directory reads a bounded amount of code without needing to trace through multiple layers of abstraction. The downside is that common logic sometimes appears in slightly different forms across projects, and refactoring one project's implementation does not automatically improve others.

The top-level repository structure has two main directories: scenic/ and setup.py. Within scenic/, the shared library code sits alongside a projects/ subdirectory containing all the individual model projects. The setup.py docstring shows the development install path:

```bash
pip intall -e . .[testing]
```

Note that the setup.py docstring contains a typo: "intall" instead of "install". This is the exact text in the setup.py file. Running the corrected form `pip install -e .` is the standard Python editable install pattern; verify against the project's README Getting Started section for the current recommended command.

Scenic is developed in JAX and uses Flax for neural network definitions. Both are dependencies that must be compatible with your hardware and CUDA version before training will run correctly.

## Models and Projects Available in the Repository

Scenic includes reproductions of several established baselines alongside research models developed and published using Scenic itself.

Baseline reproductions include ViT (An Image is Worth 16x16 Words), DETR and Deformable DETR for object detection, CLIP, MLP-Mixer, BERT, U-Net for segmentation, PCT for point clouds, and Masked Autoencoders.

Research models developed using Scenic include ViViT (A Video Vision Transformer), OmniNet for omnidirectional representations, TokenLearner, Vid2Seq for dense video captioning, AVATAR for audiovisual speech recognition, Pixel Aligned Language Models, and MatFormer for elastic inference, among others. The full list in the README runs to over 30 projects.

For object detection specifically, CenterNet and the Segment Anything Model (SAM) are also reproduced. The segmentation side includes Location-Aware Self-Supervised Transformers and U-Net.

Not all projects are equally maintained. The research projects reflect the state of published work; some may not have been updated after publication. Baseline reproductions are generally more stable since they track well-understood architectures.

## Getting Started: Development Installation and Project Structure

Scenic requires JAX and Flax as its core dependencies. The setup.py lists these as install requirements along with ott-jax and, for projects that use SimCLR augmentations, an additional download step is triggered by a custom setup command.

The setup.py docstring shows the documented development install path:

```bash
pip intall -e . .[testing]
```

Note that the docstring contains a typo in "intall". The testing extra installs the packages needed to run the test suite. Because JAX's GPU support depends on matching CUDA and cuDNN versions, JAX itself is typically installed separately before this step, following JAX's own hardware-specific install guide.

Once installed, the recommended path is to navigate to the specific project directory you want to use under scenic/projects/ and read that project's README or configuration files. The top-level scenic/ directory houses both the shared library code and a projects/ subdirectory containing all the individual model implementations.

The setup.py and its install requirements are the canonical dependency sources. The Apache-2.0 license applies to the full repository.

## What Scenic Does Not Cover: Deployment, Production, and Non-JAX Environments

Scenic is not designed for production deployment. There are no model export utilities for ONNX, TensorFlow SavedModel, or TorchScript formats. Serving a Scenic-trained model in production requires writing the export and serving logic yourself.

The codebase is JAX-only. PyTorch users who want to adopt a specific model architecture from Scenic must re-implement it in PyTorch independently. This is a firm boundary: the shared libraries are JAX and Flax idioms that have no direct PyTorch equivalents.

PyTorch-based alternatives for large-scale vision research include Detectron2 from Facebook AI Research for detection and segmentation, and the Hugging Face Transformers library for transformer-based architectures. The key difference is that Detectron2 and Transformers are packaged libraries with stable APIs and model hubs, while Scenic is a research codebase where the API may change and models may require adaptation to run correctly.

Scenic also does not include data annotation tools or dataset management utilities beyond input pipeline code for a predefined set of standard datasets.

## Maintenance and License

The last push to the main branch was on 2026-09-27, indicating the project receives regular updates. There are no formal GitHub releases; versioning is tracked through commits and the changelog in the setup.py copyright header, which shows 2026 as the current year.

The license is Apache-2.0, which permits commercial use, modification, and redistribution with attribution. The CONTRIBUTING.md file in the repository root provides guidance for submitting changes.

Since the repository hosts both shared library code and multiple research projects, the scope of maintenance differs by area. The shared libraries are more likely to remain current since they underpin all projects. Individual project subdirectories may reflect the state of a specific publication and receive fewer updates afterward.

## Conclusion

Scenic is a reasonable starting point for researchers who want working JAX implementations of attention-based vision models and plan to fork and modify them. The philosophy of preferring fork-and-copy over abstraction means project code is readable but not designed as a reusable library. Engineers who need production-grade deployment or PyTorch-based pipelines will find Scenic unsuitable. Before adopting it, verify that your hardware runs JAX and Flax at the versions the project requires, and check whether the specific project subdirectory you want has been updated recently, since not all projects within the repository receive the same maintenance cadence.

## FAQ

### What is Scenic and what is it used for?

Scenic is a JAX and Flax research codebase from Google Research for developing attention-based computer vision models at multi-device scale. It includes shared training utilities and a collection of fully implemented project pipelines for classification, detection, segmentation, video, audio, and multimodal tasks.

### What models are included in the Scenic repository?

Scenic includes reproductions of ViT, DETR, Deformable DETR, CLIP, MLP-Mixer, BERT, U-Net, Masked Autoencoders, SAM, and CenterNet, alongside research models such as ViViT, TokenLearner, Vid2Seq, AVATAR, and MatFormer. The full list in the repository runs to over 30 projects.

### Can Scenic be used for production deployment or with PyTorch?

No. Scenic is a JAX-only research codebase with no export utilities for serving frameworks. It does not support PyTorch, and there are no model export paths for ONNX or TensorFlow SavedModel format.

## Sources

- [google-research/scenic on GitHub](https://github.com/google-research/scenic)
- [Issues](https://github.com/google-research/scenic/issues)
- [License: Apache-2.0](https://github.com/google-research/scenic/blob/main/LICENSE)
- [README](https://github.com/google-research/scenic/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-research-scenic
