# TransFuser: camera and lidar fusion trained by imitation in CARLA

> TransFuser is the reference implementation behind the PAMI 2023 paper on transformer-based sensor fusion for end-to-end driving, a journal extension of the CVPR 2021 Multi-Modal Fusion Transformer. Getting it running means a branch named 2022, CARLA 0.9.10.1, and wheels pinned to CUDA 11.3.

**autonomousvision/transfuser** — [PAMI'23] TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving; [CVPR'21] Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

- Repository: https://github.com/autonomousvision/transfuser
- Stars: 1,607 · Forks: 245
- Language: Python
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/autonomousvision-transfuser

## What problem this solves: one policy from raw camera and lidar

TransFuser is code for a specific paper, and knowing which one matters. It accompanies the PAMI 2023 journal extension TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving, which is itself an extension of the CVPR 2021 paper Multi-Modal Fusion Transformer for End-to-End Autonomous Driving. The earlier conference code was not overwritten; it sits on a `cvpr2021` branch, while the default branch of this repository is named `2022`.

The problem it addresses is modality fusion for end-to-end driving. An end-to-end policy is trained to map sensor input straight to driving output, which means the network has to decide for itself how to combine what the cameras and what the lidar see. TransFuser's answer is a transformer over the fused representation, trained by imitation learning rather than by a hand-written cost function.

The audience is therefore narrow and research-shaped: people reproducing a published result, or extending that architecture. There is no inference package here, no vehicle interface, and no runtime for a real car. The published venue is also the whole of the project's claim; the repository carries no release history at all, so there is no version series to track.

## The 210GB dataset, the autopilot, and eight CARLA towns

The training data is generated, not collected. A privileged agent, the autopilot, lives at `/team_code_autopilot/autopilot.py` and drives through 8 CARLA towns using the routes and scenario files that ship in the leaderboard's training folder. The detailed documentation for those routes and scenarios sits in `tools/dataset`.

If you only want the published data rather than your own, the download is two commands:

```bash
chmod +x download_data.sh
./download_data.sh
```

That fetch is 210GB, which is the first number to plan around. The layout below each scenario is organised by town and route, and each route directory carries the camera images, matching depth images, matching segmentation images, the lidar point cloud in `.npy` format, topdown segmentation maps, 3D bounding boxes for vehicles in `label_raw`, and a `measurements` file holding the ego agent's position, velocity and other metadata. Depth, semantics and the bounding boxes are what make imitation learning on this data possible at all, since the policy's targets come from the privileged agent rather than from a human driver.

Regenerating the data yourself needs a running simulator first:

```bash
./CarlaUE4.sh --world-port=2000 -opengl
```

Once the server is up, the generation script takes the CARLA root and the path of this checkout:

```bash
./leaderboard/scripts/datagen.sh <carla root> <working directory of this repo (*/transfuser/)>
```

The two variables you set in that script are `SCENARIOS` and `ROUTES`. If your machine has no display, the README points you at CARLA's own headless documentation rather than guessing at flags.

## Setup is pinned to a 2021-era wheel index

Installation is a clone, a simulator, and a conda environment:

```bash
git clone https://github.com/autonomousvision/transfuser.git
cd transfuser
git checkout 2022
chmod +x setup_carla.sh
./setup_carla.sh
conda env create -f environment.yml
conda activate tfuse
pip install --only-binary=torch-scatter torch-scatter -f https://data.pyg.org/whl/torch-1.12.0+cu113.html
pip install --only-binary=mmcv-full mmcv-full==1.6.0 -f https://download.openmmlab.com/mmcv/dist/cu113/torch1.12.0/index.html
pip install mmsegmentation==0.25.0
pip install mmdet==2.25.0
```

Read the pins rather than skim them. The scatter library comes from a wheel index built for torch 1.12.0 and CUDA 11.3, and mmcv-full 1.6.0 comes from an OpenMMLab index for the same cu113 and torch1.12.0 combination. mmsegmentation 0.25.0 and mmdet 2.25.0 are pinned by version. The simulator is CARLA 0.9.10.1, and `setup_carla.sh` is expected to handle it once it is executable.

So there is no floating dependency set here. Four of the five Python packages arrive from fixed indexes built for one CUDA generation, which means a machine with a newer driver and a newer PyTorch build is not the target and will not be the easy path. The environment name is `tfuse`. The README does not document a Docker image, a Windows path, or a way to point these pins at a different CUDA version, and the top-level file list contains no continuous integration configuration to tell you which combinations are tested.

## One GPU first, torchrun for many

Training code is `train.py` in `team_code_transfuser`, and the README is clear that the main function's own documentation holds more options than the quick start shows. The minimal single-machine invocation is:

```bash
cd team_code_transfuser
python train.py --batch_size 10 --logdir /path/to/logdir --root_dir /path/to/dataset_root/ --parallel_training 0
```

Read that as a smoke test rather than a recipe for a result: a batch size of 10 and a serial run are what you use to find out whether the data loader and the model agree. `parallel_training 0` is the switch that separates the single-process path from the distributed one.

On a multi-GPU node the invocation changes shape, and four environment variables move with it:

```bash
cd team_code_transfuser
CUDA_VISIBLE_DEVICES=0,1 OMP_NUM_THREADS=16 OPENBLAS_NUM_THREADS=1 torchrun --nnodes=1 --nproc_per_node=2 --max_restarts=0 --rdzv_id=1234576890 --rdzv_backend=c10d train.py --logdir /path/to/logdir --root_dir /path/to/dataset_root/ --parallel_training 1
```

`CUDA_VISIBLE_DEVICES` enumerates the GPUs you want. `OMP_NUM_THREADS` should match the CPU count on the node, and `OPENBLAS_NUM_THREADS=1` is there to stop threads spawning threads. `--nproc_per_node` takes the number of available GPUs, `--max_restarts=0` disables restarts, and the rendezvous pair `--rdzv_id` with `--rdzv_backend=c10d` is what lets several processes find each other on one host.

One consequence is worth flagging before you plan an evaluation: the evaluation agent is built to evaluate models trained with multiple GPUs, and the README says a single-GPU-trained model needs one line removed from `submission_agent.py`. That is a source edit, not a flag.

## Longest6 and a modified leaderboard

Evaluation runs on the Longest6 benchmark, and the repository does not use the stock CARLA leaderboard code for it. The README states that minor modifications were made for that benchmark, and the details are not on this page. In practice that means the evaluation harness in `leaderboard/` is a patched copy rather than something you can point at an upstream tag.

The directory layout reflects how much of the pipeline ships in the repository. `scenario_runner/` and `leaderboard/` hold the simulation and scoring harness, `team_code_autopilot/` holds the privileged agent that produced the dataset, `team_code_transfuser/` holds training and the submission agent, `tools/` holds the dataset tooling, and `results/` holds outputs. For a research codebase that is unusually complete, and it is the reason the 210GB download and the simulator are part of the setup rather than something you supply.

What is absent is as informative. There are no unit tests to run, no container definition, and no scheduled job that proves the pinned dependency set still resolves. The README also does not give expected runtime, expected memory, or the training schedule length, so you are left estimating both from the batch-size example and the dataset size.

## PlanT, NEAT and carla_garage take different paths

The same group lists four other repositories, and reading them as alternatives says more about TransFuser's design than the architecture section does. PlanT, from CoRL 2022, is explainable planning transformers built on object-level representations, so it reasons over detected objects rather than over fused sensor features. NEAT, from ICCV 2021, uses neural attention fields, which renders what the network attends to instead of fusing features into a driving output. Different representations, different explanations, same problem.

KING, from ECCV 2022, generates safety-critical driving scenarios, and carla_garage, from ICCV 2023, looks at hidden biases in end-to-end driving models. Those two are not competing models but the other half of the work: they ask how to break and how to judge a policy like TransFuser rather than how to build one. If your interest is failure modes rather than fusion, carla_garage is the closer fit.

Against a conventional stack, the difference is where the perception lives. A modular pipeline puts camera and lidar handling in separate components with a hand-written interface between them, and TransFuser puts both modalities through one learned fusion. What you give up for that is inspectability: the repository offers no attention visualisation task and no evaluation that isolates the contribution of one sensor, so the fused model's behaviour is harder to attribute than a pipeline's.

## MIT licence, no releases, and a last push on 2026-09-21

The licence file is MIT, with no copyleft obligation on your model weights, your fork, or anything you train on top of it.

Maintenance is where the friction lives. The last push was on 2026-09-21 and the repository is not archived, yet it has no GitHub releases, so there is no tag to pin and nothing to install from a package index. The only version lock is `environment.yml` plus those four explicit pip commands, and the four `git checkout 2022` in the setup instructions is a reminder that the default branch is a snapshot name rather than `main`.

That combination, an active repository with no release discipline, is normal for academic code and it shapes your upgrade cost in one sentence: a CUDA or PyTorch upgrade is yours, because nothing in the repository will tell you a new combination works. Budget for reading `environment.yml` yourself, and for the disk and time that regenerating data with `SCENARIOS` and `ROUTES` of your own design would cost before you trust the published split.

## Conclusion

Use TransFuser if you want to reproduce the PAMI 2023 sensor-fusion results or extend that architecture, and you are willing to keep a CARLA 0.9.10.1 server and a 210GB dataset on disk. Do not reach for it as a driving stack: there is no release artefact, nothing in the tree describes continuous integration, and the dependency set is pinned to CUDA 11.3 wheels. Read environment.yml and the SCENARIOS and ROUTES variables in leaderboard/scripts/datagen.sh before you budget the disk and the GPUs.

## FAQ

### What is TransFuser used for in autonomous driving research?

TransFuser is the code for the PAMI 2023 paper TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving, a journal extension of the CVPR 2021 Multi-Modal Fusion Transformer. It trains a single end-to-end policy that fuses camera and lidar input by imitation learning, and the earlier conference code lives on the cvpr2021 branch.

### How do I install TransFuser and its dependencies?

Clone the repository, check out the 2022 branch, run setup_carla.sh after making it executable, then create and activate the conda environment named tfuse from environment.yml. Four pinned pip installs follow: torch-scatter and mmcv-full 1.6.0 from cu113 and torch1.12.0 wheel indexes, plus mmsegmentation 0.25.0 and mmdet 2.25.0.

### How large is the TransFuser training dataset and where does it come from?

The download is 210GB and comes from running download_data.sh. The data itself is generated by the privileged autopilot agent at team_code_autopilot/autopilot.py through 8 CARLA towns using the supplied route and scenario files, so you can regenerate it with leaderboard/scripts/datagen.sh by setting SCENARIOS and ROUTES.

### How do I train TransFuser on multiple GPUs?

Use torchrun with --parallel_training 1, setting CUDA_VISIBLE_DEVICES to the GPUs you want and --nproc_per_node to their count, OMP_NUM_THREADS to the CPU count and OPENBLAS_NUM_THREADS=1. The README notes that the evaluation agent targets multi-GPU trained models, so a single-GPU run needs one line removed from submission_agent.py.

### Can I evaluate a TransFuser model on a benchmark other than the provided one?

The evaluation described is the Longest6 benchmark, which uses minor modifications to the CARLA leaderboard code rather than the stock harness. Nothing in the repository describes an unmodified leaderboard path, so the patched copy under leaderboard/ is what you would adapt.

### Is TransFuser available as a pip package or release?

No. The repository has no GitHub releases, the setup path is a clone plus a conda environment, and the only version lock is environment.yml with four pinned pip installs against CUDA 11.3 wheel indexes. Nothing describes a package index entry or a Docker image.

## Sources

- [autonomousvision/transfuser on GitHub](https://github.com/autonomousvision/transfuser)
- [Issues](https://github.com/autonomousvision/transfuser/issues)
- [License: MIT](https://github.com/autonomousvision/transfuser/blob/2022/LICENSE)
- [README](https://github.com/autonomousvision/transfuser/blob/2022/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/autonomousvision-transfuser
