NVIDIA Physical AI AV Dataset Devkit: Python Tools for a Large Multi-Sensor Driving Dataset
Devkit and documentation for the NVIDIA Physical AI Autonomous Vehicles Dataset
At a glance
- What is it?
- physical_ai_av is a Python developer kit for working with the NVIDIA Physical AI Autonomous Vehicles Dataset hosted on Hugging Face. It provides loading utilities and interactive notebooks for researchers building end-to-end autonomous driving systems from multi-sensor data.
- Who is it for?
- The physical_ai_av devkit is the supported entry point for the NVIDIA Physical AI Autonomous Vehicles Dataset. It is the right tool for AV researchers who need structured access to a large, geographically diverse multi-sensor dataset and are comfortable with the Hugging Face authentication flow and the dataset's license agreement.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 70 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Devkit Is and What Dataset It Accesses
The NVIDIA Physical AI Autonomous Vehicles Dataset is a collection of multi-sensor data for training end-to-end driving systems. The README describes it as one of the largest and most geographically diverse collections of this type. The data is hosted on Hugging Face at nvidia/PhysicalAI-Autonomous-Vehicles, not in this repository.
The physical_ai_av Python package provides the loading and processing utilities for that dataset. The repository also contains interactive Jupyter notebooks in the notebooks/ directory and a wiki for extended documentation. The package version is 0.2.2 according to pyproject.toml.
The target audience is AV researchers working on Physical AI, which in the context of this project means end-to-end neural driving systems that process raw sensor input rather than relying on modular perception and planning pipelines. The README references the Alpamayo NV Developer Forum for usage questions about the dataset and the devkit.
Installing the Devkit and Authenticating with Hugging Face
Installing the Python package itself is straightforward:
pip install physical_ai_avAccessing the dataset requires three additional steps. First, create a Hugging Face account if you do not have one. Second, visit the dataset card at huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles and agree to the NVIDIA Autonomous Vehicle Dataset License Agreement displayed at the top of the page. Third, create a Hugging Face User Access Token and configure authentication using one of the methods the dataset card describes.
The devkit uses huggingface-hub to retrieve data, so the authentication token must be available in the environment or configured via the Hugging Face CLI before loading any data. Without completing the license agreement step, the dataset download will be refused regardless of token validity.
The package requires Python 3.11 or above. The pyproject.toml specifies requires-python = ">=3.11".
Core Dependencies and Optional Demo Requirements
The base installation brings in av, huggingface-hub, numpy, pandas, pyarrow, scipy, and tqdm. These cover video and audio decoding (av), dataset access (huggingface-hub), numerical processing (numpy, scipy), structured data (pandas, pyarrow), and progress reporting (tqdm).
Running the interactive notebooks requires the demo dependency group defined in pyproject.toml. The demo group adds ipykernel, ipywidgets, matplotlib, mediapy, and opencv-python-headless. These cover notebook execution, visualization, and computer vision operations needed for the notebook examples.
A separate dev group adds pytest and pre-commit for contributors working on the devkit itself. The pyproject.toml uses uv_build as the build backend and specifies uv_build>=0.9.7,<0.10.0. The default groups setting installs all dependency groups by default, so running uv sync without specifying groups will include the demo and dev extras.
The Repository Layout and How to Use the Notebooks
The repository structure is conventional for a Python project. The src/ directory holds the package source. The notebooks/ directory contains interactive examples for working with the dataset. The tests/ directory holds the test suite.
The pyproject.toml includes a ruff configuration with a line length of 100 characters, indicating that ruff is used for linting and formatting. The .pre-commit-config.yaml confirms pre-commit hooks are configured for contributors.
The notebooks are the primary user-facing documentation alongside the wiki. The README directs usage questions to the Alpamayo NV Developer Forum and bug reports to GitHub Issues using provided templates. Security vulnerabilities must be reported through NVIDIA's Vulnerability Disclosure Program, not as public GitHub issues.
The package name is physical_ai_av and it is imported as such. The dataset itself is multi-sensor, but the devkit's specific sensor types and data formats are documented in the wiki rather than in the README.
Limitations: License Gate, Hugging Face Dependency, and Python Version
The dataset requires accepting a license agreement on Hugging Face before any data can be downloaded. This is a real friction point for automated or institutional access. Researchers who need to provision cloud compute jobs that pull the dataset at startup must handle the authentication and license acceptance as part of their setup, not just the pip install.
All data access goes through Hugging Face. If the Hugging Face Hub is unavailable or if NVIDIA removes or changes the dataset there, the devkit cannot retrieve data. There is no offline fallback documented in the README.
Python 3.11 is the minimum version. Teams still on Python 3.10 cannot install the package without upgrading.
The most widely used public autonomous driving datasets for comparison are nuScenes and Waymo Open Dataset. Both have their own Python devkits and are also gated behind license agreements. The difference here is that the NVIDIA dataset uses the Hugging Face platform for distribution and access control, while nuScenes and Waymo use their own registration portals. All three require human acceptance of terms before data access.
Reporting Issues and Security Disclosures
The README distinguishes between three channels for feedback. Usage questions and general discussion about the dataset go to the Alpamayo NV Developer Forum. Code-level bugs, documentation problems, and feature requests go to GitHub Issues using the templates provided: Bug report, Documentation request, or Feature request. Each template auto-assigns the relevant NVIDIA responder.
Security vulnerabilities must not be filed as public GitHub issues. The README directs reporters to NVIDIA's Vulnerability Disclosure Program at intigriti.com. This separation is standard for NVIDIA projects and means that embargo periods and coordinated disclosure are handled outside the public issue tracker.
The SECURITY.md file in the repository contains the formal security policy.
Maintenance and License
The last push to the repository was on 2026-07-22. The package is version 0.2.2. The devkit itself is under the MIT license, which permits use, modification, and redistribution. The dataset it accesses is under the NVIDIA Autonomous Vehicle Dataset License Agreement, which is separate and more restrictive than the MIT devkit license.
The author listed in pyproject.toml is NVIDIA, with a contact email of [email protected]. The project homepage points to the Hugging Face dataset card.
The uv_build build backend and the use of uv for dependency management reflect current NVIDIA Python tooling practices. The uv.lock file pins exact dependency versions for reproducible installs.
Editorial conclusion
The physical_ai_av devkit is the supported entry point for the NVIDIA Physical AI Autonomous Vehicles Dataset. It is the right tool for AV researchers who need structured access to a large, geographically diverse multi-sensor dataset and are comfortable with the Hugging Face authentication flow and the dataset's license agreement. Teams that need a dataset they can use without a Hugging Face account or a license agreement, or who work primarily with well-known public AV benchmarks, may find the access requirements a barrier. Verify that you can accept the dataset license at the Hugging Face dataset card before beginning the setup.
Frequently asked questions
Do I need a Hugging Face account to use the physical_ai_av devkit?
Yes. Accessing the dataset requires a Hugging Face account, accepting the NVIDIA Autonomous Vehicle Dataset License Agreement on the dataset card, and creating a User Access Token for authentication. The README lists all three steps.
What Python version does the physical_ai_av package require?
The pyproject.toml specifies requires-python = ">=3.11". Python 3.10 and below are not supported.
Where should I ask questions about using the NVIDIA Physical AI AV Dataset?
The README directs usage questions to the Alpamayo NV Developer Forum. Code bugs and documentation issues go to GitHub Issues using the provided templates. Security vulnerabilities go to NVIDIA's Vulnerability Disclosure Program and must not be filed as public GitHub issues.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-physical-ai-av)