Model or dataset
google-research-datasets/Objectron avatar
google-research-datasets/Objectron

Objectron: object-centric AR video with 3D boxes and camera poses

Objectron is a dataset of short, object-centric video clips. In addition, the videos also contain AR session metadata including camera poses, sparse point-clouds and planes. In each video, the camera moves around and above the object and captures it from different views. Each object is annotated with a 3D bounding box. The 3D bounding box describes the object’s position, orientation, and dimensions. The dataset contains about 15K annotated video clips and 4M annotated images in the following categories: bikes, books, bottles, cameras, cereal boxes, chairs, cups, laptops, and shoes

2,352 stars268 forksJupyter NotebookNOASSERTION

At a glance

What is it?
Fifteen thousand short videos of household objects filmed from many angles, each frame carrying a 3D bounding box, a camera pose, a point cloud and surface planes, plus the notebooks that turn it into TensorFlow records.
Who is it for?
Objectron is worth your attention if your problem is 3D pose of a known object rather than 2D detection of unknown objects, because the annotation is a full oriented 3D bounding box and the AR metadata gives you the camera trajectory that a single image cannot.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the recordings contain and what is annotated

Objectron is built around a specific capture setup: someone picks up a phone, points it at one object, and walks around that object while the camera moves through different angles. What comes out is a short video clip per object plus AR session metadata recorded alongside it, which includes camera poses, a sparse point cloud, and a characterization of the planar surfaces in the surrounding environment. The headline figures are 15,000 annotated videos and more than 4 million annotated images.

Every frame carries a manually annotated 3D bounding box describing the object's position, orientation, and dimensions. That is the part that distinguishes this dataset from most video sets. A 2D box tells you where the object is in the image; an oriented 3D box tells you where it is in the room and how it is turned, which is the input a pose-aware model needs.

The nine categories are bikes, books, bottles, cameras, cereal boxes, chairs, cups, laptops, and shoes. The per-class table in the README lists 476 videos and 150k frames for bikes at the small end, and 2,204 videos with 546k frames for cups at the large end. Adding up the video column gives 14,588 and the frame column gives roughly 3.9 million, so the 15K and 4M figures in the summary are rounded up rather than exact. If a per-class count matters to your split, read it off the table instead of the headline.

There is one naming wrinkle. The dataset categories include `cups`, but the reference 3D object detection models that Google shipped for this data cover shoes, chairs, mugs, and cameras. Nobody converted the cup class into a mug class, so a reader looking for pretrained mug weights is looking at a MediaPipe model rather than at anything in this repository.

The data lives in a Cloud Storage bucket, not in git

Nothing here is committed to the repository. The repository holds schemas, notebooks, and an index. The payload sits in a public Google Cloud Storage bucket, and the README points you at `notebooks/Download Data.ipynb` for the mechanics. The sizes are worth internalizing before you plan anything: raw video plus annotations come to 1.9TB, and the full dataset including the prepared records and sequences reaches 4.4TB.

Three tiers of data are available, and picking the wrong one is the expensive mistake.

The first tier is the raw capture. Video sequences sit at predictable paths, and the annotation protobufs for each clip live beside them:

text
/videos/class/batch-i/j/video.MOV
/videos/class/batch-i/j/geometry.pbdata
/v1/records_shuffled/class/
/v1/sequences/class/

The second and third of those paths are the prepared tier, which is where most people should start. `tf.records` of annotated frames in `tf.example` format live under `/v1/records_shuffled/class/`, and videos packaged as `tf.SequenceExample` live under `/v1/sequences/class/`. Those are sharded and shuffled, and they exist specifically so you can feed a data pipeline without writing the parsing layer yourself. They work in TensorFlow and PyTorch.

If you do need the raw tier, the annotation format is a protobuf defined in `objectron/schema/object.proto`, and the AR metadata follows `objectron/schema/a_r_capture_metadata.proto`. The README does not explain either schema inline; it links to `notebooks/Parse Annotations.ipynb` and `notebooks/objectron-geometry-tutorial.ipynb` for worked parsing examples. That is where the actual field-by-field documentation lives, so treat the schema files as the source of truth and the notebooks as the tutorial.

Notebooks for TensorFlow, PyTorch, NeRF, and evaluation

The `notebooks/` directory is the real user interface of this project, and it covers the ground a reader would expect plus a few things they would not. There is a download notebook, a Hello World example for loading examples in TensorFlow, a PyTorch loading tutorial, a parsing notebook for raw annotations, a geometry notebook for the AR metadata, and a SequenceExample tutorial.

Two of them are more interesting than they sound. `notebooks/3D_IOU.ipynb` covers evaluation with the 3D IoU metric, and the README lists supporting scripts for running that evaluation alongside the data. Objectron is one of the few public datasets where the evaluation tooling ships with the data rather than being reverse-engineered from a paper appendix, and 3D IoU for oriented bounding boxes is the metric that actually tells you whether a pose regression improved.

`notebooks/Objectron_NeRF_Tutorial.ipynb` trains a NeRF on the data. That connects back to the camera pose refinement in the 0.1.0 release: the team ran offline global bundle adjustment over each video to improve the camera poses, and the release notes say explicitly that the updated poses are useful for applications needing accurate and consistent camera trajectories, naming NeRF as the example. If you are doing novel view synthesis, that refinement is the reason to use this dataset rather than raw phone video.

There is also a set of supporting scripts described in the README for loading the data into TensorFlow, JAX, and PyTorch and visualizing it, including Hello World examples, plus Apache Beam jobs for processing the dataset on Google Cloud infrastructure. Beam support matters more than it first appears, since at this scale a serial download loop is not a plan.

The index directory and the train/test split

The `index/` directory at the repository root holds one annotation index per category, named after the class, such as `index/bike_annotations` and `index/shoe_annotations`. Alongside them the README describes an index of all available samples plus train and test splits, described as being there for easy access and download.

That last point deserves emphasis. At 4.4TB, the ability to enumerate what exists before fetching any of it is what makes the dataset tractable. Reading the split files first tells you how many clips your chosen category actually contributes to training versus evaluation, and reading the index tells you whether a sample is complete. A dataset distributed as raw video in a bucket with no manifest would be a much harder project to plan against.

The sample counts vary enough between classes that this matters in practice. Books give you 2,024 videos and chairs 1,943, while bikes give 476 and cameras 815. A method that needs many distinct object instances will struggle on the smaller categories, where 476 videos might mean only a few dozen physical objects filmed repeatedly from different angles. Nothing in the README breaks down instance count, which is the number you would want. Treat it as an open question rather than assuming the video count equals the object count.

Two releases in 2021 and a repository that has not moved much since

This project has exactly two tagged releases and both are from 2021. Version 0.1.0, published 2021-04-07, is where the substantive work happened. Its notes cover the CVPR 2021 paper, the arrival of Python and Web APIs for the Objectron models as part of the prebuilt MediaPipe package, and the camera pose refinement via offline bundle adjustment. It also notes that the MediaPipe Python package was published to PyPI for Linux, macOS, and Windows, and that Objectron became available through the ActiveLoop hub.

Version 0.1.1, published 2021-07-07, is smaller. It made the full set of models available for download at the Objectron bucket, covering EfficientNet and MobilePose backbones, for predicting 3D object poses from RGB images. Example usage for Mobile, Python, and Web lives in MediaPipe, not in this repository.

The repository itself shows 2,352 stars, 268 forks, and 31 open issues, with `master` as the default branch. The last recorded push is dated 2026-03-06, which is recent enough to rule out abandonment but a long way from an active cadence for a dataset whose model and API story lives in a separate framework. Read that as what it is: a stable, mostly frozen artifact that has already done its work, paired with a MediaPipe that does move. Anyone planning a long-lived pipeline should plan around the data being effectively fixed and the surrounding tooling being what evolves.

One documentation note. The README in this material ends partway through the BibTeX citation block, so if you need the full citation you will want the paper on arXiv, identifier 2012.09988, which the README does link.

Licensing under C-UDA and how this compares to the alternatives

Objectron ships under the Computational Use of Data Agreement 1.0, C-UDA-1.0, with a copy of the license in the repository's `LICENSE` file. This is worth reading properly rather than skimming, because it is not a permissive license in the way Apache or CC-BY would be. The C-UDA is Microsoft's attempt at a middle ground: it grants broad rights to use, modify, and redistribute the data while restricting uses that it judges harmful, such as surveillance applications and biometric identification. A permissive license such as the one that covers the reCAPTCHA PHP client would not carry that kind of restriction. If your intended application touches any of the excluded categories, the license is the first thing to check, not an afterthought.

Against the usual alternatives, Objectron's distinguishing feature is that the objects are known and named. COCO and Open Images give you detection and segmentation labels at scale, but nothing about how a mug is oriented in three dimensions. Datasets built for novel view synthesis tend to be small single scenes captured under controlled conditions. Objectron sits between them: thousands of real household objects in real rooms, each with a full pose annotation and a camera trajectory.

What it does not offer is scale in the way a pretraining corpus does, and it does not offer instance variety, since each category is a narrow set of physical objects filmed repeatedly. What it does offer, and what is genuinely hard to reproduce, is a consistent oriented 3D box per frame plus synchronized AR metadata, packaged with parsing notebooks and an evaluation script so that you can start from something that already works.

Editorial conclusion

Objectron is worth your attention if your problem is 3D pose of a known object rather than 2D detection of unknown objects, because the annotation is a full oriented 3D bounding box and the AR metadata gives you the camera trajectory that a single image cannot. The practical friction is not the schema, which is clean and documented through notebooks, but the volume: 4.4TB total, a Google Cloud Storage bucket rather than a release artifact, and no selective download beyond the per-category index. Open `notebooks/Download Data.ipynb` first and look at what you can afford, and read the per-class index under `index/` before assuming a class has enough video for your training split.

Frequently asked questions

How do I download the Objectron dataset?

The data is hosted in a public Google Cloud Storage bucket rather than committed to the repository. The README points to the Download Data notebook in the `notebooks/` directory for the access mechanics.

How large is the Objectron dataset?

Raw video plus annotations total 1.9TB, and the complete dataset with prepared records and sequences comes to 4.4TB. The summary figures are roughly 15,000 annotated videos and more than 4 million annotated images.

What annotation format does Objectron use?

Each frame carries a manually annotated oriented 3D bounding box, stored as protobuf files beside the videos and defined in `objectron/schema/object.proto`. AR metadata such as camera poses and point clouds follow a separate schema, `a_r_capture_metadata.proto`.

Which object categories are in the Objectron dataset?

Nine categories: bikes, books, bottles, cameras, cereal boxes, chairs, cups, laptops, and shoes. Sample counts range from 476 videos for bikes to 2,204 for cups.

Can I use Objectron for commercial or surveillance purposes?

Objectron is released under the Computational Use of Data Agreement 1.0, which grants broad usage rights while restricting uses the license considers harmful, including surveillance and biometric identification. The license text is in the repository's LICENSE file.

Official sources

  1. google-research-datasets/Objectron on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-research-datasets-objectron.svg)](https://hysenlabs.com/projects/google-research-datasets-objectron)