BEHAVIOR-1K: a household benchmark where the hard part is the simulator
BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx
At a glance
- What is it?
- A thousand everyday activities drawn from human time-use studies, packaged as one repository containing the simulator, the task definitions, the robot drivers and the evaluation harness.
- Who is it for?
- BEHAVIOR-1K is best evaluated as infrastructure rather than as a leaderboard, because the value is in having a fixed, physically plausible version of a thousand tasks that anyone can run the same way. The repository is monolithic on purpose and the release history shows the maintenance burden that implies, with the Docker path and the robot abstraction both being actively reworked.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the benchmark actually measures
The framing is human-centered, and that word does real work in the design. The thousand activities were selected from human time-use surveys and preference studies rather than invented by robotics researchers, and they cover the ordinary categories of domestic life: cleaning, cooking, organizing. The stated goal is testing embodied AI agents on tasks that people actually do, at a scale where a single success rate would be meaningless.
A thousand tasks is a benchmark, not a single task. The interesting design question is how you aggregate across that many household situations, and the repository is largely the machinery for doing so consistently: the task definitions, the simulator, the robot drivers and the evaluation harness all live together so a result in one paper is comparable to a result in another.
The project describes itself as a platform for accelerating embodied AI research rather than as a dataset. The distinction matters. Publishing task definitions without a stable simulator would let anyone report numbers that cannot be reproduced, so the simulation environment is treated as part of the benchmark rather than as an implementation detail.
A monolithic repository with named subsystems
The tree is the best documentation the project provides, and it reads as a map of the whole stack. `OmniGibson/` is the simulator layer, the environment the agents act in. `bddl3/` holds the task definition language, which is how a household activity is specified as machine-readable state, objects and goals. `joylo/` is robot support, and the release history shows it becoming a single robot class driven by YAML definitions rather than a class per robot.
Around those sit the pieces that turn experiments into comparable numbers. `eval-jobqueue/` handles distributed evaluation, which is what you need when a thousand tasks times multiple seeds times multiple robots is a queue rather than a script. `asset_pipeline/` covers the 3D content, `datasets/` the released data, `knowledgebase/` the object and scene knowledge, and `docs/` the site built by the mkdocs configuration at the root.
There is also `docker/` for the container path, a Windows setup script alongside the shell one, and a lint configuration in ruff.toml with a pre-commit config to match. The presence of both AGENTS.md and CLAUDE.md at the root is a modern touch: the repository ships instructions aimed at coding agents, which tells you something about how contributions are expected to be made.
The README is a citation stub, the website is the manual
It is worth being blunt about this: the README is very short. It states what the benchmark is, shows a splash image, points at the main website, mentions that an installation script handles dependencies with modular installation, and then provides a BibTeX entry. There is no quick start command, no system requirements list, no example task, and no API surface in the repository's own documentation.
Everything operational lives on the documentation site, with the installation guide as the specific entry point for setup. For a project of this size that is a defensible choice, since a README that tried to summarize a thousand tasks and a simulator would be worse than a pointer. But it does mean you cannot evaluate the installation cost from the repository alone, and you should go to the site before deciding whether this fits your hardware.
The citation itself is worth noting for a different reason. The reference is an arXiv preprint from 2024 with a very long author list that includes Fei-Fei Li, and the paper title carries the full framing of a human-centered embodied AI benchmark with a thousand everyday activities and realistic simulation. If you write about embodied agents, citing this rather than the repository is the expected form of attribution.
Docker is being rebuilt from scratch
The most informative thing in the release history is what happened to Docker. Version 3.9.0, in July 2026, includes a change titled update Docker setup to create image from scratch, use setup script, which is the kind of fix that tells you the previous container image was built in a way that did not reproduce cleanly. Two related changes in the same release add conda activation to the GitHub Actions entrypoint and use the default environment inside the actions runner.
Later, version 3.9.2 adds a workflow to manually dispatch the Docker build from non-main branches. That is a small change with a clear purpose: it lets a contributor build and test an image from a feature branch without merging first, which is how you avoid the situation where CI only validates main.
Taken together, these read as a project that has been making its container setup reproducible, and that has automated the parts that were previously done by hand. If you are coming to BEHAVIOR-1K fresh, the current releases are a better starting point than older tutorials, and building the image from the current branch is safer than following an archived setup guide.
Robot support became configuration
A change in version 3.9.0 refactors robot subclasses into a single robot class plus YAML definition files. This is a structural change with real consequences for anyone extending the benchmark, and it went further than the title suggests: a companion change sets a fixed base to true for non-mobile robots, and another fixes a joint limit in a holonomic motion planner.
The reason to make the change is easy to grasp. A benchmark with one thousand tasks wants to be run on as many robots as possible, because the interesting question is whether a method generalizes across embodiments rather than exploiting the quirks of one arm. Subclass-per-robot makes that expensive, since every new robot means new code paths and new places for bugs. Configuration-per-robot moves the differences into data, where they are easier to share and to review.
Version 3.9.0 also added new robots outright and a code and report release column to the leaderboard. That last detail is a good sign for a benchmark: it means submissions are expected to publish their code, not just a number, and the leaderboard surface is built to make that visible.
The 3.9.1 and 3.9.2 releases are more domestic. Hugging Face download handling was updated, torch thread counts became configurable, observation wrapper handles are refreshed before loading, and the evaluation script was changed to use the rooms specified in task metadata. Each of those is the kind of fix that only surfaces once a lot of people run the thing.
Activity signals and the practical caveats
The repository is not archived and was last pushed on 2026-09-19, so this is active work. Open issues stand at 310, which is high in absolute terms and unremarkable for a research benchmark with a user community, a Discord and an annual challenge. GitHub reports MIT for the license, and a LICENSE file sits at the root, so the code can be used and modified with the notice retained. The project also runs an August 2026 challenge, and the release notes show a dedicated change for it, which is a reasonable cadence for a community benchmark.
Two practical cautions. First, simulation benchmarks have version coupling. Results depend on the simulator version, the task definition version and the asset set, so a number without those pinned is not reproducible. The B100 task naming in the evaluation fix from 3.9.2 is a small illustration of how quietly a task set can shift underneath a result.
Second, the honest scope. A benchmark of a thousand household activities is not a claim that household work is solved by any score on it. Success rates on realistic domestic tasks with many objects and long horizons are low across the board, and the value of the benchmark is in making progress measurable rather than in making the tasks easy. Anyone reading a leaderboard result should check which task subset was used and which simulator version produced it, because an aggregate over one thousand tasks can hide as much as it reveals.
The project site, the documentation and the installation guide are where those details live. The repository gives you the citation, the layout and the change history, and for a platform of this size that is about the right division of labor.
Editorial conclusion
BEHAVIOR-1K is best evaluated as infrastructure rather than as a leaderboard, because the value is in having a fixed, physically plausible version of a thousand tasks that anyone can run the same way. The repository is monolithic on purpose and the release history shows the maintenance burden that implies, with the Docker path and the robot abstraction both being actively reworked. Start from the installation guide on the project site rather than from the README, which is a citation stub, and expect the real configuration decisions to live there. If you are evaluating an embodied agent, the per-task definition language and the simulator version are the two things you must pin down before your numbers mean anything.
Frequently asked questions
What is behavior 1K?
It is a simulation benchmark for embodied AI agents covering one thousand everyday household activities, including cleaning, cooking and organizing. The activities were selected from human time-use surveys and preference studies rather than invented by researchers, and the repository ships the simulator, task definitions, robot support and evaluation tooling together.
What simulator does BEHAVIOR-1K run on?
The repository's `OmniGibson/` directory holds the simulation layer the agents act in. The README itself points to the documentation site for setup details rather than describing requirements, so the simulator configuration and its version are documented at behavior.stanford.edu.
How do I install BEHAVIOR-1K?
The project ships an installation script that handles dependencies and components and supports modular installation, so you can install only the parts you need. The repository has `setup.sh` and a PowerShell equivalent at its root. The README does not document the steps itself and directs you to the installation guide on the project website.
Can I add my own robot to the benchmark?
Support for that is the direction the project moved in. A 3.9.0 release refactored the robot subclasses into a single robot class driven by YAML definition files and added new robots, which lowers the cost of adding an embodiment since the differences now live in configuration rather than in new subclasses.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/stanfordvl-behavior-1k)