VirtualHome writes a program and renders it in Unity
API to run VirtualHome, a Multi-Agent Household Simulator
At a glance
- What is it?
- VirtualHome is a household activity simulator whose entire abstraction is a pair of objects: a program, which is a sequence of actions, and a graph, which is a definition of the environment. Give it both and one of two simulators runs them, either a Unity build that produces video and needs a platform-specific executable downloaded by hand, or a pure-Python graph simulator that needs nothing and still does not support every action the Unity side does.
- Who is it for?
- VirtualHome fits embodied AI research that needs scripted household activities with ground truth rather than photographs of them, and that wants the scene described as data. It does not fit a pipeline that needs photorealism or unrestricted object interaction, since both are on the project's own in-development list.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 139 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Programs, graphs, and nothing else to learn
The whole data model is two components, and the documentation does not bury them.
A program is the sequence of actions that composes an activity. A graph is a definition of the environment where the activity takes place. Given both, the simulator executes the program and produces either a video of the activity or a sequence of graphs showing how the environment evolves as it happens.
That is the entire abstraction, and it is why the platform can be called from Python with a handful of instructions. An activity such as picking up an object, switching an appliance on, or opening it is written as a sequence of steps rather than driven through a real-time control loop, which is what makes the output reproducible.
The point of the project is stated in terms of training. It can simulate multi-agent activities and serve as an environment for training agents on embodied tasks, and it can stream ground truth alongside the render: time-stamped actions, instance and semantic segmentation, optical flow and depth. That last capability is the reason a synthetic source is worth tolerating over a captured one.
You can also choose between different agents and environments, and modify environments on the fly rather than only loading them as authored.
Two simulators with a documented coverage gap
There are two simulators, and choosing between them is the first decision a user makes.
The Unity Simulator is built in Unity and generates videos of activities. Using it means downloading the appropriate executable for your platform and running it through the Python API, with a notebook demonstrating the loop.
The Evolving Graph runs entirely in Python and generates a sequence of graphs when a program executes. No executable, no graphics, no download.
The documentation attaches one important warning to the second option: some of the objects and actions it supports are not yet supported in the Unity Simulator. So the two are not interchangeable, and the cheaper one is not a strict subset.
That gap decides things in practice. An experiment that only needs scene state, for example to compute a graph-structured target or to check what an agent can reach, runs entirely in Python and is fast and headless by construction. An experiment that needs pixels, segmentation masks or optical flow has no choice but to take on the executable, and with it the display server problem described further on.
The Unity executable is a manual download
Installing the Python package is one command:
pip install virtualhomeThe simulator is a separate step, and it is manual on purpose. There are prebuilt executables for Linux on x86-64, macOS and Windows, published for the current version, and you move the one you downloaded into the simulator directory inside the package before anything works.
That arrangement has a practical consequence: the version of the executable and the version of the Python package are pinned to each other by hand, and a mismatch is not something pip will resolve for you.
Running it locally is straightforward. Double-clicking the executable is enough to start it, and when running it from a terminal the documentation recommends windowed mode so the simulator does not take over the screen, using two flags, one to disable fullscreen and one to set screen quality.
A notebook with a demo and starting code ships alongside, so the first thing to do after the download is run the demo rather than write your own first script.
Headless means an X server, and the helper wants your GPUs
Testing the simulator on a machine with no monitor, which includes most compute servers, has two documented routes: Docker or an X server.
The X server route has more moving parts than it should, and the documentation is explicit about them. With an X server you run the executable in batch mode. On Linux you start the X server first, from a helper script that takes the display number and a list of GPUs, defaulting to all the GPUs available on the machine. Then, on a separate terminal, you launch the executable against that display:
DISPLAY=:display_num ./{path_sim}/{exec_file}.x86_64 -batchmodeThere is also a higher-level path for Linux, where you construct a communication object and hand it the executable file, a port and a display, and it opens the executable on the correct screen for you.
That path has the property that matters for research use: you can open several executables at once and drive them from multiple processes, which is how you generate a dataset faster than one render at a time or train a model against several simulators concurrently.
Reinforcement learning is a first-class consumer
The reinforcement learning story is older than the rest of the platform, which tells you what it is for.
There is a set of OpenAI Gym-like environments for training reinforcement learning agents, documented in the unity environment class. An earlier update added a demonstration of that, along with an example of combining environments with Ray, which is the specific detail worth remembering: the intended scaling path is many simulator processes behind a queue rather than one simulator doing more.
There is also an API for adding characters to a scene, for placing fixed cameras, and for recording from those cameras. Fixed cameras matter more than they sound, because a learned policy needs observations from a consistent viewpoint, and a moving viewpoint is a different problem.
Alongside the environments there is a dataset section and a section on modifying VirtualHome itself, which are the two other reasons to read past the quick start: one for consuming what others have generated, one for changing the simulator's own assets and scripts.
2.3 added procedural generation and a day/night cycle
Version 2.3 is the most recent tagged release, and its list is the best summary of where the project has been spending effort.
The headline addition is procedural generation, so agents can explore a much larger set of unique environments than the hand-authored set. Alongside it came more custom designed environments and enhanced simulated physics in them.
The second cluster is lighting and time. A time management system with synchronised day and night, new outdoor terrain with accurate sunlight and shadows, improved indoor realtime lighting, and more realistic rooms. That is a coherent group of work: making the environment vary over time in ways that affect what an agent can see.
The rest is housekeeping with real content behind it, namely significant performance enhancements, asset optimisations and stability improvements, documentation updates, and bug fixes for existing environments.
The list of what is in development is short and candid: further procedural generation enhancements, photorealism, more actions and object interactions, and human interaction. Two of those four are prerequisites for anything resembling a video dataset, which is the clearest statement of where this simulator is not yet.
A 2022 release with a 2026 commit
The version history is worth reading before you pin anything.
The tags are v2.1.0 in January 2021, v2.2.0 in March 2021 and v2.3.0 in March 2022. The package metadata also declares 2.3.0, so the published package and the newest tag agree. The last push to the master branch is dated May 20, 2026, which is recent activity on a release line that has not moved since early 2022.
The dependency pins explain part of why a major version has not come. The install requirements are a mix of exact pins and floors, and the exact ones are not all current: a specific certificate library version, a specific character detection library, an older internationalised domain library, a specific network library release, a specific OpenCV build, a specific plotting library release, a debugger and a terminal colour library. Nothing here is unreasonable for research software, but it does mean an existing environment with newer versions of those packages will not simply install.
The Python requirement is 3.10 or newer, and the repository is a small one: the package itself, an assets directory, a Docker directory, and the packaging files.
Editorial conclusion
VirtualHome fits embodied AI research that needs scripted household activities with ground truth rather than photographs of them, and that wants the scene described as data. It does not fit a pipeline that needs photorealism or unrestricted object interaction, since both are on the project's own in-development list. Before you plan a dataset around it, check the coverage gap between the two simulators, pick one deliberately, and note that the latest tagged release dates from March 2022 even though the repository is still being pushed.
Frequently asked questions
what is virtual home
VirtualHome is an interactive simulator for complex household activities, driven by a Python API. You write an activity as a sequence of instructions, pair it with a graph describing the environment, and the simulator renders it, letting agents pick up objects, switch appliances on and off and open them, with ground truth streamed alongside.
What is the difference between the Unity Simulator and the Evolving Graph in VirtualHome?
The Unity Simulator is built in Unity and generates videos, which means downloading the platform executable yourself. The Evolving Graph runs entirely in Python and produces a sequence of graphs as the program executes, though some objects and actions it supports are not yet supported in the Unity Simulator.
How do I run the VirtualHome simulator without a monitor?
Use Docker, or run an X server and launch the executable in batch mode. On Linux the X server is started through a helper script that takes a display number and a GPU list, defaulting to all available GPUs, and the executable is then launched against that display.
Can I run multiple VirtualHome simulators at the same time?
Yes. You can construct a communication object that is told which executable to open, on which port and display, and open several executables simultaneously to train models or generate data with multiple processes.
Does VirtualHome support reinforcement learning?
Yes. It includes OpenAI Gym-like environments for training agents, documented in the unity environment class, and an earlier example shows how to combine environments with Ray for scaling across processes.
What did VirtualHome 2.3 add?
Procedural generation for varied environments, more custom designed environments, enhanced simulated physics, a time management system with synchronised day and night, outdoor terrain with accurate sunlight and shadows, improved indoor realtime lighting, and a round of performance, asset and stability work.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xavierpuigf-virtualhome)