Framework
facebookresearch/ai4animationpy avatar
facebookresearch/ai4animationpy

AI4AnimationPy: character animation research that dropped the Unity dependency

A Python framework for AI-driven character animation using neural networks.

2,107 stars257 forksPythonNOASSERTION

At a glance

What is it?
A Python port of Meta's AI4Animation work that keeps the game engine architecture, adds a raylib renderer, and lets you backpropagate through inference. The catch is a CC BY-NC licence and a very tight Python floor.
Who is it for?
AI4AnimationPy earns its place for animation research where the loop between training and looking at the result is the bottleneck, since backpropagating through inference is impossible in the Unity original and the built-in renderer means a training run and its visualization happen in the same process.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 56 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why the Unity dependency was the problem

The README states the motivation as friction, not novelty. Research on AI driven character animation has required juggling multiple disconnected tools: model work happens in Python while visualization needs specialised software, and bridging the two means building a communication pipeline between them. The original AI4Animation project leaned heavily on Unity, which the README credits with being useful for visualization and runtime inference while also noting that talking to PyTorch had to go through ONNX or data streaming.

That detour is what this project removes. AI4AnimationPy fuses everything into one framework running only on NumPy and PyTorch, so training, inference and visualization share a backend. The README puts the old workflow in a table with concrete claims: generating training data for 20 hours of mocap takes under 5 minutes here against more than 4 hours in the Unity version, and setting up a new experiment takes about 10 minutes against more than 4 hours. Those are the project's own numbers, so treat them as claims rather than measurements, but the shape of the difference is what you would expect once the bridge disappears.

Two consequences in that table matter more than the timings. Visualizing inputs and outputs during training is built in rather than requiring streaming, and backpropagating through inference is supported, which the Unity version cannot do at all. That second one is a research capability, not a convenience.

Three execution modes and what manual control buys

The framework runs in one of three modes. Standalone uses the built-in rendering pipeline. Headless runs without a window, which is the mode for server side training. Manual execution lets you decide when and how often the update loop runs, and the README notes this is what enables running code locally or remotely on a server.

The distinction between Standalone and Headless is that both invoke automatic update callbacks, whereas Manual gives you the callback. That is the hook a researcher needs for stepping a simulation frame by frame to inspect an intermediate state, and it is the difference between a framework you observe and one you interrogate.

This design only makes sense because of the ECS choice. The feature table lists an Entity-Component-System with modular architecture and lifecycle management, plus an update loop with game engine style callbacks named Update, Draw and GUI. Borrowing that structure from game engines is not incidental: it gives the framework a stable place to hang per-frame work, which is what makes a renderer, an IK solver and a neural network step coexist in one process without one owning the loop.

The trade is that you inherit game engine conventions. If your background is data pipelines rather than engines, the ECS vocabulary is a cost, and the architecture overview page in the documentation is the place to absorb it.

Loading motion from GLB, FBX and BVH into npz arrays

Motion capture import is the entry point most people need first. The framework reads mesh, skin and animation data from GLB, FBX and BVH, and the README gives the API directly:

python
from ai4animation import Motion

motion = Motion.LoadFromGLB("character.glb")
motion = Motion.LoadFromFBX("character.fbx")

There is a matching `Motion.LoadFromBVH` for BVH files. BVH is the interesting one for motion capture work specifically, since BVH is the interchange format most motion capture tools emit.

The internal representation is worth knowing because it determines what you can do with NumPy and PyTorch directly. The internal motion format is `.npz`, storing 3D positions and 4D quaternions for each skeleton joint per frame. So a loaded motion is arrays of positions and rotations indexed by joint and frame, which is exactly the shape a tensor library wants. There is no conversion step between loading a file and running a network over it.

The three importers matter because they cover the two ends of a pipeline. FBX arrives from a DCC tool, GLB arrives from a real time export, and BVH arrives from a capture session. Having all three in one framework means you can iterate on a motion without a round trip through a separate tool.

What is built in: raylib rendering, IK, and the neural network types

The rendering stack is the most unusual thing about this project. Rather than outsourcing visualization to a browser or a separate viewer, the framework includes a real time renderer with deferred shading, shadow mapping, SSAO, bloom and FXAA, plus GPU accelerated skinned mesh rendering. That is a substantial graphics stack to find in a research framework, and it is what allows training and visualization to share a process.

The rendering dependency is explicit in `setup.py`, which pins `raylib==5.5`. That same file gives you the packaging picture: the distribution is named `ai4animation` at version 1.0.0, the importable package directory is `ai4animation`, and `python_requires` is `>=3.12.12`, an unusually specific floor for a Python 3.12 project. One console script is registered, `convert`, pointing at `ai4animation.Import.BatchConverter:main`, so batch conversion of imported assets has a command line entry point.

The rest of the stack is a conventional research toolkit: a math library with vectorized forward kinematics, quaternions, axis-angle, matrices and mirroring; neural network types listed as MLP, autoencoder and codebook matching with training utilities; inverse kinematics via a FABRIK solver; animation modules for joint contacts and root and joint trajectories; and a camera system with free, fixed, third person and orbit modes with smooth blending.

Inverse kinematics deserves a note. FABRIK is a popular iterative solver, and having it built in means you can correct a generated pose to hit a target without leaving the process, which is the kind of thing that otherwise needs a separate solver bolted on.

What is listed as not done yet

The feature table is honest about its gaps, using a planned marker for three items: physics simulation covering rigid bodies and collision, path planning and spline tooling, and audio support.

Those three gaps cluster in one direction. Everything the framework currently does is about producing and looking at motion: import it, run a network over it, pose it, correct it, render it. Anything that requires the motion to interact with a world, which is to say physics, navigation and sound, is not there. That is a coherent boundary rather than an unfinished list, and it tells you what kind of problem this framework is aimed at.

Audio support is the one to watch, because `setup.py` pulls in `torchaudio`, `sounddevice` and `soundfile` even though the feature table marks audio as planned. The dependencies are present for work in progress rather than for a finished capability.

The dependency list also shows the intended workload. `numpy==1.26.4` and `torch>=2.0.0` with matching `torchvision` and `torchaudio`, plus `scipy`, `scikit-learn`, `einops`, `matplotlib` and `tqdm`. The exact numpy pin and the bare `tqdm` sit oddly together, and the exact pins will be a maintenance consideration if you need to share an environment with other work.

The demos section points to a set of specific examples: a stylized biped locomotion controller trained on style100, a quadruped controller with gait transitions, future motion anticipation with training visualization, the ECS itself, inverse kinematics, a motion editor, and motion capture import. The motion editor for browsing and feature visualization is the one that tells you this is aimed at inspecting data, not just running a model.

The licence rules out commercial use as published

The licence badge at the top of the README reads CC BY-NC 4.0, and the LICENSE file at the repository root is the reference. The non-commercial clause is the part that matters for anyone beyond research.

This is not an unusual licence for research code and not a defect in it. A framework published by a research group, with authors named as Paul Starke and Sebastian Starke, and hosted under the facebookresearch organisation, is usually released so that others can build on the research rather than so that a studio can ship it. But it does change who can use it. If your plan involves character animation in a commercial product, a CC BY-NC 4.0 codebase is not a starting point, and no amount of its technical fit compensates for that.

The rest of the practical picture: the repository is not archived, the last push was on 2026-08-14, and it has no tagged releases at all, so there is no version to pin beyond the `1.0.0` string in `setup.py`. Documentation is on GitHub Pages and split into installation instructions, a quick start guide, an architecture overview, demo programs and an API reference, which is a better documented arrangement than most research repositories manage.

There is also a hosted set of web demos, plus video demos in the repository's `Media/` directory and a `Demos/` folder in the tree, so you can see what the locomotion controllers actually do before installing anything.

Editorial conclusion

AI4AnimationPy earns its place for animation research where the loop between training and looking at the result is the bottleneck, since backpropagating through inference is impossible in the Unity original and the built-in renderer means a training run and its visualization happen in the same process. It is a poor fit for commercial use as shipped, because the licence is CC BY-NC 4.0, and it is a poor fit if you need physics, path planning or audio, all three of which the feature table marks as not yet done. Install on Python 3.12.12 or newer, import motion with `Motion.LoadFromGLB`, and read the architecture overview before wiring anything into a pipeline that assumes those planned systems exist.

Frequently asked questions

What is AI4AnimationPy used for?

It is a Python framework for neural network driven character animation, covering motion capture processing, training, inference and animation engineering. It brings the AI4Animation project to Python, removing the Unity dependency and keeping a game engine style architecture.

How do you load a motion capture file in AI4AnimationPy?

Import the Motion class from `ai4animation` and call one of the loaders, such as `Motion.LoadFromGLB("character.glb")`, `Motion.LoadFromFBX` or `Motion.LoadFromBVH`. Internally a motion is stored as npz arrays of 3D positions and 4D quaternions per joint per frame.

Does AI4AnimationPy need a graphics card or a display?

It has three execution modes: Standalone with the built-in raylib renderer, Headless for server side training, and Manual for controlling the update loop yourself. The renderer uses deferred shading, shadow mapping, SSAO, bloom and FXAA, and can be skipped entirely in Headless mode.

What license does AI4AnimationPy use and can I use it commercially?

The README badge and the LICENSE file state CC BY-NC 4.0, which is a non-commercial licence. It suits research and personal experimentation; commercial use is not covered by it.

What features are not yet implemented in AI4AnimationPy?

The feature table marks physics simulation for rigid bodies and collision, path planning and spline tooling, and audio support as not yet available. Present capabilities cover ECS, the update loop, math, neural networks, rendering, skinned meshes, inverse kinematics, animation modules, cameras and motion import.

Official sources

  1. facebookresearch/ai4animationpy on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-ai4animationpy.svg)](https://hysenlabs.com/projects/facebookresearch-ai4animationpy)