# XR Animator: webcam mocap in a web worker, and a repository with no licence file

> XR Animator turns a single webcam into full body and face motion capture that drives a VRM or MMD avatar, running pose detection in a web worker on an offscreen canvas so the browser stays at sixty frames a second, and it publishes real frame rate figures against a named entry level graphics card. What an adopter has to weigh alongside that is administrative: there is no licence file in the repository, the assets ship as an encrypted archive, and the maintainer has written plainly that the project may not continue.

**ButzYung/SystemAnimatorOnline** — XR Animator, AI-based Full Body Motion Capture and Extended Reality (XR) solution, powered by System Animator Online

- Repository: https://github.com/ButzYung/SystemAnimatorOnline
- Website: https://sao.animetheme.com/XR_Animator.html
- Stars: 1,899 · Forks: 174
- Language: JavaScript
- License: not declared
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/butzyung-systemanimatoronline

## A repository that is also a website, with a page per browser engine

The top-level file list is the most surprising thing about this project and it explains a great deal. There is no source tree and no build configuration in the conventional sense. What there is, is a set of HTML entry points, one per host environment: a WebKit page, a page for an embedded Chromium framework, a page for the retired Internet Explorer engine, an HTML application for that same engine, a XUL page, a Silverlight page, and a separate page for a multiplayer build. Alongside them sit three vendored JavaScript renderers, one for each of the two main 3D library generations plus a MMD-specific module, a committed node_modules directory, a stylesheet directory, and a sound directory. The application is a static page, and shipping it inside five different shells is a matter of pointing a different wrapper at the same files. Two more files explain the hosting. One disables the static site generator that would otherwise process a repository served as a website, and one names the custom domain, which is the project's own site. So the same repository is the source, the build output and the published website, and the multiple page variants are variants of the same static application. That has a real consequence for you as a user: you are not installing a package, you are loading a page, and the differences between the versions are which shell hosts it rather than what the code does.

## The shell disagreement, and a window configured for a virtual camera

One small discrepancy is worth surfacing because it is exactly the kind of thing that costs an hour. The readme says the desktop application is powered by one particular Node-plus-Chromium shell. The manifest at the top of the repository names a different one in its very first field. Both approaches embed a Chromium engine in a Node process, so the practical difference between them is small, but the documentation and the manifest disagree and you should check which one the current release actually uses before you write a build or packaging script against it. The rest of that manifest is more revealing than the shell name, because it is configured for one specific job: presenting an avatar over a live camera feed. The window is declared frameless, with no toolbar and not resizable, and transparent, so the 3D output composites directly over whatever is behind it. The chromium arguments enable audio input and ignore the graphics blocklist, the second of which matters on exactly the machines people use this on. And node remote is bound to a local scope. Put together with the feature list, where a frameless transparent window usable as a source in broadcasting software is called out as a feature, the design intent is clear: this is meant to sit on top of a camera preview and look like the person speaking is the avatar, which is the whole premise of the format. The icon field points at a specific image rather than a generic one, which is the last sign that this is a finished product rather than a demo.

## Pose detection in a worker on an offscreen canvas, and what it costs

The technical claim is specific and checkable. Pose detection comes from a published machine learning solution that estimates three dimensional joint positions from a live camera image, and the readme states the architecture that makes it fast enough to be usable in a browser: on engines that support both a web worker and an offscreen canvas, the application achieves sixty frames a second of visual rendering and thirty frames a second of pose detection on what it calls a mediocre personal computer. That is a real design decision rather than a slogan, because the alternative is running inference on the main thread, which would block the render loop and give you a slideshow. The face side is different in kind, and the readme is precise about it: the application supports fifty-two facial blend shapes matching the augmented reality toolkit standard, which is why any avatar authored for that standard works without retargeting. Tracking is configurable across three regions, the face, the body, and the hands, in any combination, so you can run a full body pass and disable the face for performance. On phones the readme advises the same thing in plainer terms, limiting the load to face tracking, which is honest, because a phone has no discrete graphics card and inference is the expensive half. There is also a mode that is not pose at all: tracking real objects seen by the camera and mapping them onto 3D props, which is a different model doing a different job and which is worth knowing about if you want a held prop to appear in the scene.

## Frame rates against a named graphics card, and the dual-GPU trap

This is one of the few readmes that quantifies its own performance claims, and the specificity is what makes it useful. The claim is that system requirements are low enough for laptops and even phones, and the supporting data names the hardware: an entry level personal computer with a graphics card from a particular generation of consumer part, running full body capture, and the expected results are twenty frames a second or better for pose and finger tracking, forty or better for face tracking with an explicit cap at thirty, and sixty frames a second for the 3D rendering. Three details in those numbers are worth more than the numbers themselves. The face figure being capped at thirty says the author knows inference above that rate adds nothing, which is an engineering judgement rather than a hardware limit. The separation of pose from rendering tells you which part of the frame budget is the constraint, because the rendering is comfortably at the display rate and the inference is not. And naming a specific card rather than a class of hardware means you can compare against what you actually own. The troubleshooting note is the most practically useful paragraph in the readme. On a laptop the application may end up on the slower integrated graphics, particularly on a machine with two graphics processors, and the fix is in the system's graphics settings rather than in the application, with a link to a guide for doing it on one desktop operating system. That is the failure most people will hit, it is invisible from inside the app, and the readme names it before you lose an afternoon to it.

## Interoperability is the real argument: five motion formats and a live protocol

If you have an existing animation pipeline, this is the section that decides it. The readme lists what the application can read and what it can write, and the lists overlap in a useful way. It can record captured motion and export it in three formats, and it can load four. It can also convert between them, taking a motion in one of three formats and writing the result in another. The formats are the ones the MMD and avatar community actually uses, which means a capture session produces something your existing tools can open rather than a proprietary file. The second half of the interop story is a live protocol, and it is the more interesting one: the application can drive a 3D model in a different application entirely, and the readme names three that speak it. That turns this from an editor into a source, so a capture running here can animate a model in another program while you work. The catch is stated plainly in the same bullet: the live protocol works only in the desktop application, not in the browser version. So the tiering across the whole project is consistent, and it is worth internalising before you choose a version. The browser build is the portable one, works everywhere including phones, and gives you capture, export and import. The desktop build is the capable one, and adds the live protocol, a transparent frameless window for broadcasting, a desktop wallpaper mode, and running a depth-based backdrop as a separate window. If your work needs any of those, the desktop application is not a convenience, it is the only option.

## A flat image becomes a 3D backdrop from a generated depth map

There is a second machine learning model in this project and it is unrelated to bodies. Given a flat image, the application can generate a depth map from it automatically and use the result to turn the image into a three dimensional backdrop, so a photograph or a painting acquires parallax when the camera moves. That is a different technique from pose estimation and it solves a different problem, and it is the feature most likely to be unfamiliar to you. The readme describes three ways of using it. In the browser it turns a 2D image into a 3D image viewer, so the depth effect is something you can look at rather than something you have to set up. In the desktop application it can run that depth-based backdrop independently on the Windows desktop as a wallpaper gadget, which means the parallax follows the desktop rather than a window. And in the ordinary editing mode you can assign any image as the scene backdrop. The scene in general is customisable with flat images or video, spherical panoramas, and 3D objects in two file formats, so the backdrop is one layer of a scene you compose rather than a fixed background. It is also worth noting what the depth map is and is not. It is a single-channel estimate produced by a model, not measured geometry, so it will be wrong at edges and it will not know that a foreground object is a person. For a backdrop behind an avatar that is a completely reasonable trade, and for anything where the geometry has to be correct it is not, and the readme does not claim otherwise.

## Augmented reality is a chain of four dependencies, and the readme gives the checklist

The augmented reality mode is the feature most likely to fail on a given phone, and the readme is honest about why by spelling out the requirements as a numbered list rather than a promise. You need a phone on the published list of devices that support the platform's augmented reality technology, you need that platform's services for augmented reality installed from the app store, you need the Chrome browser, and you need the browser to expose the newer immersive web interface. That is four conditions, three of them supplied by one vendor, and any one of them being absent means the mode simply does not appear. The interaction itself is described precisely enough to be reproducible. Once the page has fully loaded you activate the mode from a small phone button in one of two menu positions, the camera view appears, you point it at the surface where you want the model, a white circle appears to indicate a candidate placement, and a double tap confirms it. A second double tap brings the circle back so you can place it elsewhere. Note that the placement is a tap on a detected surface rather than a scale-and-rotate gesture, which tells you the implementation is surface detection with a fixed model placement and not a full augmented reality editing tool. The readme also links two demonstration playlists for the mode, which is the right thing to do for a feature this dependent on hardware. My honest summary is that this is a well documented feature with a real dependency chain, and the documentation is the best part of it.

## No licence, an encrypted asset archive, and a repository that keeps its history in place

This is the section that decides whether you should use this at all, and none of it is in the feature list. There is no licence file in the repository. The description of the project records no licence either. That means there is no grant of any kind, so there is no stated basis on which you may copy this code, modify it, or redistribute it, and no stated basis on which you may redistribute the assets, which is a separate question from the code. The assets also ship as a single encrypted archive file at the top level, which is consistent with a project that has not resolved redistribution and has chosen obscurity instead of a licence. Three vendored JavaScript libraries and a committed node_modules directory sit alongside, each with its own copyright, in a repository that has not documented any of them. So the legal position of a repository that a user is invited to download and run is unresolved, and that is a conversation to have with the author before you put it in front of anyone. The second half of the story is about continuity. The support section states that family circumstances have significantly increased the author's financial burden, that financial return from the work was minimal, and that reality forces an evaluation of the project's sustainability, with the possibility of having to give up stated plainly. The primary distribution channel offered in response is a paid membership, and its headline benefit is access to new versions at least nine months ahead of the public release, alongside insider material. A named list of individual sponsors follows. Read together, the position is a capable and actively developed project with an unresolved licence and a maintainer who has publicly warned that it may end. Use it, enjoy it, export your work, and do not build a dependency you cannot survive losing. The naming is worth sorting out because a search for the repository name will not find the product. The repository is named for the online version of the earlier project, the product is called something else entirely, and the readme says it is inherited from a previous desktop gadget project with a third name. All three refer to the same codebase at different stages, and the marketing name is the current one. The release record is on the marketing name, at version zero point three four point three in August 2026, preceded by zero point three four point two in June and zero point three four point oh in May, so roughly one release every two months on a version line that is still below one. The last push was on 2026-09-18, so the work is current. Distribution is two channels: a web application served from the project's own domain, and a desktop application published as releases for three operating systems, with the readme pointing at both. And then there is the accumulation, which is the clearest sign of a long-lived personal project. There are two readme files whose names differ only in case, a plain text readme alongside the markdown one, a separate readme for the multiplayer build, a plain text changelog rather than a markdown one, a blank pair of pages kept for the browser to load, a committed temporary directory, a configuration file holding a default path for one shell, a Windows application directory, an application manifest for a legacy Windows shell, and a project descriptor. None of that is a criticism of a developer who has shipped a real product for years. It is a signal about the shape of the project: one person, a long history, files kept rather than removed, and a rewrite of the front end over time with the old entry points left in place so old links keep working.

## Conclusion

XR Animator is the right tool if your content is a VTuber avatar or an MMD model and you want motion capture from a webcam with no depth camera, no mocap suit and nothing installed, because MediaPipe pose estimation from a single camera is exactly that and the interop is real rather than aspirational, with VMD, BVH, glTF, FBX and VRMA both readable and writable plus a live protocol for driving other applications. Two things have to be settled before you build a pipeline on it. The licence, because there is none in the repository, which means there is no grant to rely on and no basis for redistribution, and that conversation is the first gate rather than a formality. And the maintenance outlook, because the maintainer has said in the readme that financial pressure may force the project to end, the primary distribution channel is a paid membership offering access nine months ahead of the public release, and the version line is still below one with releases every few months. Given that, use the web version, keep your own copies of your exported motion, and treat the desktop application and its extra features as something you may lose rather than something you depend on.

## FAQ

### What does XR Animator need to run in a browser?

For the performance the readme quotes, an engine supporting both a web worker and an offscreen canvas, which is how detection is kept off the main thread. On a browser without those, or on a phone with limited processing power, the readme advises limiting the load to face tracking. The web version needs no installation at all.

### Which motion formats can it read and write?

It can record capture and export in VMD, BVH and glTF, load VMD, FBX, BVH and VRMA, and convert from the latter three into VMD. There is also a live protocol that can drive a model in another application, and that one is available only in the desktop application rather than the browser version.

### What frame rates should I expect?

On an entry level personal computer with a GTX 1650 class graphics card running full body capture, the readme states 20 or more frames per second for pose and finger tracking, 40 or more capped at 30 for face tracking, and 60 for the 3D rendering. On a laptop with two graphics processors, the common cause of poor performance is the application running on the slower integrated one, fixed in the system graphics settings.

### What is the difference between the web version and the desktop application?

The desktop application, for Windows, Linux and macOS, adds the live interop protocol, a frameless transparent window for use as a source in broadcasting software, a desktop wallpaper mode, and the ability to run a depth based backdrop as its own window. The browser version is portable, works on phones, and covers capture, import and export.

### Can I use this commercially?

That is not something the repository settles. There is no licence file and no licence is recorded for the project, so there is no grant of any kind to rely on, and the bundled assets ship as an encrypted archive. The question to put to the author is whether you may use, modify and redistribute it, and until that is answered in writing the answer is not yes.

## Sources

- [ButzYung/SystemAnimatorOnline on GitHub](https://github.com/ButzYung/SystemAnimatorOnline)
- [Issues](https://github.com/ButzYung/SystemAnimatorOnline/issues)
- [Project website](https://sao.animetheme.com/XR_Animator.html)
- [README](https://github.com/ButzYung/SystemAnimatorOnline/blob/master/README.md)
- [Releases](https://github.com/ButzYung/SystemAnimatorOnline/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/butzyung-systemanimatoronline
