Mosaico: a data platform that turns robot sensor logs into a queryable archive
Mosaico - The data platform for Physical AI
At a glance
- What is it?
- A Python SDK and Rust daemon built around zero-copy reads, an immutable layer model, and an ML bridge that aligns unsynchronized sensor streams into tensors for training.
- Who is it for?
- Mosaico is a good fit for a robotics team drowning in `.bag` and `.mcap` files, where the real cost is not recording the data but answering questions about it six months later. The zero-copy retrieval model and the automatic stream indexing are the parts that pay for themselves, and the ML bridge removes the CSV staging step most teams hand-roll badly.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem it names: sparse asynchronous logs versus dense synchronous tensors
The README frames the whole product as a mismatch between two data regimes. Classical robotics, in its description, runs on an event-driven world where data is asynchronous and sparse, stored in monolithic sequential files such as ROS bags. A lidar might fire at 10Hz, an IMU at 100Hz and a camera at 30Hz, all drifting relative to one another.
Physical AI needs the opposite shape: synchronous, dense, tabular data, because models expect fixed-size tensors arriving at a constant frequency. The README's example is a batch of state vectors at exactly 50Hz.
Most of the engineering cost in a robotics ML team sits in that gap, not in the model. Aligning streams that never shared a clock, resampling them to a common rate, flattening the result, and writing it somewhere a training loop can read it is unglamorous work that every team rebuilds. Mosaico's stated contribution is to make that a platform concern.
The pitch is deliberately narrow. The README says Mosaico does one thing well, which is transforming monolithic sensor logs into a structured, queryable archive built for multi-modal data. The topics on the repository are data-platform, physical-ai and robotics, which is a consistent signal that this is infrastructure for a specific domain rather than a general-purpose time-series database with robotics examples.
Zero-copy retrieval and why bag file storage makes you parse everything
The storage claim is the technical centre of the project, and it is worth understanding precisely rather than accepting the adjective.
Mosaico describes its approach as a modern data lake layout with zero-copy architecture, which it says removes serialization overhead and allows direct, random access to a specific signal without parsing an entire file. The contrast is with `.bag` and `.mcap` storage.
That contrast is real and it is the reason to care. In a bag file, a single IMU channel is typically interleaved in a sequential stream with camera frames and lidar points, so extracting one signal means reading past everything around it. If you want 200 episodes of camera data and the IMU is interleaved in all of them, you read all 200. With an indexed, chunked store where streams are laid out independently and indexed automatically, you seek to the chunks you need and skip the rest.
The README also states that streams are automatically indexed, which allows the query engine to prune irrelevant data and stream only the precise slices required. Combined with the random access claim, that turns the archive from a write-once dump into something you can query without a batch job.
This is also where the comparison to a conventional time-series database becomes interesting rather than promotional. What Mosaico is optimising for is many wide, independent, multi-modal streams where you usually want a subset. If your workload is a single scalar metric at high resolution, the advantage is much smaller, and a purpose-built time-series engine may serve you better.
A strictly typed ontology instead of generic byte arrays
Underneath the storage format sits a semantic layer, and the README is direct about why: rather than treating data as generic byte arrays, Mosaico enforces a semantic understanding of every object. It claims this guarantees the data is validatable, optimised for transport and queryable by physical values.
That last phrase is the one to sit with. Queryable by physical values means a query can reason about a unit and a physical dimension, not just about a named column of floats. In practice that is the difference between a schema that records an acceleration in metres per second squared and one that records three unnamed float arrays.
The cost of a typed ontology is that you must map your sensors onto it, and the README does not document what happens to the ones that do not fit. This is the most likely place a real evaluation stalls. If your lab has a bespoke test rig with an unusual frame of reference, forcing it into the ontology is work, and forcing the ontology to accept it is probably a contribution back upstream. Worth asking the maintainers directly with your sensor list in hand rather than assuming the mapping is trivial.
The typing also pays off in the least obvious place: transport. If a stream knows its own type and unit, compression can be chosen per stream rather than applied uniformly, and a query planner can reject nonsense predicates before touching disk.
Immutable layers, and the versioning decision that will rule you in or out
This is the design choice most likely to determine whether Mosaico fits your organisation, and it is easier to state plainly than to argue around.
The README says Mosaico targets durable long-term storage and strict data lineage, and that it rejects traditional versioning because versioning introduces query ambiguity. Instead it uses immutable data layers, and the stated effect is horizontal growth where query history remains deterministic and immutable.
The reasoning has internal consistency. If a query written in January is re-run in October, immutable layers mean it returns what it returned in January. Under a mutable-versions model, the same query can return different data depending on which version the query engine resolves, and reproducing an old result becomes an exercise in archaeology.
The practical consequence is that there is no way to say give me the January version of this stream. What you can say is give me the immutable layer that existed in January. That is a narrower guarantee and it is the right one for reproducibility of results, but for a regulated setting where an auditor asks to see exactly what was captured before a correction, an unmutable dataset can be a problem rather than a feature.
Data lineage, which the README pairs with durability, is the term to interrogate here. The README asserts strict lineage as a design property but does not show a lineage API in the repository, so the practical answer to how you audit a specific training example back to its source frame probably lives in the documentation rather than in this page.
Monorepo layout: a Python SDK in front of a Rust daemon
The repository holds both halves of the system and states the reason: a monorepo configuration to simplify testing and reduce compatibility issues. The tree shows the split clearly with `mosaico-sdk-py/` for the client library and `mosaicod/` for the server daemon, alongside `docker/`, `docs/`, `scripts/` and a `CONTRIBUTING.md` with a `.pre-commit-config.yaml` at the root.
The split follows a standard client and server division. `mosaicod` is the central hub handling data conversion, compression and organised storage. The SDK is what you import into your scripts, and it manages the communication while abstracting the implementation so your API usage stays stable as the platform evolves underneath. That last guarantee is the one worth testing, since a data platform whose client API breaks every minor release is a real tax on everyone who writes code against it.
Mosaico calls itself strictly code-first, and the reasoning is that the project did not want to force another SQL-like sublanguage on users just to move data. Instead you get native SDKs, starting with Python. For teams already living in Python that is the right default. For a robotics stack written primarily in C++ or Rust, there is no second client listed, which means the Python SDK plus a local bridge is currently the whole story.
The daemon is distributed as a container, and the tree places the Docker material in the repository rather than pointing at a hosted service. Deployment is therefore yours to solve.
What the release history says about stability
There are three releases, and reading what they contain tells you more about the project's stage than the version numbers do.
v0.5.0, published 2026-05-28, is the only one with release notes. It lists features including a background cleanup routine, ordering support for strings, support for a list of fingerprints in revoke, an `api-key purge` command, an expanded `--version` option and CLI colouring. The bug fixes are more revealing: an inverted min and max comparison in `ChunkQueryBuilder::compile_clause` for `Op::Eq`, which is a query planner returning wrong results rather than erroring, plus a stream size increase on download, a broken Docker build workflow, and TCP bind failures from random port collisions in tests.
The `ChunkQueryBuilder` fix is the one to sit with. A chunked, indexed query path is exactly the kind of code where an inverted comparison produces plausible wrong answers instead of a failure, and it shipped and was caught. The performance tuning in that release is listed without numbers, which is normal but not verifiable.
The refactors section is where the instability shows. It lists decoupling the resource locator from the path in the store, moving both remote and local store configuration to environment variables, and changed logic for `create`, `finalize`, `do_put` and `delete` actions on sequences. Configuration moving to environment variables and action logic changing together is the profile of a project still finding its own shape.
Since then, v0.6.0 on 2026-08-07 and v0.6.1 on 2026-09-08 shipped with empty release bodies. The last push to the repository was 2026-09-18, so work is continuing, and 36 open issues is consistent with a pre-1.0 platform in active discussion rather than a settled one. Pin a version rather than tracking `main`.
Editorial conclusion
Mosaico is a good fit for a robotics team drowning in `.bag` and `.mcap` files, where the real cost is not recording the data but answering questions about it six months later. The zero-copy retrieval model and the automatic stream indexing are the parts that pay for themselves, and the ML bridge removes the CSV staging step most teams hand-roll badly. It is the wrong choice if you need data versioning, because the platform deliberately rejects it on the grounds that versioning makes historical queries ambiguous, and a team whose compliance process depends on being able to point at the exact bytes used six months ago will find that unacceptable. It is also pre-1.0, so the API is still moving. Check three things before committing: whether `mosaicod` runs under your deployment constraints given the Docker setup lives in the repository rather than on a hosted service, whether your sensor set maps cleanly onto the ontology rather than being forced into it, and whether a breaking release lands between your evaluation and your production cutover.
Frequently asked questions
What does Mosaico do with robot sensor data?
It transforms monolithic sensor logs into a structured, queryable archive built for multi-modal data. Streams are automatically indexed so the query engine can prune irrelevant data and stream only the slices required, with direct random access to a specific signal instead of parsing whole files.
Does Mosaico support data versioning?
No, and this is deliberate. The README states that Mosaico rejects traditional versioning because it introduces query ambiguity, using immutable data layers instead so query history stays deterministic. You can reference the layer that existed at a point in time, but you cannot retrieve an earlier mutable version of the same stream.
How do I use Mosaico, Python or a query language?
The project describes a strictly code-first approach with native SDKs, starting with Python, rather than another SQL-like sublanguage. The Python SDK is what you import into scripts; the mosaicod daemon handles conversion, compression and storage behind it. There is no second client SDK listed.
How does Mosaico handle unsynchronized sensor streams for model training?
Its ML module ingests raw unsynchronized data and transforms it on the fly into aligned, flattened formats ready for model training. The README frames this as removing the need for large intermediate CSV files between recording and training.
Is Mosaico open source and what license does it use?
Yes, under Apache-2.0. The repository holds both the mosaico-sdk-py Python SDK and the mosaicod Rust backend in one monorepo, chosen to simplify testing and reduce compatibility issues between the two.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mosaico-labs-mosaico)