# Mirobody: a self-hosted LOINC resolver and health data engine

> Mirobody turns lab reports, wearables and genomics into LOINC-coded, UCUM-normalized, FHIR-ready records, and ships an agent that reads the original documents. The library installs in two packages; the full stack runs from one deploy script.

**thetahealth/mirobody** — The AI-native health data engine — collect, standardize, and reason over labs, wearables & genomics.

- Repository: https://github.com/thetahealth/mirobody
- Website: https://mirobody.ai
- Stars: 1,349 · Forks: 236
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/thetahealth-mirobody

## The problem Mirobody actually solves: two reports, one test, two spellings

The README states the problem in one line: two reports, same test, different spelling and units. A lipid panel printed in Shanghai and the same panel exported from a US lab may describe total cholesterol with the same words but different units, and a Chinese report may name hemoglobin in characters that share no substring with the English term. Any downstream reasoning step, whether a rules engine or a language model, has to compare those readings, and it cannot do that until the names and units collapse onto one identity.

Mirobody is aimed at that collapse. The intended audience is developers and healthcare-industry builders who receive heterogeneous health data and want a canonical form: LOINC codes, UCUM-normalized units, FHIR-recognized code systems. The repository classifiers list the intended audience as Developers and Healthcare Industry, with a development status of Beta. It is not a consumer app. The README notes that the engine powers Theta Wellness, a live consumer product, but the repository itself is the engine and the tooling around it.

The scope is deliberately three stages. Collect pulls signals from three device providers, a SQL source, seven file formats, and Apple Health through a signed iOS client that POSTs data in. Translate resolves readings to canonical codes. Answer runs an agent over the original documents. Everything in the codebase, the docs and the contributing guide is organized around those three stages, which is a useful signal: the project has one story rather than a pile of features.

## How the resolver works, and why it refuses to guess

The core mechanism is a concept graph, and the README gives its shape: 440,961 nodes and 22,044,110 cross-vocabulary edges, with 595,746 source ids distilled into canonical concepts across LOINC, SNOMED CT and RxNorm. On top of that sit 49,253 multilingual aliases, including 22,578 in Chinese and 16,809 in Japanese. The resolver runs offline, which matters if you are handling health data and do not want a network round trip per term.

The design decision worth noticing is abstention. The README's example uses 血脂, which means lipids. That is a category, not an observation, so the resolver returns nothing rather than a plausible wrong code. The same principle applies to the return object: an empty code is described as a gap worth a second look, while method="refused" is described as a decision. Those are different outcomes and the API distinguishes them. For a clinical data pipeline, a wrong code that looks right is worse than a gap, so this is the right default even though it pushes work back onto the caller.

The unit handling is where the graph earns its keep. LOINC encodes the unit and the result type into the identity, so the same name resolves to different codes depending on context. The README shows total cholesterol resolving to 2093-3 in mass per volume and 14647-2 in moles per volume, and a neutrophil percentage resolving to 26511-6 while a neutrophil count resolves to 26499-4. There is a molar-mass bridge keyed by LOINC code to convert between those. The practical consequence is that you should pass the value and the unit whenever you have them, because the unit picks the code. A resolver that only saw the name could not make that distinction.

Two more details from the README are worth flagging. Traditional Chinese gets a shipped 3,336-character zh-Hant to zh-Hans fold table plus curated rows, and a curated row always beats a fold. And the bundle is pinned: mirobody.BUNDLE_VERSION reports loinc-2.82+2026.08.28-af2524b7a285, a release, cut date and a digest over the bundle's own members. The README explicitly addresses why 2.82 and not 2.83, so if you need 2.83 codes, read that document before assuming the gap is accidental.

## Installing Mirobody and resolving your first panel

The library install is small on purpose. The README states that pip install mirobody is two packages and 52 MB, on numpy only. There is a hard floor of Python 3.12 or newer, and pyproject.toml explains why: the codebase uses PEP 701 f-strings with nested same-type quotes, which are a SyntaxError on 3.11, so the package cannot even be imported there. If you are on 3.11, upgrade before you try.

The quickest first use needs no key, no config and no network. The README gives this command:

```bash
pip install mirobody
mirobody resolve "LDL cholesterol" 血红蛋白 ヘモグロビン "空腹血糖(GLU)" 血脂
```

You should see the Chinese and Japanese terms for hemoglobin land on the same code as the English word, LOINC 718-7, and the fasting glucose term resolve to its own code. 血脂 returns nothing, because it is a category rather than an observation.

From Python, the same resolver is importable, and the README's example shows the unit-sensitive behaviour directly:

```python
from mirobody.engine import resolve, resolve_reading

resolve("血红蛋白").loinc                                # '718-7'
resolve("total cholesterol").loinc                     # '2093-3'  [Mass/volume]
resolve_reading("total cholesterol", "5.0", "mmol/L")   # '14647-2' [Moles/volume]
resolve_reading("total cholesterol", "193", "mg/dL")    # '2093-3'  the unit picks the code
```

If you want the rest of the platform, the README's path is a clone, a Git LFS pull for the data bundles, and a deploy script that brings up Postgres with pgvector:

```bash
git clone https://github.com/thetahealth/mirobody.git && cd mirobody
git lfs install && git lfs pull
./deploy.sh
```

The LFS step is not optional. The README warns that a fresh clone holds pointer stubs until you run it, so the engine's data bundles are not present otherwise. The compose file pins pgvector/pgvector:0.8.6-pg17-trixie and exposes Postgres on 127.0.0.1:18062 and Redis on 127.0.0.1:18069, both loopback only, with placeholder passwords that the file itself tells you to replace in production.

## Where Mirobody is the wrong tool

The README is unusually candid about one boundary: the opt-in semantic tier cannot abstain. The default lexical path refuses when it does not know a term, but if you turn on the semantic tier you trade that refusal for coverage. In a clinical pipeline that trade is not free, and the README points to a standardization document rather than pretending the choice does not exist. If your use case cannot tolerate an occasionally wrong code, stay on the default path and handle the gaps yourself.

The coverage claim is also narrower than it first reads. The README states 211 out of 211 on the panels an ordinary checkup prints, in English, Chinese (Simplified and Traditional) and Japanese, verified by a test file in the repository. That is a checkup panel, not all of laboratory medicine. The README separately raises what LOINC covers of the wearable world, which is a hint that the answer is not everything. If your data is mostly continuous sensor streams rather than named lab observations, the resolver is not the part of this project you need.

The project describes itself as Beta in its own classifiers. The README does not document rollback for deploy.sh, and the compose file's passwords are placeholders by design, so treating the default deployment as production-ready would be a mistake. There is also a version floor you cannot negotiate around: Python 3.12 or newer, or the package will not import.

## Mirobody compared with FHIR-first tooling and LLM extraction

The obvious alternative approach is to skip the resolver and let a language model read the report directly, mapping names to codes as it goes. That is faster to prototype and needs no terminology bundle. The difference is failure behaviour. A model asked for a LOINC code will usually produce one, including when the term is a category like 血脂 that has no single observation code. Mirobody's default path returns nothing in that case and labels it as a refusal. The README also argues that AI can only reason over what it can read, which is why the agent stage reads the original documents through a virtual filesystem and cites the page it read from, rather than relying on a pre-extracted summary.

A second alternative is FHIR-first tooling that assumes the data already arrives in FHIR. That works when your sources speak FHIR. Mirobody's premise is that they do not: the README's Collect stage lists three device providers, a SQL source, seven file formats, and Apple Health as receive-only. If your inputs are already FHIR resources, the Translate stage is doing work you have already done, and you would be adopting a large terminology bundle for a smaller gain.

What Mirobody adds on top of either is the packaging. It is Apache-2.0, self-hosted, and the same tools are served over MCP to Claude Desktop, Cursor or your own agent. The MCP surface is the part that is hardest to replicate with a prompt, because it exposes the resolver and the reading tools rather than asking a model to remember codes.

## Maintenance, licensing and the upgrade surface

The repository is not archived, and the last push was on 2026-09-10, with releases 1.4.0 on 2026-09-07, 1.3.0 on 2026-08-31, and a terminology build data release on 2026-09-10. That is a recent cadence, and the terminology data ships as its own release stream, which is the right structure: you can update the concept graph without changing the engine version.

The licence is Apache-2.0, and there is a separate LICENSE-3RD-PARTY file at the repository root. That second file is the one to read before commercial use, because the terminology bundles are distilled from LOINC, SNOMED CT and RxNorm, and those vocabularies carry their own terms that Apache-2.0 does not override. The README does not spell out those obligations, so treat the third-party licence file as the authoritative list rather than assuming the top-level licence covers everything. This is not legal advice; it is a pointer to the file that answers the question.

The upgrade cost has two axes. Engine upgrades are ordinary Python version bumps, with the caveat that base dependencies use floors only: pyproject.toml states that an upper cap needs a reproduced breakage documented next to it, because unjustified caps create resolver conflicts downstream. That is a deliberate policy and it means you own your own lockfile. Terminology upgrades are the other axis, and they are pinned by digest. A code that resolves under loinc-2.82+2026.08.28-af2524b7a285 may resolve differently or not at all under a later bundle, so if you store resolved codes, record the bundle version alongside them.

## Conclusion

Adopt Mirobody if you ingest multilingual lab reports or wearable exports and need canonical LOINC codes plus a FHIR-shaped output you can host yourself. Do not adopt it expecting a finished clinical product: it is Beta, its semantic tier cannot abstain, and the README does not document rollback for the deploy script. Verify first that your Python is 3.12 or newer, that `mirobody resolve` returns codes for your own panel names, and that your LOINC use case is covered by the pinned 2.82 bundle rather than 2.83.

## FAQ

### What does Mirobody mean and what is the project?

Mirobody is the name of the engine; the README describes it as an AI-native health data engine that collects, standardizes and reasons over labs, wearables and genomics. It resolves readings to LOINC codes, normalizes units to UCUM, and produces FHIR-ready records. The same name covers the Python library on PyPI and the self-hosted stack in the repository.

### What Python version does Mirobody need?

Python 3.12 or newer. pyproject.toml sets requires-python to >=3.12 and explains that the codebase uses PEP 701 f-strings with nested same-type quotes, which are a SyntaxError on 3.11, so the package cannot be imported there.

### Why does mirobody resolve return nothing for some terms?

The resolver abstains rather than guessing. The README's example is 血脂, which is a category rather than an observation, so no single LOINC code applies and the resolver returns nothing instead of a plausible wrong code. An empty result is a gap; a result with method="refused" is a separate decision.

### Can I run Mirobody entirely offline?

The resolver runs offline, and the README's first example needs no key, no config and no network. The full stack is self-hosted through deploy.sh, which brings up Postgres with pgvector and Redis on loopback ports. The README does not state that every feature works without network access.

## Sources

- [License: Apache-2.0](https://github.com/thetahealth/mirobody/blob/main/LICENSE)
- [Project website](https://mirobody.ai)
- [README](https://github.com/thetahealth/mirobody/blob/main/README.md)
- [Releases](https://github.com/thetahealth/mirobody/releases)
- [thetahealth/mirobody on GitHub](https://github.com/thetahealth/mirobody)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thetahealth-mirobody
