# Synthea: a synthetic patient population simulator for FHIR test data

> Synthea generates synthetic patient records from birth to death and exports them as FHIR, C-CDA, CSV or CPCDS. It suits teams that need realistic test data without touching real patients, and it is a poor fit for anyone who wants a hosted service.

**synthetichealth/synthea** — Synthetic Patient Population Simulator

- Repository: https://github.com/synthetichealth/synthea
- Website: https://synthetichealth.github.io/synthea
- Stars: 3,367 · Forks: 945
- Language: Java
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/synthetichealth-synthea

## What Synthea generates and who needs it

Synthea is a Synthetic Patient Population Simulator. The README states the goal plainly: output synthetic, realistic but not real patient data and associated health records in a variety of formats. That sentence is the whole product thesis. You are not anonymising real records or sampling from a production database. You are running a simulation in which each patient is born, accumulates conditions, medications, encounters, observations, procedures and care plans, and eventually dies.

The audience follows from that. Engineers building or testing FHIR servers need payloads with plausible structure and volume. Analysts testing cohort logic need populations large enough to produce edge cases. Educators and demo builders need records they can publish without a privacy review. Anyone whose test fixtures are currently a handful of hand-written JSON bundles is the target reader.

What Synthea is not is a data source with clinical authority. The output is generated from modular rules and configuration-based statistics, with Massachusetts Census data as the default demographic baseline per the README. The realism is bounded by what those rules encode.

## The modular rule system behind each simulated patient

The architecture visible from the repository is a Java engine plus a rule layer. The README describes a Modular Rule System with drop-in Generic Modules and the option of custom Java rules modules for additional capabilities. Modules are the disease and encounter logic; they are what decides that a patient with a given history develops a given condition and receives a given treatment.

The lifecycle runs birth to death, and encounters are typed as primary care, emergency room, or symptom-driven. That distinction matters for testing: a symptom-driven encounter is a different entry point into the record than a scheduled primary care visit, and downstream logic that assumes one shape will break on the other.

Output formats are the second half of the design. FHIR comes in R4, STU3 v3.0.1 and DSTU2 v1.0.2, which means the same engine can feed a modern FHIR R4 pipeline and a legacy DSTU2 consumer. Bulk FHIR is emitted as ndjson. C-CDA, CSV and CPCDS are also available. The README is explicit that by default Synthea does not generate CCDA, CPCDA, CSV or Bulk FHIR, so the format surface is opt-in rather than automatic.

There is a debugging affordance worth noting: rendering rules and disease modules with Graphviz, exposed as ./gradlew graphviz. When a module produces a record you did not expect, the graph is the fastest way to see which branch fired.

## Installing Synthea and generating a first population

The README splits its instructions in two. People who just want to run Synthea are pointed at the wiki page Basic Setup and Running. The Developer Quick Start is for those who want to examine the source, extend it, or build locally. The steps below are the developer path, because that is what the README documents in full.

System requirements first: Java JDK 17 or newer. The README strongly recommends an LTS release, specifically 17 or 25, and warns that issues may occur with more recent non-LTS versions. Clone the repository and run the Gradle build, check and test tasks together:

```bash
git clone https://github.com/synthetichealth/synthea.git
cd synthea
./gradlew build check test
```

That command compiles the project and runs the test suite. If it fails immediately, check the JDK version before anything else, since a non-LTS JDK is the failure mode the README calls out.

With the build working, generate patients. Running the script with no arguments produces the population one at a time:

```bash
./run_synthea
```

To control the run, pass a seed, a population size, and optionally a state and city. The usage line from the README is:

```bash
run_synthea [-s seed] [-p populationSize] [state [city]]
```

A concrete invocation with all three, taken from the README examples, is:

```bash
./run_synthea -s 987 Washington Seattle
```

The seed makes the run reproducible, which is the property that matters most if you plan to check generated fixtures into a repository. Output goes to ./output, and the README states Synthea will output patient records in C-CDA and FHIR formats there.

If you need a format that is off by default, either edit src/main/resources/synthea.properties or override a setting on the command line. The README gives this example for turning on FHIR export explicitly and redirecting the output directory:

```bash
./run_synthea -p 10 --exporter.fhir.export=true
./run_synthea --exporter.baseDirectory="./output_tx/" Texas
```

Bulk FHIR follows the same pattern: set exporter.fhir.bulk_data = true to activate it. C-CDA, CSV and CPCDS use exporter.ccda.export, exporter.csv.export and exporter.cpcds.export respectively. The README also points to a guided customizer tool at synthetichealth.github.io/spt/#/customizer if you would rather not hand-edit properties.

## Where Synthea stops being the right tool

The most important limitation is stated by the project itself: the data is realistic but not real. That is not a caveat you can engineer around. If your test depends on a specific real-world comorbidity distribution, a particular coding quirk from a specific EHR vendor, or the messy inconsistencies that appear in production extracts, Synthea will not reproduce them, because it generates from rules and configured statistics rather than from observed records.

Demographics are the second boundary. The README says configuration-based statistics and demographics default to Massachusetts Census data. A population generated without changing that configuration reflects a Massachusetts baseline. If your application is being validated against a different region, the default demographics are the wrong input, and the README does not document a turnkey path for substituting another census source in the text available here.

The third limitation is operational. This is a Java project you build and run, not a service you call. There is no hosted API in the README. Teams that want synthetic records delivered as an endpoint will need to run Synthea themselves and store the output, which adds a build and storage step to their pipeline.

Finally, format defaults are a trap for the impatient. A first run produces FHIR and C-CDA in ./output, and nothing else. If you expected CSV or ndjson and did not set the corresponding exporter property, the absence is a configuration problem, not a bug.

## Synthea compared with record-level de-identification

The obvious alternative approach is de-identification: take real records and strip or perturb the identifiers. The difference in method is fundamental. De-identification preserves the clinical distribution of the source data, because the clinical content came from real patients, and it inherits the source's gaps, coding conventions and outliers. Its risk is re-identification, and its cost is the review and tooling needed to make that risk defensible.

Synthea inverts both properties. There is no real patient behind any record, so the privacy question largely disappears, and the README's framing of synthetic data is the reason teams reach for it. In exchange, the clinical content is only as good as the modules and configuration. You trade distributional fidelity for safety and reproducibility.

A second alternative is hand-authored fixtures. Those are precise and small, and they are the right choice when you need to assert on one exact bundle. They do not scale to population-level questions, and they drift as the schema evolves. Synthea's value is the middle ground: thousands of records with internal consistency, generated on demand from a seed, in whichever of the supported formats your pipeline consumes.

The practical decision rule is about what you are testing. Schema conformance, pipeline throughput, cohort query logic and UI rendering are all well served by synthetic populations. Clinical decision support tuned to real prevalence is not.

## Maintenance, licence and the cost of upgrading

Synthea is not archived. The last push to the default branch was on 2026-08-18, which is recent enough that the repository is being touched, though the README does not describe a release cadence or a support policy. The release history shows v3.4.0 on 2025-11-05, v4.0.0 on 2026-03-05, and a master-branch-latest tag matching the 2026-08-18 push. The gap between v3.4.0 and v4.0.0 is a major version step, so anyone pinning to v3 should read the release notes before moving.

Upgrade cost concentrates in two places. First, the Java baseline: the README requires JDK 17 or newer and recommends LTS releases, so an environment stuck on an older JDK needs work before it can build current Synthea at all. Second, the module format. Because modules are the rule layer, a major version that changes module semantics can invalidate custom modules a team has written. The README documents drop-in Generic Modules and custom Java rules modules but does not document a module compatibility policy in the text available here.

On licensing, Synthea is Apache-2.0. The README carries the standard Apache text, copyright 2017-2025 The MITRE Corporation, and states the software is distributed on an AS IS BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND. Apache-2.0 is a permissive licence that generally allows commercial use and modification with attribution and notice retention, but the README does not address the status of generated patient data, and nothing here should be read as legal advice. If you plan to redistribute generated datasets, have counsel review the NOTICE file and the licence terms rather than assuming the output inherits the code's permissions.

## Conclusion

Adopt Synthea if you need reproducible patient populations for FHIR pipelines, analytics testing or demos, and you are comfortable running a Java build. Do not adopt it if you need production-grade clinical realism, an official hosted dataset, or a service with an SLA; the README points to the wiki and a guided customizer instead. Before committing, verify two things yourself: that your JDK is 17 or 25, since the README warns that non-LTS versions may cause issues, and which exporters you actually need, because C-CDA, CSV, CPCDS and Bulk FHIR are all off by default and must be turned on in src/main/resources/synthea.properties.

## FAQ

### What is Synthea data?

It is synthetic patient data produced by the Synthea simulator: records for simulated patients that run from birth to death, including conditions, medications, encounters, observations, procedures and care plans. The README describes the output as realistic but not real, and it can be exported in FHIR, C-CDA, CSV and CPCDS formats.

### Is Synthea free to use?

The repository is licensed under Apache-2.0, copyright 2017-2025 The MITRE Corporation, with the standard Apache terms in the README. The README does not state a separate fee or paid tier for the software.

### What programming language is Synthea written in?

Synthea is a Java project, listed with Java as its primary language. The README requires Java JDK 17 or newer and recommends an LTS release such as 17 or 25, and it also supports custom Java rules modules for additional capabilities.

### Is Synthea HIPAA compliant?

The README does not make a compliance claim, and it does not discuss HIPAA. What it does state is that the generated data is synthetic and not real, which is the property teams usually cite when using it in place of real records. Compliance determinations are not something the available documentation addresses.

### How do you install and run Synthea?

The README's Developer Quick Start says to clone the repository and run ./gradlew build check test with JDK 17 or newer installed. After that, ./run_synthea generates a population, and arguments such as -s seed, -p populationSize and a state and city control the run. Output lands in ./output.

### What can I use instead of Synthea?

The README does not name alternatives. The closest different approach is de-identifying real records, which preserves the source clinical distribution but carries re-identification risk, whereas Synthea generates records from modular rules and configuration-based statistics so no real patient is involved.

## Sources

- [License: Apache-2.0](https://github.com/synthetichealth/synthea/blob/master/LICENSE)
- [Project website](https://synthetichealth.github.io/synthea)
- [README](https://github.com/synthetichealth/synthea/blob/master/README.md)
- [Releases](https://github.com/synthetichealth/synthea/releases)
- [synthetichealth/synthea on GitHub](https://github.com/synthetichealth/synthea)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/synthetichealth-synthea
