# cartography-cncf/cartography: Infrastructure Assets and Relationships in a Neo4j Graph

> Cartography is a Python tool that syncs AWS, GCP, Azure, Kubernetes, GitHub, Okta and other platforms into Neo4j as nodes and edges. Its value is in the joins, and its cost is in the graph you have to operate.

**cartography-cncf/cartography** — Cartography is a Python tool that pulls infrastructure assets and their relationships into a Neo4j graph database.

- Repository: https://github.com/cartography-cncf/cartography
- Website: https://docs.cartography.dev/
- Stars: 4,105 · Forks: 577
- Language: Python
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/cartography-cncf-cartography

## The cross-provider question Cartography is built to answer

Most cloud inventory tools answer questions inside one provider's boundary. Cartography's README frames the problem differently, listing questions such as which identities have access to which datastores across multiple tenants or providers, which compute instances are exposed to the internet, and what network paths run in and out of an environment. Those are joins. A single AWS account view can tell you an S3 bucket exists; it cannot tell you that a GitHub team, an Okta group and an Entra ID application all resolve to the same human who can read that bucket.

The tool is aimed at security and infrastructure engineers who already run several platforms and want one graph instead of several consoles. The supported platform list in the README runs past thirty entries, including AWS, GCP, Azure, Kubernetes, GitHub, Okta, Entra ID, CrowdStrike, Google Workspace, Keycloak, Cloudflare, Duo, Kandji, Huntress and CVE metadata from NVD and FIRST.org. A second audience is the detection engineering team: the repository ships a cartography-rules command that runs named security rules against the populated graph, and the README shows listing and running a rule called object_storage_public.

It is not an agent, not a scanner and not a remediation system. Cartography reads APIs and writes nodes and relationships. Nothing in the README suggests it changes state in the source platforms.

## How the sync works: modules, intel jobs and a Bolt write to Neo4j

The architecture visible in the repository is a Python package with per-platform modules. You select modules on the command line with --selected-modules, and each module is responsible for calling that platform's API and emitting nodes and relationships. The README's AWS example covers a long list of services, from EC2, ECS, ECR and EKS to IAM, KMS, RDS, S3, Lambda, GuardDuty, Inspector, Security Hub, Secrets Manager and Bedrock. The graph is not a flat dump: resources hang off account and organization nodes, and relationships such as RESOURCE connect an AWSAccount to the things inside it, which is what makes the Cypher examples in the README work.

Writes go to Neo4j over Bolt. The neo4j Python driver is a declared dependency with a floor of 6.0.0, and the project offers an extra, cartography[neo4j-rust], that swaps in Neo4j's Rust Bolt codec. The README states this cuts sync time by roughly 20-30%, and the Dockerfile repeats that figure. The Dockerfile also explains why it is an extra rather than a hard dependency: a platform without a pre-built wheel would need a Rust toolchain to build it, while the linux/amd64 and linux/arm64 images the Dockerfile targets are covered by manylinux wheels.

There is a second class of job beyond the platform modules. The README lists CVE Metadata as a module that enriches the graph with CVSS, EPSS scores and CISA KEV data from NVD and FIRST.org, and notes that the older NIST CVE module is deprecated in its favour. That enrichment is what turns a graph of assets into something you can filter by exploitability rather than by presence. The AIBOM module goes the other direction, linking AI component detections to ECR images.

## Installing Cartography and running a first AWS sync

The README's quick start is three steps: install the package, start Neo4j, run a sync. Cartography requires Python 3.11 or newer according to pyproject.toml, and it is published on PyPI under the name cartography.

```bash
pip install cartography
```

The README notes that installing cartography[neo4j-rust] instead swaps in Neo4j's Rust Bolt codec for a stated 20-30% reduction in sync time. The next step is a database. The quick start uses the Neo4j 5 community container with authentication disabled and both the HTTP and Bolt ports published.

```bash
docker run -d --publish=7474:7474 --publish=7687:7687 -v data:/data --env=NEO4J_AUTH=none neo4j:5-community
```

Confirm that http://localhost:7474 responds before going further. Then make sure your AWS credentials and default region are available, for example through AWS_PROFILE, AWS_DEFAULT_REGION, or ~/.aws/config, and run the AWS module against the local Bolt endpoint.

```bash
cartography --neo4j-uri bolt://localhost:7687 --selected-modules aws
```

When the run finishes, open http://localhost:7474 and query the graph. The README's first example finds unencrypted RDS instances by account.

```cypher
MATCH (a:AWSAccount)-[:RESOURCE]->(rds:AWSRDSInstance{storage_encrypted:false})
RETURN a.name, rds.id
```

The property names matter: storage_encrypted, exposed_internet, instanceid and publicdnsname appear in the README's examples, and the data schema page is the reference for the rest. If you would rather start from a rule than a query, the repository also installs a cartography-rules command. Against the no-auth container above, no password is needed.

```bash
cartography-rules list
cartography-rules list object_storage_public
cartography-rules run object_storage_public
```

For an authenticated Neo4j, the README says to set NEO4J_PASSWORD or use one of the other password options in the rules documentation.

## The Neo4j instance is the operational cost centre

Cartography's own footprint is small; the graph database is not. The repository's docker-compose.yml sets NEO4J_server_memory_pagecache_size and both heap settings to 1G, with a comment saying to raise memory limits. That is a development default. A graph covering several cloud accounts, their IAM policies, container images and CVEs will outgrow it, and page cache is the setting that decides whether relationship traversals stay in memory.

The compose file also enables the APOC and Graph Data Science plugins, sets NEO4J_PLUGINS to ["graph-data-science", "apoc"], and allowlists gds.* and apoc.* procedures. Those plugins are what make some of the more interesting graph analysis possible, but they are a dependency you carry. The healthcheck polls http://localhost:7474 with wget every ten seconds, which is a reasonable signal for a local stack and not a substitute for monitoring a production instance.

The compose file is explicit that it is a starting point. Its header comment says it is intended to help you quick-start or develop Cartography, that it is a good starting point for your own customizations, and that upstream changes to it are hard to make support as many users as possible. Treat it as an example, not as a deployment topology. The same applies to NEO4J_AUTH=none: fine on a laptop, wrong anywhere reachable.

## Where Cartography is the wrong tool

Cartography is a batch sync into a graph, not a real-time control plane. If you need an alert within seconds of a bucket becoming public, a graph that is refreshed on your schedule is the wrong shape for that job. The README does not document incremental sync semantics, conflict handling or rollback, so you should not assume a failed run leaves the graph in a state you can reason about without checking.

Coverage is uneven by design. The platform list is long, but the depth varies enormously: AWS gets dozens of services while DigitalOcean and Jumpcloud appear as entries without a service breakdown in the README. If your estate is mostly on a platform that gets a thin module, you will be writing the module yourself, and the repository's CONTRIBUTING.md and AGENTS.md are where that conversation starts.

There is also a maturity signal worth reading literally. pyproject.toml classifies the project as Development Status :: 4 - Beta. That is the project's own label, not a criticism from outside. Combined with a version line still in the 0.14x range, it means schema and module behaviour can move between releases.

Finally, credentials. Cartography needs read access to everything it inventories, which for AWS means a broad IAM policy. The README points at the AWS credentials documentation for configuration but does not prescribe a least-privilege policy inline, so scoping that role is your problem, not the tool's.

## Alternatives and the difference in approach

The closest comparison is a cloud security posture management product, commercial or open source, that evaluates resources against a rule set and reports findings. Those tools typically own both the collection and the verdict: they ship the checks, they rank the findings, and they hand you a queue. Cartography owns neither. It gives you the graph and a rules runner, and the verdict is a Cypher query you write. The trade is explicit: more assembly work, but you can ask a question the vendor never anticipated, such as joining an Okta group to an ECR image to a CVE.

A second alternative is a general graph or asset inventory you already run, where you would write your own collectors. That keeps one store, but you inherit the maintenance of every API integration, including the ones Cartography already covers, such as ECR image layers and attestations or Entra ID federation to AWS Identity Center.

A third is the cloud provider's own inventory and query service, which is cheap to start and accurate within its boundary. It will not join across providers or tenants, which is the exact case the README leads with. If your questions never cross a provider boundary, that is the cheaper answer and Cartography is overhead.

## Licence, maintenance and what upgrades cost you

Cartography is licensed under Apache-2.0, with a LICENSE file at the repository root and license-files declared in pyproject.toml. Apache-2.0 permits commercial use and modification and includes an explicit patent grant; it also requires that you preserve notices and state significant changes. That is a description of the licence text, not legal advice, and if you redistribute a modified Cartography inside a product, your counsel should read the notice and modification clauses rather than this paragraph.

The maintenance picture is active. The repository is not archived, and the last push was on 2026-08-11, the same day as the 0.140.0 release. The two releases before it, 0.139.1 and 0.139.0, landed on 2026-07-20 and 2026-07-13. The project also carries governance and maintainer files, GOVERNANCE.md and MAINTAINERS.md, and is published under the CNCF organisation, which is a signal about stewardship rather than about release cadence.

Upgrade cost is mostly schema drift. Dependencies use loose lower bounds by convention, with exact pinning handled by uv.lock, so installing from PyPI resolves versions you did not choose. If a Cypher query in your runbooks depends on a node label or property, a minor release can break it. The cheapest defence is to pin the version you run and read the release notes before moving, then re-run your saved rules with cartography-rules run to see whether the graph still answers them. The deprecation of the NIST CVE module in favour of CVE Metadata is a concrete example of a module you may need to migrate.

## Conclusion

Adopt Cartography if your questions are cross-platform joins, such as which identities reach which datastores or which compute is exposed to the internet, and you can operate a Neo4j instance with enough page cache for your graph. Do not adopt it if you need a hosted product, if you want per-resource remediation rather than a graph, or if your only source is one cloud whose native inventory already answers your questions. Before committing, verify that the modules you need exist under docs/modules, that your credentials can be read-only, and that the Neo4j memory settings in docker-compose.yml match the size of the graph you intend to sync.

## FAQ

### What is cartography-cncf/cartography?

It is a Python tool that pulls infrastructure assets and their relationships into a Neo4j graph database. The README lists AWS, GCP, Azure, Kubernetes, GitHub, Okta, Entra ID, CrowdStrike and more than thirty other platforms as sources.

### How do I install Cartography?

Install it from PyPI with pip install cartography, or pip install cartography[neo4j-rust] to swap in Neo4j's Rust Bolt codec, which the README says cuts sync time by roughly 20-30%. It requires Python 3.11 or newer.

### How do I run a first sync against AWS?

Start a Neo4j 5 community container with the Bolt port published, configure your AWS credentials and default region, then run cartography --neo4j-uri bolt://localhost:7687 --selected-modules aws. Query the result at http://localhost:7474.

## Sources

- [Official documentation](https://docs.cartography.dev/)
- [Official README](https://github.com/cartography-cncf/cartography#readme)
- [Project repository](https://github.com/cartography-cncf/cartography)
- [Release notes](https://github.com/cartography-cncf/cartography/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cartography-cncf-cartography
