# Embedding Atlas: interactive visualization for large embeddings

> Apple's Embedding Atlas renders millions of points in the browser with WebGPU, clusters and labels them automatically, and links charts to an embedding view. Here is what the repository documents, and where it stops.

**apple/embedding-atlas** — Embedding Atlas is a tool that provides interactive visualizations for large embeddings. It allows you to visualize, cross-filter, and search embeddings and metadata.

- Repository: https://github.com/apple/embedding-atlas
- Website: https://apple.github.io/embedding-atlas/
- Stars: 4,960 · Forks: 327
- Language: TypeScript
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/apple-embedding-atlas

## What Embedding Atlas is for, and who ends up using it

The project targets a specific moment in a machine learning workflow: you have computed embeddings for a corpus, you have metadata columns alongside them, and you want to see structure rather than read a nearest-neighbor table. The README frames it as a tool to "visualize, cross-filter, and search across your data," which puts it in the exploratory data analysis slot rather than the production serving slot. The people who benefit are researchers and engineers inspecting a projection after UMAP or similar dimensionality reduction, and analysts who received a table of vectors and want to know whether the clusters mean anything. It is not a vector database and it does not answer queries at serving time. The audience is also explicitly split: the README markets the tool twice, once "For embeddings" and once "For any tabular data," which is an honest description of a viewer that happens to have a strong embedding mode rather than an embeddings-only product.

## How the rendering and clustering pipeline is put together

The repository is a monorepo, and its shape explains the architecture better than the README does. The npm workspace lists packages/utils, packages/component, packages/viewer, packages/umap/umap-wasm, packages/density-clustering, packages/embedding-atlas, packages/examples, packages/backend and packages/docs. The Cargo workspace mirrors the compute-heavy parts in Rust: density_clustering, density_clustering_wasm, nndescent, umap and umap-wasm. So the browser-side viewer is TypeScript and Svelte (the root package.json depends on svelte 5), while clustering and the UMAP nearest-neighbor graph are compiled to WebAssembly from Rust. That split is the reason the README can claim a few million points: heavy work runs in wasm, and rendering goes through WebGPU. Two rendering details are named in the README and are worth noting because they are uncommon. Kernel density estimation with density contours distinguishes dense regions from outliers, and order-independent transparency keeps overlapping points readable instead of letting draw order decide which point wins. Automatic clustering and labeling is the third piece, backed by a separate paper cited in the README, "A Scalable Approach to Clustering Embedding Projections."

## Installing Embedding Atlas with pip and running a first dataset

The README's Python path is two commands. The first installs the package from PyPI, the second launches the viewer against a dataset path. Note that the README writes the argument as <your-dataset> without spelling out accepted formats, so check the documentation site before pointing it at a directory.

```bash
pip install embedding-atlas

embedding-atlas <your-dataset>
```

Inside a Jupyter notebook the same package exposes a widget. The README imports EmbeddingAtlasWidget from embedding_atlas.widget and passes a data frame; the viewer renders inline rather than opening a separate page.

```python
from embedding_atlas.widget import EmbeddingAtlasWidget

# Show the Embedding Atlas widget for your data frame:
EmbeddingAtlasWidget(df)
```

For a web application, the npm package exports EmbeddingAtlas and EmbeddingView, with entry points for React and Svelte. The README shows the import paths but not the props, so the overview page is the place to look for the component API.

```js
import { EmbeddingAtlas, EmbeddingView } from "embedding-atlas";

// or with React:
import { EmbeddingAtlas, EmbeddingView } from "embedding-atlas/react";

// or Svelte:
import { EmbeddingAtlas, EmbeddingView } from "embedding-atlas/svelte";
```

## Cross-filtering, chart specs and the MCP endpoint

The tabular half of the tool is what separates it from a scatter plot library. The README lists bar, line, bubble, count plot and eCDF charts, plus "a composable chart spec for building custom charts like heatmaps and average-line overlays." Charts can be configured to cross-filter, meaning a selection in one chart constrains the embedding view and the other charts at the same time. Multimodal viewers are built in for text, image, audio, numeric, categorical and time columns, so an image column renders as thumbnails rather than as a string of bytes. The newest surface is agent access: the README states that AI agents can query the schema, run SQL, create charts and capture screenshots via Model Context Protocol. That is a notable design choice. It means the viewer exposes a query interface rather than only a mouse interface, and it is the part of the project most likely to change shape between releases, since the release cadence is fast (v0.22.0 in July 2026, v0.23.0 and v0.24.0 in August 2026).

## Where Embedding Atlas stops being the right tool

The README ties the scale claim to a specific backend: "Up to a few million points, powered by WebGPU." That is a constraint, not a footnote. WebGPU support varies by browser and platform, and the README does not document a fallback renderer or state what happens on a machine without it. If your users are on locked-down browsers or thin clients, this is the wrong choice. The second limitation is the shape of the data. Embedding Atlas wants a table: vectors plus columns. If your embeddings live behind an API and you cannot materialize them into a data frame or file, there is nothing for the viewer to load. Third, this is a local or embedded viewer, not a shared service. The README describes a Python package, a notebook widget and npm components. It does not describe user accounts, permissions or a hosted multi-tenant deployment, so teams expecting a URL they can hand to a stakeholder will need to build that layer themselves. Finally, automatic clustering and labeling is a convenience that also hides a parameter choice; the README does not document how to tune the cluster count, so if the labels look wrong you are dependent on what the cited clustering paper and the documentation site expose.

## Embedding Atlas compared with Nomic Atlas and Streamlit-style dashboards

The most common comparison is Nomic Atlas, which appears in the related searches for this project. The difference in approach is where the computation lives. Nomic Atlas is a hosted platform: you send data to a service and it handles storage, projection and the web interface. Embedding Atlas runs the viewer in your own browser and notebook, with the clustering and UMAP code compiled locally to WebAssembly, and your data stays on your machine. That is a real trade-off in both directions. You avoid an upload step and a service dependency, and you take on the browser compatibility and memory limits yourself. The second comparison is a hand-built Streamlit or Dash dashboard, also visible in the search data. A Streamlit app gives you arbitrary Python on the server and any chart you can write, but it does not give you a WebGPU point cloud with density contours and order-independent transparency, and it does not give you automatic clustering with labels. Embedding Atlas gives you that rendering and the cross-filter wiring for free, and asks you to accept its chart vocabulary and its data-frame-shaped input.

## Licence, release cadence and the cost of upgrading

The repository is MIT licensed, and the Cargo workspace declares license = "MIT" for the Rust crates as well, so the same terms cover the compute code and the TypeScript packages. MIT imposes no copyleft obligation on your application, though the usual caveat applies: this is a description of the licence file, not legal advice, and you should read LICENSE and any third-party notices before shipping. On maintenance, the last push to the default branch was on 2026-09-15, and the three most recent releases are v0.22.0 on 2026-07-07, v0.23.0 on 2026-08-18 and v0.24.0 on 2026-08-20. Two releases in three days suggests active iteration, which cuts both ways for adopters. You get fixes quickly, and you also get a moving target: pin the pip and npm versions in your lockfiles rather than tracking latest, especially if you build on the MCP interface or the composable chart spec, which are the parts of the README most likely to shift. The development instructions live at the documentation site and in packages/docs/develop.md; the root package.json shows the test split across JavaScript, Python and Rust, so contributing means a multi-language toolchain.

## Conclusion

Adopt Embedding Atlas if you already have vectors plus a metadata table and want to inspect clusters, outliers and nearest neighbors without writing a dashboard. Skip it if you need a hosted service with accounts, or if your GPU and browser cannot run WebGPU, since the README ties the few-million-point figure to that backend. Before committing, run the pip package on your own dataset and confirm the widget renders inside your notebook environment, because the README shows the widget call but does not document memory limits or a fallback renderer.

## FAQ

### How do I install Embedding Atlas in Python?

Run pip install embedding-atlas, then launch the viewer with embedding-atlas <your-dataset>. The README gives these two commands as the command line path.

### Can I use Embedding Atlas inside a Jupyter notebook?

Yes. The README shows importing EmbeddingAtlasWidget from embedding_atlas.widget and calling it with a data frame to render the widget inline.

### How many points can Embedding Atlas render?

The README states up to a few million points, powered by WebGPU. It does not document a fallback renderer for environments without WebGPU support.

### Is Embedding Atlas the same as Nomic Atlas?

No. Nomic Atlas is a hosted platform, while Embedding Atlas ships as a Python package, a notebook widget and npm components that run the viewer locally with WebAssembly-based clustering and UMAP.

## Sources

- [apple/embedding-atlas on GitHub](https://github.com/apple/embedding-atlas)
- [License: MIT](https://github.com/apple/embedding-atlas/blob/main/LICENSE)
- [Project website](https://apple.github.io/embedding-atlas/)
- [README](https://github.com/apple/embedding-atlas/blob/main/README.md)
- [Releases](https://github.com/apple/embedding-atlas/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apple-embedding-atlas
