# Hugging Face dataset-viewer: The Backend That Powers Dataset Pages on the Hub

> The Hugging Face dataset-viewer is the Apache-2.0 backend service behind every dataset page on the Hugging Face Hub, serving pre-computed table data, search results, and statistics through a public API. The frontend viewer is not open-source; only this backend is.

**huggingface/dataset-viewer** — Backend that powers the dataset viewer on Hugging Face dataset pages through a public API.

- Repository: https://github.com/huggingface/dataset-viewer
- Website: https://huggingface.co/docs/dataset-viewer
- Stars: 902 · Forks: 133
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/huggingface-dataset-viewer

## What dataset-viewer does and why it exists

When you open a dataset page on the Hugging Face Hub, the page displays a table of the first rows, column types, basic statistics, and a search interface. Computing these things on the fly for every page load would be slow and expensive: some datasets contain billions of rows and require format conversion and statistical analysis before any data can be shown.

dataset-viewer solves this by running a background processing pipeline that ingests each dataset on the Hub, pre-computes the rows, statistics, and search index, and stores the results. The public API at https://huggingface.co/docs/dataset-viewer serves these pre-computed results to any client, including the Hub's own frontend and third-party tools.

The README is explicit about scope: the frontend viewer component is not part of this repository and is not open-source. This repository is the backend only. Teams who want to add dataset-viewer functionality to their own infrastructure can run this service, which exposes the same API the Hub uses.

## Service architecture: workers, cache, queue, and specialized services

The system is not a single application. The docker-compose.yml and Makefile expose at least seven services with distinct roles, running on separate ports:

The API service (port 8180) serves the public-facing endpoints. The rows service (port 8182) handles row-level data retrieval. The search service (port 8183) powers full-text search over dataset contents. The SSE API service (port 8185) handles server-sent events for live updates. The worker service (port 8186) processes background jobs from the queue. The webhook service (port 8187) receives events from the Hub when datasets change. An admin service (port 8181) handles internal operations. A reverse proxy on port 8100 routes external traffic to the appropriate service.

All services share MongoDB for both caching processed results and maintaining the job queue. The docker-compose.yml uses CACHE_MONGO_URL and QUEUE_MONGO_URL environment variables, which can point to the same MongoDB instance or separate ones. File assets (images, audio clips extracted from datasets) are stored either on local disk or in S3, configured through ASSETS_STORAGE_PROTOCOL.

## Running dataset-viewer locally with Docker Compose

The Makefile provides the primary entry points for local development. Starting the full stack:

```bash
make start
```

This runs docker-compose with the .env file, building and starting all services with --force-recreate and waiting up to 20 seconds for health checks to pass.

For development with debug settings:

```bash
make dev-start
```

This adds the .env.debug overrides on top of .env. Stopping the stack:

```bash
make stop
```

This removes containers, networks, and volumes. The ASSETS_BASE_URL and CACHED_ASSETS_BASE_URL environment variables are required and have no defaults; the README's .env section notes they must be set to values matching the reverse proxy configuration, for example http://localhost:8100/assets when running locally.

The e2e test suite runs separately through the e2e/ directory. The individual service Makefiles under libs/ and services/ handle per-service installation, quality checks, and tests.

## The Rust component: libviewer and why it exists

One component in the stack is written in Rust rather than Python. The Dockerfile builds a libviewer component in a separate stage before the main Python services. It uses maturin, a build tool for Rust-backed Python extensions, to compile the Rust code into a Python wheel.

The build stage installs the Rust toolchain, compiles libviewer from libs/libviewer, and outputs the wheel to /tmp/dist. The Python services then install this wheel as part of their own dependency setup.

The README does not explain exactly which tasks libviewer handles, but the build structure indicates it is a performance-sensitive component that the Python service layers call into. Having a Rust component means the build process requires Rust toolchain installation. The Dockerfile handles this automatically inside the Docker build, but local non-Docker builds need the Rust toolchain installed before running any Python install steps for services that depend on libviewer.

## The public API: what it exposes and how to use it

The public-facing API documentation is at https://huggingface.co/docs/dataset-viewer. The API exposes endpoints for retrieving rows from a dataset split, getting dataset configuration information, running full-text search within a dataset, and accessing computed statistics.

This API is what the dataset pages on the Hub call, and it is what third-party integrations should use for reading dataset contents. It is not necessary to run the self-hosted backend to use the API; the Hub operates its own instance of this service.

The LeRobot project, which was listed in the related searches for this repository, uses the dataset-viewer API format for its robotics datasets, which shows up in the context of tools like Stardustai's dataset viewer that are built to consume the same API endpoints.

## Limitations: frontend is closed, MongoDB required, no standalone mode

The README states the frontend is not open-source. This means self-hosting the backend gives you the API and the processing pipeline, but not the table UI that Hub users see. Building a UI on top of the API requires separate frontend work.

MongoDB is a hard dependency for both cache and queue. There is no SQLite or in-memory mode for development or small-scale deployments. Running without a MongoDB instance requires modifying the codebase.

The Python version in the Dockerfile is 3.14.5, which is a recent release. If your organization has constraints on Python versions, verify compatibility before deploying.

The latest GitHub release is 0.21.0, dated 2023-02-14, but the last push to the repository was on 2026-09-25. The project is under active development on the main branch rather than through tagged releases. The gap between the last release and the last push is significant; contributors and operators should track the main branch rather than the release tags.

## Alternative: Hugging Face datasets library and the public Hub API

For the common use case of reading data from a Hugging Face dataset, the datasets Python library is the right tool. It handles downloading, caching, and iterating over dataset contents without requiring any server infrastructure. Installation is a single pip command and reading a dataset is a few lines of Python.

dataset-viewer is the wrong starting point for that use case. It is a server-side preprocessing and serving system designed to run continuously as a background service, not a library for reading datasets in a script. The difference in scope is fundamental: one is an infrastructure component that processes millions of Hub datasets for web display; the other is a client library for loading data into a training or analysis pipeline.

Organizations that want to run their own dataset catalog with the same viewer capabilities as the Hub need the dataset-viewer backend. Everyone else should use the datasets library or the public API directly.

## License and contribution model

The repository is released under the Apache-2.0 license. Contributing code or documentation involves following the CONTRIBUTING.md guide, which links to a DEVELOPER_GUIDE.md for setting up the backend locally. Bug reports for the public Hub viewer should be opened as discussions on the dataset's page on the Hub rather than as GitHub issues; the README says this is the most efficient path for fixing issues visible on specific dataset pages. The last push was on 2026-09-25, and the repository is not archived.

## Conclusion

Organizations that need to self-host the dataset viewer backend to power their own dataset catalog, or that want to contribute to the service that runs on the Hugging Face Hub, will find this repository directly relevant. It is not appropriate as a starting point for reading individual datasets programmatically; the Hugging Face datasets Python library or the public API documented at https://huggingface.co/docs/dataset-viewer are the right tools for that. Before self-hosting, review the architecture diagram and the DEVELOPER_GUIDE.md to understand the service count and MongoDB dependency.

## FAQ

### Is the Hugging Face dataset viewer frontend open source?

No. The README states that the frontend viewer component is not part of this repository and is not open-source. Only the backend API service is open-source under Apache-2.0.

### What database does the dataset-viewer backend require?

MongoDB is required for both the cache and the job queue. The docker-compose.yml uses separate CACHE_MONGO_URL and QUEUE_MONGO_URL environment variables, which can point to the same MongoDB instance or different ones.

### How do I report a bug in what I see on a Hugging Face dataset page?

The README recommends opening a discussion on the dataset's page on the Hub and tagging @lhoestq in the discussion. This is described as the most efficient path for fixing issues on specific datasets. GitHub issues are for reporting bugs in the backend service itself.

## Sources

- [huggingface/dataset-viewer on GitHub](https://github.com/huggingface/dataset-viewer)
- [License: Apache-2.0](https://github.com/huggingface/dataset-viewer/blob/main/LICENSE)
- [Project website](https://huggingface.co/docs/dataset-viewer)
- [README](https://github.com/huggingface/dataset-viewer/blob/main/README.md)
- [Releases](https://github.com/huggingface/dataset-viewer/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huggingface-dataset-viewer
