Datashader: rasterizing millions of points into a fixed-size grid
Quickly and accurately render even the largest data.
At a glance
- What is it?
- Datashader is a Python rasterization pipeline for plotting datasets too large for conventional scatter plots. It trades interactive per-point rendering for a projection, aggregation and transformation pipeline that produces a fixed-size image, and it is best judged on that trade-off.
- Who is it for?
- Adopt Datashader when a scatter plot of your full dataset either hangs a browser or hides density, and when you can accept a fixed-resolution aggregate image instead of individually addressable points. Do not adopt it if you need per-point hover, exact point identity or vector output, because the pipeline deliberately compresses records into bins.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The plotting problem Datashader was built to remove
Conventional plotting libraries draw marks. One record becomes one point, one line or one polygon on a canvas, and the renderer has to keep every mark addressable so it can be styled, selected or hovered. That model works until the record count outgrows the display. A scatter plot of a hundred million rows cannot put a hundred million distinguishable pixels on a screen, so the renderer either drops points, aliases them into solid blocks, or stops responding. The README frames Datashader's purpose directly: it exists to automate "creating meaningful representations of large amounts of data."
The intended user is someone with a large tabular or array dataset who wants to see its density, its spatial shape or its distribution rather than the identity of individual rows. Census blocks, GPS traces, taxi pickups and drop-offs, and point clouds are the kinds of inputs the repository's example images point at. If you need to click a point and read its fields, this is not the tool. If you need to know where the mass is, it is.
Projection, aggregation, transformation: the three stages
Datashader does not draw your data. It bins it. The README describes three stages, and the order matters because each stage shrinks the working set.
Projection assigns each record to zero or more bins of a nominal plotting grid, based on a specified glyph. A glyph is the shape a record contributes: a single point falls into one bin, a line segment falls into every bin it crosses, a polygon covers the bins it overlaps. This is why the same pipeline handles scatter, line and area data without a separate renderer per chart type.
Aggregation then computes reductions per bin, compressing the dataset into what the README calls a much smaller aggregate array. The output of this stage is not an image; it is an array. That distinction is the whole design. Because the intermediate result is a regular grid, its size is bounded by the grid shape you chose, not by the number of input records. Ten million rows and ten billion rows both produce an aggregate of the same dimensions.
Transformation processes those aggregates into an image. Colour mapping, shading and other adjustments happen here, after the data has already been reduced to grid scale. The practical consequence is that the expensive part of the pipeline scales with input size while the visual part scales with output resolution. Datashader can also be used as a pre-processing stage inside another plotting library, which is how it extends that library's reach to datasets it could not otherwise draw.
Installing Datashader and producing a first aggregate
Datashader supports Python 3.10 through 3.14 on Linux, Windows and Mac, and the README gives two installation routes. The conda route is the one the project recommends for performance, because conda packages are built against numerical libraries optimized for the target platform.
conda install datashaderIf you are not using conda, pip works as well. The README also notes a pyviz channel for the latest releases and a dev-labelled channel for pre-release versions, so a pinned production environment and a bleeding-edge environment are separate commands rather than conflicting advice.
pip install datashaderOnce installed, the package ships a command that fetches the example notebooks and their data. Running it creates a directory named datashader-examples in the current working directory.
datashader examples
cd datashader-examplesThe examples need extra dependencies beyond the core install. Inside an active conda environment, the README says to update from the fetched environment.yml; otherwise you create a fresh environment from it.
conda env update --file environment.ymlWhat you should see after these steps is a local directory of notebooks and datasets that run against the installed package. The examples are the practical entry point here, because the README itself does not walk through an API call; it points to the documentation site for API details and for papers and talks about the approach.
Where the pipeline breaks down
The aggregate is the product, and that is also the limitation. Because records are reduced into bins, the individual row is gone by the time an image exists. Hover tooltips that report a specific record's values, linked brushing that selects exact rows, and any interaction that depends on point identity cannot be served from the aggregate alone. You can rebuild those features by re-querying the source data for the selected region, but that is application work the pipeline does not do for you.
Resolution is a second constraint. The grid shape is chosen up front, so detail below one bin's footprint is not represented. Zooming in without recomputing the aggregation against a finer grid shows you the same coarse bins magnified, not more data. Datashader is therefore not a substitute for a zoomable vector renderer over a small dataset; it is a substitute for failing to render a large one.
There is also a dependency cost. numba is a hard dependency in pyproject.toml, which means your Python version has to be one numba supports, and the first execution of a code path pays compilation time before it runs at full speed. The README's preference for conda over pip is a hint that binary compatibility of the numerical stack is a real concern rather than a stylistic choice. On top of that, the core dependency list is broad: pandas, xarray, numpy, scipy, param, colorcet and others arrive with the install.
Datashader against matplotlib and plotly
The difference from matplotlib is architectural, not cosmetic. Matplotlib draws every artist you hand it, so a scatter call with tens of millions of points asks the renderer to place tens of millions of marks. Datashader never places marks at all; it bins records into a grid and produces an array whose size you control. That is why the two are not really competitors for the same job. Matplotlib is the right answer when the dataset fits and you want precise control over every element. Datashader is the answer when the dataset does not fit and you want density.
The comparison with plotly is about where the work happens. Plotly renders in a browser and ships data to the client, so the practical ceiling is set by what the browser can hold and draw. Datashader reduces the data before any image exists, so the client receives a fixed-size raster regardless of input size. The cost moves to the server or notebook kernel, and the client loses per-point fidelity. Datashader is also designed to sit underneath a higher-level library rather than replace it, so a stack that pairs it with a HoloViz tool is a supported arrangement rather than a workaround.
An alternative in the same space is to pre-aggregate with a database or dataframe groupby and plot the result. That works, but you own the binning, the coordinate handling and the glyph logic yourself, and you will reimplement the projection stage badly before you reimplement it well.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-23. Releases are frequent enough to plan around: v0.19.1 on 2026-05-19, v0.19.0 on 2026-03-20, and v0.18.2 on 2025-08-05 before that. The project classifies itself as Production/Stable, and the presence of a CHANGELOG.md, a ROADMAP.md and a pre-commit configuration at the repository root suggests a maintained release process rather than ad hoc tagging.
Upgrade cost is mostly the numerical stack, not Datashader's own API. Because numba, numpy, pandas and xarray are direct dependencies, a Python version bump or a major pandas release can move the whole environment at once. Pin the Datashader version and read CHANGELOG.md before moving, particularly if you rely on a specific glyph or reduction.
The licence is BSD-3-Clause, declared both in the repository's LICENSE.txt and in pyproject.toml under license-files. That is a permissive licence, which generally means you can use and redistribute the library with the copyright notice and disclaimer intact, including in commercial products. This is a description of what the repository states, not legal advice; if the licence terms matter to your organisation, have counsel read LICENSE.txt rather than a summary.
Editorial conclusion
Adopt Datashader when a scatter plot of your full dataset either hangs a browser or hides density, and when you can accept a fixed-resolution aggregate image instead of individually addressable points. Do not adopt it if you need per-point hover, exact point identity or vector output, because the pipeline deliberately compresses records into bins. Before committing, verify that numba supports your Python version, since Datashader lists numba as a hard dependency, and check the CHANGELOG and release tags for the version you pin.
Frequently asked questions
datashader vs matplotlib: what is the difference?
Matplotlib draws every mark you pass it, so a scatter plot of tens of millions of points asks the renderer to place that many marks. Datashader instead bins records into a grid and produces a fixed-size aggregate array, so the output size does not grow with the input record count.
datashader vs plotly: what is the difference?
Plotly renders in the browser and ships data to the client, so the practical ceiling is what the browser can hold. Datashader reduces the data before an image exists, so the client receives a fixed-size raster and the cost moves to the kernel or server, at the price of per-point fidelity.
What is the datashader alternative?
Datashader compresses records into bins, so the individual row is not available in the aggregate. Pre-aggregating with a database or dataframe groupby and plotting the result keeps point identity, but you own the binning, coordinate handling and glyph logic.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/holoviz-datashader)