# duckdb-wasm: running DuckDB OLAP queries inside the browser

> duckdb-wasm compiles the DuckDB OLAP engine to WebAssembly so SQL over Parquet, CSV and JSON runs client side. It is genuinely useful for in-browser analytics, but single-threaded by default, sandboxed against out-of-core work, and dependent on CORS for every remote file.

**duckdb/duckdb-wasm** — WebAssembly version of DuckDB

- Repository: https://github.com/duckdb/duckdb-wasm
- Website: https://shell.duckdb.org
- Stars: 2,131 · Forks: 230
- Language: C++
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/duckdb-duckdb-wasm

## An OLAP engine inside the page

DuckDB is an in-process SQL OLAP database management system, meaning no separate server process sits between your code and the data. DuckDB-Wasm takes that engine and compiles it to WebAssembly so the same query engine runs inside a browser tab or inside Node.js, with no server to provision. The README states the package has been tested with Chrome, Firefox, Safari and Node.js, and the project points to a VLDB publication and a recorded talk for the design rationale.

The audience is a specific one: teams that want users to run real aggregation over files in the client, rather than shipping a pre-computed answer set. The data stays where the browser can reach it. The README describes the engine as speaking Arrow fluently and reading Parquet, CSV and JSON files, with those files backed either by Filesystem APIs or by HTTP requests.

The gap it is filling is that JavaScript alone has no SQL engine for columnar data, and a remote query service adds a network round trip and a deployment surface. Putting the engine in the client removes both, at the cost of download weight and a sandbox, which is where the interesting constraints live.

## What changes between the native build and the Wasm build

The README keeps a list of relevant differences, and it is worth reading as a list of trade-offs rather than a bug list.

The default HTTP stack is different between native and Wasm builds. `LOAD httpfs` opts into using the same HTTP logic as the native version, but re-implemented in JavaScript. Without it, the Wasm build has its own default path, and requests are always upgraded to HTTPS. Any resource you fetch also needs the server to allow Cross Origin access, so a third-party endpoint that serves data without CORS headers will fail in the browser for a reason that has nothing to do with SQL.

Extension handling is lazy. `INSTALL extension_name FROM 'https://repository.endpoint.org';` only registers the extension, and the fetch is deferred to the first `LOAD extension_name;`. `INSTALL x FROM community;` is supported as a shorthand.

Two constraints shape performance rather than correctness. The Wasm builds are optimized for download speed, so core extensions that are usually bundled into native DuckDB binaries, including autocomplete, JSON, Parquet and ICU, are autoloaded at runtime in the browser, which means fetching them. And DuckDB-Wasm is sandboxed, so it may not have the same level of support for out-of-core operations and filesystem access. The default mode is single threaded, with multithreading still experimental.

## Reading a DuckDB file over HTTP with ATTACH

The simplest demonstration in the README is that a DuckDB database file made available by URL can be attached and queried as if it were local. Databases written by DuckDB are compatible to be read from DuckDB-Wasm.

```sql
ATTACH 'https://blobs.duckdb.org/data/test.db'; FROM db.t;
```

That single statement is the whole mechanism. No copy step, no import wizard, no schema translation. The file stays on whatever host serves it, and the browser fetches what the query touches. The same pattern reaches Parquet files directly, which is how the spatial demo in the README creates a table from a remote file and then joins it against itself to rank the closest station pairs by aerial distance.

The catch is the one from the previous section. The engine may not reach outside the browser sandbox for out-of-core work, so a query that streams a large scan from an HTTP source can hit limits that a local file would not. The README does not quantify those limits, which makes this the area where you would want to measure against your own file sizes before committing.

## Extensions are downloaded on demand, at a real byte cost

DuckDB delegates functionality to extensions, and DuckDB-Wasm keeps that model. Core extensions live at `https://extensions.duckdb.org` and community extensions at `https://community-extensions.duckdb.org`.

```sql
LOAD icu;

DESCRIBE FROM read_parquet('https://blobs.duckdb.org/stations.parquet');

INSTALL h3 FROM community;
INSTALL sqlite_scanner FROM 'https://extensions.duckdb.org';
LOAD h3;
LOAD sqlite_scanner;
```

There are three ways to get an extension in, and they behave differently. `LOAD` fetches and activates in one step. `INSTALL` registers a source without fetching, deferring the download until a later `LOAD`. And autoloading happens implicitly when you touch a function or setting that needs an extension, which is what the `DESCRIBE FROM read_parquet(...)` line above triggers.

The README puts a number on this. Trying the extension demo requires about 3.2 MB of compressed Wasm files transferred over the network on first visit, and caching might help. Extension sizes vary with the functionality provided and the toolchain used.

One caveat is documented plainly: for ICU, autoloading does not work correctly in all cases, and an explicit `LOAD icu;` might be needed to reproduce the same behaviour as native. That is the kind of detail that turns into a debugging session if you skip it. Checking what is loaded takes one query, `FROM duckdb_extensions() WHERE loaded;`, which the README says lists h3, icu, parquet, quack and sqlite_scanner after the demo.

## Building the engine from source

Building DuckDB-Wasm is not a package install. The engine source lives in a git submodule, and the build needs Emscripten, which the Makefile will fall back to a container for. The README gives five commands.

```shell
git clone https://github.com/duckdb/duckdb-wasm.git
cd duckdb-wasm
git submodule init
git submodule update
make apply_patches
make serve
```

The submodule steps matter. `submodules/duckdb` holds the engine itself, and the Makefile derives a hash from it with `git reflog -n 1 | head -c 10`, so a shallow or stale submodule produces a different build identity. `make apply_patches` applies the compatibility patches before compiling.

You only need this path if you want to change the engine or the extension set. For using the library, the npm package is `@duckdb/duckdb-wasm`, and the README links a separate examples repository plus API documentation hosted at shell.duckdb.org.

The top-level layout explains the shape of the project. `lib/` holds the C++ core, `packages/` holds the TypeScript packages and the shell, `tools/` and `patches/` hold build tooling, and `examples/` holds five runnable targets: bare-browser, bare-node, esbuild-browser, esbuild-node and plain-html. The workspace Cargo.toml lists only `tools/dataprep` and `packages/duckdb-wasm-shell/crate`, which is a reminder that most of the Rust here is the shell, not the engine.

## The CI container, ccache and the emscripten cache

Compiling DuckDB to WebAssembly is slow, so the repository treats caching as part of the build rather than an afterthought. The compose file mounts three paths into the build container.

```yaml
services:
  duckdb-wasm-ci:
    image: duckdb/wasm-ci:0.75
    volumes:
      - .:/wd
      - ./.ccache:/mnt/ccache
      - ./.emscripten_cache:/mnt/emscripten_cache
    environment:
      - CCACHE_BASEDIR=/wd/lib
      - CCACHE_DIR=/mnt/ccache
      - EM_CACHE=/mnt/emscripten_cache/
```

`ccache` keeps compiled C++ objects across rebuilds, and `EM_CACHE` keeps the Emscripten sysroot so the toolchain is not downloaded again. The `CCACHE_BASEDIR` setting is there because the cache keys contain absolute paths, and the mount point changes between machines; without it a cache written on one host misses on another.

The Makefile picks the environment for you. `set_environment` checks for `emcc` on the path and uses a native build if it finds one, otherwise it echoes `docker compose run duckdb-wasm-ci` so subsequent targets run in the container. The default goal is `app`, and the `build/data` target generates TPC-H tables at several scale factors, which is the fixture set the benchmark packages consume.

## The version story is confusing, and the README admits it

The README states that DuckDB-Wasm is currently based on DuckDB v1.5.4. The listed releases for this repository are v1.31.0 on 2025-09-26, then v1.32.0 and v1.33.0 both on 2025-12-16, the second roughly two hours after the first. Those numbers do not line up with the base version in the README, and neither the README nor the release notes explain the relationship between them. Treat the pinned DuckDB submodule as the authoritative engine version and the README sentence as stale documentation rather than a specification.

The last push to the repository was on 2026-07-28, and the project is not archived. The licence is MIT. A GitHub Actions badge points at `workflows/main.yml`, and the README also displays the latest DuckDB upstream release badge, so the two version tracks are surfaced side by side without being reconciled.

There is no migration or rollback documentation. If you pin the npm package version, nothing in the repository describes how to move between DuckDB-Wasm versions or what to do when an autoloaded extension changes shape underneath you.

## Where DuckDB-Wasm is the wrong tool

Four constraints decide this, and all four are stated in the README rather than inferred.

Out-of-core work is the first. The engine is sandboxed and may not have the same level of support for out-of-core operations and filesystem access, so a query designed to spill to disk on a large local scan may behave differently. The second is threading. The default is single threaded and multithreading is experimental, so a parallel analytical workload sees one core unless you turn on something the project still calls experimental. The third is the extension payload. Because builds are optimized for download speed, extensions that would be bundled natively are fetched at runtime, and the extension demo moves about 3.2 MB on first visit before any query runs.

The fourth is the network shape. Requests are always upgraded to HTTPS and every resource needs Cross Origin access from the server, which makes arbitrary third-party URLs a fragile input. The `examples/plain-html` target and the hosted shell at shell.duckdb.org exist precisely so you can see this working before committing to it, and they are worth running against your own data volume.

## Conclusion

Adopt duckdb-wasm when queries over remote or local columnar files belong in the client and the download cost is acceptable, and reject it when you need parallelism or out-of-core execution. Verify first that your endpoint serves CORS headers and that a representative scan survives the sandbox. Pin the npm package version rather than trusting the base DuckDB version stated in the README, and issue an explicit LOAD icu; if ICU-dependent results differ from native DuckDB.

## FAQ

### How do I install DuckDB-Wasm in a web app?

Add the @duckdb/duckdb-wasm package from npm; the README also links a jsdelivr distribution and five runnable example targets, from plain-html through esbuild to bare-node. The repository's build-from-source path is for changing the engine, not for consuming it.

### Is DuckDB-Wasm multithreaded?

No. The README states the default mode is single threaded and that multithreading works but is still experimental and not enabled by default.

### How do I load a DuckDB extension in the browser?

Use LOAD extension_name for an explicit fetch and load, INSTALL ... FROM 'https://extensions.duckdb.org' or FROM community to register a source without fetching, or rely on autoloading when you touch a function that needs one. The README warns that ICU autoloading does not work correctly in all cases and an explicit LOAD icu; may be needed.

### How large is the DuckDB-Wasm download?

The README gives about 3.2 MB of compressed Wasm files for the extension loading demo on first visit, and notes that caching might help. Extension sizes vary with the functionality provided and the toolchain used.

### Can I read a Parquet or CSV file directly from a URL?

Yes. The engine reads Parquet, CSV and JSON files backed by Filesystem APIs or HTTP requests, and a remote DuckDB database file can be attached and queried with ATTACH 'https://blobs.duckdb.org/data/test.db'; FROM db.t;. Note that requests are always upgraded to HTTPS and the server must allow Cross Origin access.

### What licence is DuckDB-Wasm under?

MIT. That is a permissive licence with no copyleft obligation on code that links against it, though this is not legal advice and the LICENSE file in the repository is the authoritative text.

## Sources

- [duckdb/duckdb-wasm on GitHub](https://github.com/duckdb/duckdb-wasm)
- [License: MIT](https://github.com/duckdb/duckdb-wasm/blob/main/LICENSE)
- [Project website](https://shell.duckdb.org)
- [README](https://github.com/duckdb/duckdb-wasm/blob/main/README.md)
- [Releases](https://github.com/duckdb/duckdb-wasm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/duckdb-duckdb-wasm
