# weld-project/weld: a Rust runtime that optimizes across data libraries

> Weld builds a lazy computation for a whole analytics workflow, then optimizes and evaluates it as one unit. It is a research runtime aimed at library authors, not a drop-in replacement for Pandas.

**weld-project/weld** — High-performance runtime for data analytics applications

- Repository: https://github.com/weld-project/weld
- Website: https://www.weld.rs
- Stars: 3,008 · Forks: 252
- Language: Rust
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/weld-project-weld

## The data movement problem Weld targets

Analytics code is assembled from pieces. One function reads a table, another filters it, a third groups it, a fourth trains or scores a model. Each piece can be fast on its own. The README's claim is that the combined workflow often runs an order of magnitude below hardware limits because data moves between the functions. Weld's answer is to stop treating each call as a finished computation. A library expresses its core operations in a common intermediate representation, and Weld lazily builds a computation for the entire workflow, optimizing and evaluating it only when a result is needed. The audience is therefore library and framework authors, plus researchers who want to test whether that fusion idea holds. It is not a tool you install to make an existing Pandas script faster by flipping a flag.

## How the lazy workflow is expressed and compiled

The repository is a Rust workspace with four members: weld, weld-capi, weld-repl and weld-hdrgen, while weld-python is excluded from the workspace. Building it produces two dynamically linked libraries, libweld and libweldrt, as .so files on Linux and .dylib files on macOS. The docs directory splits the surface into separate documents: language.md covers the syntax of the Weld IR, api.md covers the low-level C API, python.md covers the Python API, and tutorial.md walks through building a small vector library. The split tells you where the boundaries are. The IR is the layer where cross-library optimization happens; the C API and the Python bindings are the ways other languages hand work to it. Because weld-hdrgen is a workspace member, header generation for the C API is part of the build rather than a separate download. LLVM is the code generator underneath, which is why the build instructions spend most of their length on getting llvm-config onto the PATH before Rust ever runs.

## Installing Weld and running a first build

Weld needs the latest stable Rust and LLVM/Clang++ 6.0. Start by confirming the Rust toolchain, then install the matching LLVM. On Ubuntu 16.04 the README gives this sequence, which adds the apt.llvm.org repository and installs the development packages:

```bash
wget -O - https://apt.llvm.org/llvm-snapshot.gpg.key | sudo apt-key add -
sudo apt-add-repository "deb http://apt.llvm.org/xenial/ llvm-toolchain-xenial-6.0 main"
sudo apt-get update
sudo apt-get install llvm-6.0-dev clang-6.0
```

The dependencies look for llvm-config, so point it at the version 6 binary. The README notes that sudo may be required:

```bash
ln -s /usr/bin/llvm-config-6.0 /usr/local/bin/llvm-config
```

Run `llvm-config --version` and you should see 6.0.x or newer. zlib is also required, installed as zlib1g-dev. On macOS the equivalent is `brew install llvm@6` followed by a symlink from `brew --prefix llvm@6`/bin/llvm-config into /usr/local/bin. With LLVM in place, clone the repository, export WELD_HOME and build:

```bash
git clone https://www.github.com/weld-project/weld
cd weld/
export WELD_HOME=`pwd`
cargo build --release
```

WELD_HOME matters beyond the build. The README states that Grizzly needs it set because Grizzly locates its own native library through that variable. Finish by running `cargo test`, which executes the unit and integration tests; a substring argument filters them, so `cargo test <substring to match in test name>` runs a subset. For a first real use, the Python path is the shortest: bindings live in python/, examples in examples/python, and Grizzly, the Pandas subset integrated with Weld, is documented under python/grizzly with example workloads in examples/python/grizzly.

## Where Weld is the wrong tool

The build instructions are the first limitation, and they are not incidental. Pinning to LLVM 6.0 is unusual for a project whose last push was 2026-04-13. A machine that already carries a newer LLVM will need the 6.0 packages installed alongside it, and the symlink step deliberately overwrites /usr/local/bin/llvm-config, which can disturb other builds on the same host. That is a real cost for a single experiment. The second limitation is scope. Grizzly is described as a subset of Pandas, not Pandas. If your workflow depends on parts of the Pandas API outside that subset, Weld cannot run it, and the README does not list which parts are covered. Third, the lazy model changes when work happens. A computation is built up and only evaluated when a result is needed, so debugging is not the same as stepping through eager calls; the repository does ship an interactive REPL for inspecting and debugging programs, documented in docs/tools.md, which is a sign the authors know the model needs its own tooling. Finally, the README gives no benchmark numbers, no supported-workload list and no rollback or versioning policy. Anyone expecting a performance guarantee from the description alone is reading more into it than is written.

## Weld against Dask and Polars

The obvious comparison is with dataframe engines that also try to reduce intermediate materialization, and the difference is where the optimization lives. Dask and Polars optimize a graph or query plan built from their own operators. Weld instead asks each library to express its core computations in a shared IR, then optimizes across the libraries, which is why the repository ships a language document, a C API, a header generator and Python bindings rather than a single user-facing API. That is a bigger ask from an adopter: you are not just calling a faster function, you are agreeing to a representation. The payoff, if it materializes, is that the fusion happens across framework boundaries rather than inside one framework. The cost is that Weld only helps the parts of your pipeline that have been expressed in its IR. For a pipeline already written entirely in one engine, that engine's own optimizer is the more direct route, and Weld adds an LLVM 6.0 build dependency for no gain.

## Maintenance, releases and the BSD-3-Clause licence

The repository is not archived, and the last push was on 2026-04-13. The most recent release listed is v0.4.0 from 2020-02-13, preceded by v0.3.1 and v0.3.0 in August 2019. That gap between the last push and the last tagged release is the fact to plan around: commits have continued, but there is no recent tagged version to pin to, and the README does not describe a release cadence, a deprecation policy or an upgrade path between versions. For upgrade cost, treat the LLVM 6.0 requirement as the fixed part of your environment and budget for keeping those packages available. The licence is BSD-3-Clause, which is permissive and generally allows use in closed products with the copyright notice and disclaimer retained. It contains no copyleft obligation that would force you to publish your own code. This is a description of the licence identifier, not legal advice; read the LICENSE file in the repository and, if the code goes into a product, have counsel review it.

## Conclusion

Adopt Weld if you are building or extending a data library and want to see whether cross-library fusion pays off in your workload, or if you are studying how a common IR plus LLVM code generation handles data movement. Do not adopt it if you need a maintained general-purpose dataframe engine: the last release on the repository is v0.4.0 from 2020-02-13, and the README does not document rollback, versioning policy or a support window. Before committing, verify that LLVM 6.0 is available on your machine, that `llvm-config --version` reports 6.0.x after the symlink step, and that `cargo test` passes on your platform. Then reproduce one Grizzly example from examples/python/grizzly and compare it against the equivalent Pandas code on your own data, because the README gives no benchmark numbers and no promise about which workloads benefit.

## FAQ

### What is weld-project/weld?

It is a language and runtime for improving the performance of data-intensive applications. It expresses core computations from different libraries in a common intermediate representation so that a whole workflow can be optimized as one unit instead of function by function.

### How do I install weld-project/weld from source?

Install the latest stable Rust and LLVM/Clang++ 6.0, make sure llvm-config resolves to the 6.0 binary, then clone the repository, set WELD_HOME to the checkout directory and run cargo build --release followed by cargo test.

### What is Grizzly in weld-project/weld?

Grizzly is a subset of Pandas integrated with Weld, documented under python/grizzly with example workloads in examples/python/grizzly. The README notes that Grizzly needs the WELD_HOME environment variable set because it finds its own native library through that variable.

## Sources

- [License: BSD-3-Clause](https://github.com/weld-project/weld/blob/master/LICENSE)
- [Project website](https://www.weld.rs)
- [README](https://github.com/weld-project/weld/blob/master/README.md)
- [Releases](https://github.com/weld-project/weld/releases)
- [weld-project/weld on GitHub](https://github.com/weld-project/weld)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/weld-project-weld
