W&B Weave: tracing and evaluation for LLM applications, from the team that sells experiment tracking
Weave is a toolkit for developing AI-powered applications, built by Weights & Biases.
At a glance
- What is it?
- A Python toolkit that decorates your functions to record inputs, outputs and traces, and is currently midway through pruning its own boards product down to tracing and evaluations.
- Who is it for?
- Weave is a reasonable choice if you are already inside the Weights & Biases ecosystem, or if what you want is evaluation storage attached to traces rather than a separate experiment tracker. Outside that ecosystem the trade is sharper: a free-tier W&B account is listed as a prerequisite, so evaluating an LLM application now means sending metadata about your prompts to a hosted service.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the toolkit claims to do, in three bullets
The README is short, and its three claims are the whole pitch. You can log and debug language model inputs, outputs and traces. You can build rigorous, apples-to-apples evaluations for language model use cases. You can organize the information generated across the workflow, from experimentation to evaluations to production.
The stated goal underneath those is to bring rigor, best practices and composability to an experimental process without introducing cognitive overhead. That last clause is the one to test, because instrumentation that costs you a decorator on every function is a real ongoing expense in a codebase that changes often.
What the project is not is a model server or an agent framework. It has no opinion about which provider you call, no chain abstraction and no tool-calling helper. It instruments what you already wrote.
Two prerequisites narrow who can try it: Python 3.10 or higher, and a Weights & Biases account with a free tier available. The account requirement is the one that changes the calculus for teams who deliberately keep prompt data off third-party infrastructure, since even the free tier means your trace metadata is hosted rather than local.
Traces come from a decorator, not from monkey-patching your provider
The mechanism is a decorator called `weave.op`. Apply it to any function and Weave builds a trace tree of inputs and outputs across everything it has wrapped. The README is explicit that this works for provider calls such as OpenAI, Anthropic and Google AI Studio, for Hugging Face and other open source generation calls, and equally for your own validation functions and data transformations.
That last clause is the useful part. A typical tracing library only sees the model boundary. Weave will record a pure function that reshapes a dataframe between two model calls, which means a trace can show where the data went wrong rather than just which model answered.
The quick start is three steps:
pip install weaveimport weave
weave.init("my-project-name")@weave.op
def my_function():
# Your tracked code!
passNote the order dependency. `weave.init` names the project, and the decorator has to be applied after initialization for the trace tree to attach to the right destination.
The README's fuller example composes two traced functions and a third that calls both, and it also shows a traced function wrapping an OpenAI chat completion with a JSON response format and a temperature setting. That combination, deterministic decoding plus structured output, is the shape of the evaluation use case the README leads with.
The repository is mid-prune, and the README admits it
The Contributing section says plainly that the project is in the middle of a cleanup, and that the codebase contains a large amount of code for the Weave engine and Weave boards, which have been put on pause while the team focuses on tracing and evaluations. That sentence is more useful than any feature list, because it tells you which parts of the library you should not build a new dependency on.
The README then maps the remaining code: tracing lives mostly in `weave/trace` and `weave/trace_server`, evaluations live mostly in `weave/flow`. Those are the two directories worth reading.
The wider repository tells a larger story about the project's shape. Beyond `weave/` there are `sdks/`, `examples/` with a `remote_scorer/` case, `tests/`, `trace_server_mock/`, `rules/`, `skills/`, `dev_docs/`, a `noxfile.py`, a `uv.lock` and a `Makefile`. A `trace_server_mock` in particular implies the server can be run locally for tests, which is the detail you want if you are deciding how much of your workflow can run offline.
Two smaller signals worth noting. The default branch is `master`, which will surprise you if you script against it. And the root `Makefile` contains a target that runs `cd weave && make generate_base_object_schemas` and then reaches four directories up into a `frontends/weave` package to run `yarn generate-schemas`, which only makes sense if this repository is checked out inside a larger W&B monorepo. That is a clue about how the published package is developed, and a reason not to expect this tree to build standalone without that surrounding structure.
Packaging details that will decide whether the install is easy
The `pyproject.toml` is worth reading for anyone who has fought a dependency resolver on an LLM project. The pinned dependencies include pydantic 2, sentry-sdk, tenacity with an explicit exclusion of 8.4.0 because that version had an import bug affecting `AsyncRetrying`, click, and packaging for version parsing in integrations.
That tenacity exclusion is a small but telling detail. It means a project already on tenacity 8.4.0 cannot take this dependency without moving, and it tells you the maintainers are pinning around a known upstream defect rather than leaving it to chance.
The project classifiers are less current than the code. They list `Development Status :: 4 - Beta`, and the topic strings describe a data-driven application toolkit with database, spreadsheet and visualization framing, along with a Flask framework classifier. The README describes a generative AI toolkit instead. Both descriptions are in the repository, and the README is the one that matches what the code does.
On that mismatch, the practical question is whether to trust the package metadata or the README when you evaluate the project. Trust the README for scope, and check the dependency list for integration cost.
The package also declares its license through a file reference rather than an SPDX string, and classifies itself under the Apache Software License, which is consistent with the Apache-2.0 license recorded for the repository. For a toolkit that will send prompt and trace data to a hosted service, that combination matters: the code is permissively licensed, while the data handling is a service question rather than a licensing one.
How it compares with LangSmith, Langfuse and plain logging
The honest comparison is with other LLM observability tools, and each has a different reason to exist.
LangSmith sits inside the LangChain ecosystem and is strongest when your application is built from LangChain or LangGraph primitives. Its traces are graph-shaped because the runtime is graph-shaped. Weave has no dependency on a chain library at all, so it applies equally to hand-written code, and its `weave.flow` evaluation module is described in the README as sitting alongside `weave/trace` rather than being coupled to it.
Langfuse is the self-hostable option. That is the substantive difference from Weave, and it is larger than the feature lists suggest: Weave's prerequisite list names a W&B account, so the hosted path is the documented one. If your requirement is that trace data never leaves your infrastructure, that requirement decides the comparison before features matter, and you should read what the `trace_server_mock` directory enables before assuming either way.
Plain logging is the baseline that matters most, because it is what you already have. A decorator plus a hosted store buys you trace trees across function boundaries and evaluation results that stay attached to the traces they came from. Whether that is worth an account and a network dependency is a question about your data, not about the code.
One limitation to name plainly: the README documents the tracing and evaluation surface and defers everything else to wandb.me/weave. Server deployment, authentication, retention and cost are not described in the repository, so if you are sizing this for production you are reading external documentation.
Editorial conclusion
Weave is a reasonable choice if you are already inside the Weights & Biases ecosystem, or if what you want is evaluation storage attached to traces rather than a separate experiment tracker. Outside that ecosystem the trade is sharper: a free-tier W&B account is listed as a prerequisite, so evaluating an LLM application now means sending metadata about your prompts to a hosted service. The project is also mid-refactor, with the README stating that Weave engine and Weave boards code is on pause while the team focuses on tracing and evaluations, so judge the library by its current surface rather than by its history. Latest release is v0.53.7 from 2026-08-27 with a last push of 2026-09-18, and the default branch is `master`, not `main`. Verify first that `weave.init` plus a single `@weave.op` decorator shows you a trace tree before instrumenting a pipeline.
Frequently asked questions
What is Weave from Weights & Biases?
A Python toolkit for developing generative AI applications. It traces function inputs, outputs and call structure with a decorator, stores evaluations against those traces, and organizes the information generated from experimentation through to production. It is built and documented by Weights & Biases.
How do I install Weave and record my first trace?
Install with `pip install weave`, then call `weave.init` with a project name, then decorate the functions you want to record with `@weave.op`. The decorator must be applied after initialization. The README also shows a traced function wrapping an OpenAI chat completion to return a JSON object.
Do I need a Weights & Biases account to use Weave?
The README lists a Weights & Biases account as a prerequisite, noting that a free tier is available, alongside Python 3.10 or higher. So the hosted path assumes an account, which matters if your requirement is that prompt and trace data stay inside your own infrastructure.
What is the difference between Weave tracing and Weave boards?
Both exist in the codebase, but the README states that the Weave engine and Weave boards code has been put on pause while the team focuses on tracing and evaluations. Tracing lives mostly in `weave/trace` and `weave/trace_server`, evaluations in `weave/flow`, and those are the parts to build a new dependency on.
How does Weave compare with LangSmith for tracing LLM applications?
LangSmith is strongest when the application is built on LangChain or LangGraph, since its traces are shaped by that runtime. Weave has no dependency on a chain library and can decorate hand-written functions including pure data transformations between model calls. It is the hosted W&B ecosystem, though, and the README lists a W&B account as a prerequisite.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/wandb-weave)