Model or dataset
langfuse/langfuse-python avatar
langfuse/langfuse-python

langfuse-python: tracing, datasets and prompts from a plain Python process

🪢 Langfuse Python SDK - Instrument your LLM app with decorators or low-level SDK and get detailed tracing/observability. Works with any LLM or framework

492 stars353 forksPythonMIT

At a glance

What is it?
The official Langfuse Python SDK wraps OpenTelemetry tracing, dataset experiments, LLM-as-a-judge scores and prompt management behind one client. It installs with pip, but the v4 rewrite changed the tracing API, so version choice is the first decision.
Who is it for?
Adopt langfuse-python if your application is Python 3.10 or newer and you want tracing, datasets, evaluation scores and prompt management behind one client instead of wiring separate tools. Skip it if you only need raw OpenTelemetry spans with no evaluation or prompt layer, or if you cannot run or reach a Langfuse server.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What langfuse-python is for, and who ends up using it

The package is the Python client for the Langfuse platform, and the README frames it as covering the whole product rather than tracing alone: observability and tracing, datasets and experiments, LLM-as-a-judge and custom evaluations, prompt management, and a REST API client. That combination is the reason to pick it over a generic tracing library. A team that already logs spans somewhere still has to answer separate questions: which prompt version produced this output, how did this model change score against the previous one, and can the same input set be replayed after a prompt edit. langfuse-python puts those jobs in one dependency.

The audience is Python teams building LLM features who want per-call telemetry tied to evaluation and prompt state. The pyproject file sets requires-python to >=3.10,<4.0, so the floor is Python 3.10. If your runtime is older, this package is not an option at any version. The dependency list is deliberately small: httpx, pydantic v2, backoff, wrapt, packaging, the three OpenTelemetry packages and typing-extensions. There is no vendor SDK in the runtime dependencies, which matches the README claim that it works with any LLM or framework.

How tracing actually flows: OpenTelemetry under the decorators

The mechanism is OpenTelemetry. The runtime dependencies include opentelemetry-api, opentelemetry-sdk and opentelemetry-exporter-otlp-proto-http, all pinned to the 1.33.1 line, so the SDK builds on the standard tracing model and exports over OTLP/HTTP rather than a private protocol. The README describes the tracing layer as OpenTelemetry-based and lists OpenAI and LangChain integrations on top of it.

The user-facing surface is a context manager. The quickstart creates a client with get_client(), opens a span with start_as_current_observation(as_type="span", name="process-request"), nests a generation inside it with as_type="generation" and a model argument, and updates output on each. The README states that spans are closed automatically when their context blocks exit. A second path exists for code that does not fit context managers: the README's own description mentions instrumenting with decorators or the low-level SDK, and the repository topics list decorators alongside pydantic. Those two styles are the split to understand before writing instrumentation, because one follows the call stack and the other requires you to manage observation lifetimes yourself.

The client reads configuration from the environment. The quickstart comment names LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and LANGFUSE_BASE_URL, and get_client() takes no arguments in that example, so credentials and the server address come from those variables.

Installing langfuse-python and getting a first trace out

Installation is a single pip command. The README marks the v4 rewrite as important and points to a migration guide for moving from v3, so decide the version before you write instrumentation code.

bash
pip install langfuse

The quickstart then expects three environment variables to be present before the client is created: LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and LANGFUSE_BASE_URL. The README shows them as a comment above the example rather than as export lines, so set them the way your deployment already handles secrets. With those in place, this is the smallest useful program: it opens a span, nests a generation, records outputs, and flushes.

python
from langfuse import get_client

langfuse = get_client()

with langfuse.start_as_current_observation(as_type="span", name="process-request") as span:
    span.update(output="Processing complete")
    with langfuse.start_as_current_observation(as_type="generation", name="llm-response", model="gpt-5.6") as generation:
        generation.update(output="Generated response")

langfuse.flush()

What you should see: two observations in the Langfuse UI, one nested inside the other, with the generation carrying the model name. The README notes that spans close automatically when their context blocks exit, and that flush is for short-lived applications. That last point matters in practice. A script that exits without flushing can lose buffered events, so the flush call belongs at the end of any process that does not stay alive.

For the rest of the surface, the README points to the SDK guide at langfuse.com/docs/observability/sdk/overview, the full docs at langfuse.com/docs, a machine-readable index at langfuse.com/llms.txt, and the API reference at python.reference.langfuse.com.

The version boundary is the real adoption risk

The README states plainly that the SDK was rewritten in v4 and released in March 2026, and that the migration guide covers updating code from v3 to v4. That sentence is the most consequential line in the document. Any tutorial, Stack Overflow answer or blog post written before the rewrite is describing a different API, and the search results around this package are full of v2 and v3 phrasing. If you copy an older snippet, the failure will be at import or attribute level, not a subtle behavioural difference.

A second constraint is the server. The pyproject test markers separate unit tests, which the file describes as deterministic tests that run without a Langfuse server, from e2e tests that require a real Langfuse server or persisted backend behaviour. That split is a fair signal about the shape of the product: the client library is only half the system. The README does not document what happens to buffered events when the server is unreachable, and it does not document rollback behaviour for the migration. Those are gaps to close against the migration guide and the full docs before committing to an upgrade window.

Finally, the SDK is a client, not a store. Nothing in the README suggests local persistence of traces, so a process that cannot reach a Langfuse instance is not collecting anything useful.

Where langfuse-python is the wrong tool

If your only requirement is spans in an existing OpenTelemetry pipeline, this package adds a platform contract on top of a standard you already have. The dependency list shows the OTLP HTTP exporter is bundled, so you are not gaining transport capability you could not get from opentelemetry-exporter-otlp-proto-http directly. What you gain is the Langfuse data model: observations with a type, generations with a model field, datasets, scores and prompts. If none of those concepts appear in your workflow, the extra dependency is overhead.

A second mismatch is non-Python services. The package is the Python SDK; the README scopes it that way and points to the broader Langfuse docs for everything else. A polyglot system will end up with this client in the Python services and something else everywhere else, which is fine but means the trace model has to be consistent across both.

A third case is teams that cannot operate or reach a Langfuse server. Since the e2e test marker distinguishes tests needing a real server from those that do not, the project itself treats a running backend as a normal part of the setup, not an optional extra.

Alternatives and how the approach differs

The honest comparison is against using OpenTelemetry directly. The opentelemetry-sdk and opentelemetry-exporter-otlp-proto-http packages give you spans, context propagation and an exporter, and they are already dependencies here. The difference is what sits above the span: langfuse-python adds typed observation kinds, a generation concept with a model attribute, dataset and experiment runs, evaluation scores, and prompt management. If you need those, building them on raw OpenTelemetry means designing your own schema and your own storage for scores and prompt versions. If you do not, raw OpenTelemetry is fewer moving parts and no platform dependency.

The second alternative is a framework-native tracer, for example whatever ships with your orchestration library. The README lists LangChain and OpenAI integrations, which suggests the intended pattern is to use this SDK alongside a framework rather than instead of it. The trade-off is the same one you get with any integration layer: you inherit the SDK's release cadence. The v3 to v4 rewrite is the concrete example. A framework-native tracer moves when the framework moves; this package moves on its own schedule, documented in its own migration guide.

Licence, maintenance and upgrade cost

The licence is MIT, declared in pyproject.toml under license = "MIT" with license-files = ["LICENSE"], and the README carries an MIT badge. MIT is permissive, so the usual obligations are attribution and including the licence text; this is a description of the licence, not legal advice, and your own counsel should review how it interacts with your distribution model.

On maintenance, the repository is not archived and the last push was on 2026-09-10. Recent releases are v4.15.2 on 2026-09-09, v4.15.1 on 2026-08-28 and v4.15.0 on 2026-08-27. The patch releases inside a single month tell you the upgrade cadence is frequent, which is good for fixes and a cost for anyone pinning versions. The dependency ranges are loose in places: httpx is constrained to >=0.15.4,<1.0, pydantic to >=2,<3, and the OpenTelemetry packages to the 1.33.1 line. A loose httpx range means a resolver can pull a version you have not exercised, so pinning in your own lockfile is the practical defence.

The upgrade cost is concentrated in the v4 rewrite. The README directs readers to the migration guide, and the pyproject markers show the project maintains a real e2e suite against a server, which is the kind of suite that catches client-side breakage. Budget the migration as a code change across every instrumentation site, not a version bump.

Editorial conclusion

Adopt langfuse-python if your application is Python 3.10 or newer and you want tracing, datasets, evaluation scores and prompt management behind one client instead of wiring separate tools. Skip it if you only need raw OpenTelemetry spans with no evaluation or prompt layer, or if you cannot run or reach a Langfuse server. Before writing code, confirm whether you are starting on v4 or migrating from v3, because the tracing API changed in the v4 rewrite; the migration guide is the document to read first. Then check that LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and LANGFUSE_BASE_URL are set in every process that calls get_client().

Frequently asked questions

What is Langfuse and what does the langfuse-python SDK do?

Langfuse is a platform for LLM observability, datasets and experiments, evaluation scores and prompt management. The langfuse-python package is its Python SDK, covering tracing, datasets and experiments, LLM-as-a-judge and custom evaluations, prompt management, and a full REST API client.

Can Langfuse be run locally with the Python SDK?

The SDK reads its server address from the LANGFUSE_BASE_URL environment variable, so it points at whatever instance you configure. The pyproject test markers distinguish unit tests that run without a Langfuse server from e2e tests that require a real one, which shows a running backend is part of the normal setup.

How do you use langfuse-python in an application?

Install it with pip install langfuse, set LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and LANGFUSE_BASE_URL, then call get_client(). From there you open observations with start_as_current_observation, nesting a generation inside a span, and call langfuse.flush() in short-lived applications.

Who created Langfuse?

The README carries a Y Combinator W23 badge, and pyproject.toml lists the package authors as langfuse with the contact address [email protected].

What are the benefits of using langfuse-python instead of plain OpenTelemetry?

The SDK is built on OpenTelemetry and exports over OTLP/HTTP, so the transport is not the difference. What it adds on top is the Langfuse data model: typed observations, generations with a model field, datasets and experiments, evaluation scores, and prompt management.

Official sources

  1. langfuse/langfuse-python on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/langfuse-langfuse-python.svg)](https://hysenlabs.com/projects/langfuse-langfuse-python)