Model or dataset
google/langextract avatar
google/langextract

LangExtract grounds every extraction to a character interval, then hands you the ungrounded ones too

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

38,897 stars2,720 forksPythonApache-2.0

At a glance

What is it?
LangExtract is Google's Apache-2.0 Python library that turns unstructured text into structured records with source positions, chunking for long documents, and a self-contained HTML review file. Its guarantees are model-dependent, and filtering out example leakage is left to a line of your own code.
Who is it for?
LangExtract has the right shape for anyone who must point at the sentence behind every field, and the char_interval grounding is what makes human review possible at all. It is a library rather than a pipeline, so three things stay on you: filtering ungrounded spans, confirming which providers honour output_schema, and watching the model identifier.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Ungrounded extractions are handed back, not dropped

The central claim is source grounding: every extraction maps to an exact location in the source text, which is what lets the HTML review file highlight it. The mechanism has a deliberate gap in it. LLMs will sometimes lift content out of your few-shot examples instead of the input document, and the library detects that case and marks those extractions with char_interval = None rather than discarding them. So the object returned by lx.extract contains both real findings and example leakage, separated by nothing but a field you have to check yourself. The documented fix is a single comprehension filtering on e.char_interval. If you iterate result.extractions directly, which is the obvious line to write first, every entity the model copied out of your sample text is counted as an extraction from your document, and your recall figures are wrong in the flattering direction.

Schema enforcement stops at the models that support controlled generation

A consistent output schema is the second stated advantage, enforced from your few-shot examples, and the mechanism is named in the same breath: controlled generation, in supported models like Gemini. That qualifier carries the entire limitation. Gemini and OpenAI are the two providers said to accept an output_schema argument, with or without few-shot examples, and the specifics are deferred to docs/examples/output_schema.md. The built-in Ollama path is not named as a schema-bearing provider, so a local model is steered by imitation of your examples rather than by a decoded schema. Nothing in the library raises an error when a provider ignores the argument, which means the same code can look correct against a cloud model and return loose prose against a local one. Confirm which providers honour the schema before you build validation logic on top of the output shape.

Recall on long documents is paid for in extra model passes

Long documents are handled by a stated strategy of text chunking, parallel processing, and multiple passes, offered as the answer to the needle-in-a-haystack problem. Read that as a cost multiplier rather than a free optimisation. Every chunk is sent more than once, so the token bill for one document scales with the product of document length and pass count, and the parallel processing shortens wall-clock time without reducing spend. The quick start exposes neither knob, because lx.extract is shown taking only text_or_documents, prompt_description, examples, and model_id, so chunk size and pass count live in options the visible example never sets. The lever the documentation does offer is model tier: gemini-3.5-flash is the recommended default and gemini-3.1-flash-lite is put forward for high-volume or cost-sensitive work. Budget the two together, because they multiply.

python
# The input text to be processed
input_text = "Lady Juliet gazed longingly at the stars, her heart aching for Romeo"

# Run the extraction
result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    model_id="gemini-3.5-flash",
)

Misaligned few-shot examples raise a warning and the run continues

Your examples are treated as instructions to the model, so their shape matters and the library checks it. Each extraction_text in an example should be verbatim from that example's own text rather than paraphrased, with entities listed in order of appearance and no overlap between them. LangExtract raises Prompt alignment warnings by default when the pattern breaks. A warning is not an error: extraction proceeds with your misaligned examples still sitting in the prompt. The consequence is that a quality regression looks identical to a clean run in your logs, and the only signal is a line of warning output somebody eventually filters out. This is also why a tidied example, one with dates reformatted so the schema looks neat, is not available to you. Resolve the warnings before you judge the model, because a noisy prompt makes every later comparison meaningless.

The default model id carries a retirement date you cannot see here

Model selection is treated as a first-class decision and given a menu with an implicit cost ordering. gemini-3.5-flash is the recommended default for schema-constrained extraction, gemini-3.1-flash-lite is offered for high-volume or cost-sensitive workloads, and a current Gemini Pro model is the route for tasks needing deeper reasoning. For large-scale or production use a paid Gemini tier is suggested, to raise throughput and avoid rate limits. Underneath all of it sits a lifecycle caveat: Gemini models have defined retirement dates, and the documentation points at Google's model version page rather than printing a single date. So the example hard-codes a snapshot name, and if you pin that string in your own source, a model retirement becomes your outage. The remedy lives outside this repository, on a schedule nobody schedules for you.

A key-free local install still pulls the Google client libraries

The Ollama route needs no key and the tree ships examples/ollama/ to support it, but the declared dependencies are unconditional. google-genai>=1.39.0 and google-cloud-storage>=2.14.0 sit in the same list as pandas>=1.3.0 and PyYAML>=6.0, with no extras marker separating the local path from the cloud path. An air-gapped or key-free deployment therefore still installs the Google client libraries, which costs you disk and audit surface rather than correctness. The version floors tell a similar story: pydantic is declared as >=1.8.0, a floor from a much older era, while the same table requires google-genai>=1.39.0, so the resolver hands you a pydantic far newer than the declared floor implies. Read the floors as minimums nobody has re-tightened, not as the versions the code is written against. The package requires Python 3.10 or newer.

The Dockerfile installs from PyPI and starts a shell

The container recipe is short enough to read in full. It starts from python:3.10-slim, sets a working directory of /app, installs the package from PyPI, and sets the default command to python with no arguments. Three consequences for anyone planning to ship it. It installs langextract from the package index, so the image carries whatever version the index currently serves rather than the tree in your checkout, which matters the moment you are testing an unreleased change. The base image is pinned to Python 3.10, exactly the floor the project declares, so the only interpreter in the image is the oldest one the package claims to support. And the default command opens an interactive shell with no extraction entrypoint, no mount point for input text, and no model credentials wired in, so this is a packaging smoke test rather than a service you can deploy as written.

dockerfile
# Production Dockerfile for LangExtract
FROM python:3.10-slim

# Set working directory
WORKDIR /app

# Install LangExtract from PyPI
RUN pip install --no-cache-dir langextract

# Set default command
CMD ["python"]

Extension runs through a plugin folder and a hand-kept list

Two mechanisms extend the model layer, and they sit at different levels of support. Writing your own provider is a code path, with examples/custom_provider_plugin/ in the tree and a dedicated section on custom model providers. Third-party providers are a document, COMMUNITY_PROVIDERS.md at the repository root. That split is the part to watch, because a provider named in a hand-maintained Markdown file is not covered by the library's tests, its release history, or its version bumps, so when it breaks you are debugging someone else's integration against your own data. The rest of the tree suggests how the project is run: benchmarks/, tox.ini, a pre-commit config, a pylintrc, a skills/ directory, and a CITATION.cff alongside a Zenodo DOI. The three worked domains, full-text literature, medication extraction, and the RadExtract radiology report structurer, show the shapes it has been aimed at, and none of the three is a clinical validation. If your inputs are clinical notes, treat a demo extraction as a demo, not as accuracy evidence on your records.

Editorial conclusion

LangExtract has the right shape for anyone who must point at the sentence behind every field, and the char_interval grounding is what makes human review possible at all. It is a library rather than a pipeline, so three things stay on you: filtering ungrounded spans, confirming which providers honour output_schema, and watching the model identifier. Before committing, run the same documents through both the Ollama path and the default gemini-3.5-flash, compare what survives the char_interval filter, and check Google's model version page, because the snapshot name in the example carries a retirement date.

Frequently asked questions

what is langextract

LangExtract is a Python library that uses LLMs to extract structured information from unstructured text documents, guided by instructions you write and examples you supply. It is Apache-2.0 licensed, the package version is 1.7.0, and it requires Python 3.10 or newer. Extractions carry a character interval pointing back into the source text.

how to install langextract

The project metadata names the distribution langextract, version 1.7.0, requiring Python 3.10 or newer. The container recipe included in the repository installs it from PyPI with pip install --no-cache-dir langextract. Cloud models such as Gemini additionally need an API key, while the Ollama path does not.

how to use langextract

Call lx.extract with your input text, a prompt_description, a list of examples, and a model_id, then save the result with lx.io.save_annotated_documents to a .jsonl file. LangExtract then generates a self-contained interactive HTML file from that file so you can review the extracted entities in their original context.

Is Google LangExtract free?

The library is Apache-2.0 licensed and installs from PyPI at no cost. The model calls are a separate question: the quick start uses gemini-3.5-flash, and for large-scale or production use a paid Gemini tier is suggested to increase throughput and avoid rate limits. Local models through the built-in Ollama interface are the path that needs no key.

is langextract open source

It is. The repository is licensed under Apache-2.0, the pyproject.toml file header carries the Apache License 2.0 text, and the source sits in a langextract/ package directory beside tests/, benchmarks/, docs/, and examples/. A Zenodo DOI and a CITATION.cff file are provided for citing the work.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-langextract.svg)](https://hysenlabs.com/projects/google-langextract)