# ContextGem: a Python framework for LLM document extraction without prompt writing

> ContextGem wraps LLM document extraction in declarative Python objects, returning structured data with paragraph-level references and justifications. Here is what it does, how to install it, and where it stops being the right tool.

**shcherbak-ai/contextgem** — ContextGem: Effortless LLM extraction from documents

- Repository: https://github.com/shcherbak-ai/contextgem
- Website: https://contextgem.dev/
- Stars: 2,007 · Forks: 186
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/shcherbak-ai-contextgem

## The problem ContextGem targets: extraction code that nobody wants to maintain

Getting a structured record out of a long PDF is not one task. It is five: writing the extraction prompt, defining a validation schema, mapping each output field back to the sentence it came from, chaining several extraction steps, and counting tokens across providers. Teams usually build that scaffolding themselves, once per document type, and then maintain it as models change.

ContextGem's stated motivation is to absorb those five jobs into one declarative layer. The README puts it plainly: you describe what to extract in natural language and the framework handles how. The audience is developers, not analysts. Everything is Python objects, and the package keywords point at contract analysis, legaltech and fintech, where the extracted value is worthless without a citation back to the source text. That citation requirement is the real differentiator. A generic JSON-mode call gives you a field; ContextGem's stated output includes paragraph and sentence level references and automatic justifications, which is what a reviewer needs to trust the field.

## Aspects, concepts and the extraction pipeline

The framework is built on two abstractions. An aspect is a slice of a document defined in natural language, such as a topic, theme or category. A concept is something to pull out of that slice: an entity, a fact, a conclusion, an assessment. Aspects can contain concepts, and aspects can nest inside other aspects, which is how the project describes multi-level pipelines.

Around those two sit the parts that would otherwise be your glue code. Prompts are generated dynamically rather than written by hand. The data model is generated from your concept definitions, and the project uses Pydantic v2 for validation. Output is mapped back to source references at paragraph and sentence granularity. Documents are held in a serializable storage model, so a parsed document can be persisted and reused instead of re-parsed on every run. LLM usage is tracked across calls.

The README's quick-start example extracts anomalies from a legal document, which the project calls a complex concept requiring contextual understanding. That choice of example is telling: the framework is aimed at fields that a regex cannot find and a single prompt struggles to define consistently. The trade-off is that every one of those abstractions is another object to learn. A one-off extraction of a due date from a PDF does not need an aspect hierarchy.

## Installing ContextGem and running a first extraction

The README recommends uv and also documents pip. Python 3.10 through 3.14 are supported according to the project metadata, which declares requires-python >=3.10,<3.15. Pick one installer and stay with it; mixing them in the same environment is how you end up with two copies of the package.

```bash
uv add contextgem
```

Or, with pip:

```bash
pip install -U contextgem
```

After installation, the workflow in the documentation is to build a Document from your source content, attach aspects and concepts defined in natural language, run the extraction, and read the structured result. The README's own quick-start uses a legal document and an anomaly concept. The package ships with an example set; the repository layout includes a dev/ directory alongside contextgem/ and tests/, and the documentation site at contextgem.dev carries the full walkthrough. Start there rather than inventing your own first example, because the concept definitions are where the framework's behaviour is actually decided.

One practical note before you scale up: the extraction call goes to an LLM, so the number of aspects and concepts you define multiplies into the number of model calls. Prototype on a single short document and watch the usage tracking the framework provides before you point it at a folder of contracts.

## Where ContextGem is the wrong tool

It is a Python library. If your extraction work lives in a no-code tool, a spreadsheet, or a service written in another language, nothing here helps you; you would be adding a Python sidecar to a stack that does not want one. The project metadata classifies the release as Development Status 4 - Beta, and the version numbers reflect that: 0.27.0, with 0.26.0 and 0.25.1 in the months before it. Beta status is not a defect, but it means the API surface can move between minor versions, and you should read the changelog before every upgrade rather than assuming a patch-level bump is safe.

The deeper limitation is inherent to the approach. Extraction quality is bounded by the model you point it at, not by the framework. ContextGem removes the prompt-writing and schema-mapping work; it does not make a small model reliable on a dense 80-page agreement. If your requirement is deterministic output, a rule-based parser or a fine-tuned classifier will beat any prompt-driven pipeline on repeatability, and it will do so without a per-document inference bill. And if your documents are short and your fields are simple, the abstraction cost outweighs the benefit. This is a framework for extraction problems that are genuinely hard, not a general-purpose PDF reader.

## How it differs from LangExtract and ExtractThinker

The related searches around this project point at LangExtract and ExtractThinker, and the comparison is worth making concrete. LangExtract is Google's extraction library, and the search data shows people asking whether it is free and how it compares; the distinction that matters is ownership and model coupling. ContextGem is Apache-2.0 and provider-agnostic by design, with usage tracking across LLMs as a stated feature, so switching models is a configuration change rather than a rewrite. A library tied to one vendor's stack makes that switch a migration.

ExtractThinker takes a document-contract metaphor rather than an aspect-and-concept one. Where ContextGem asks you to describe aspects and concepts in natural language and lets the framework generate prompts and schemas, ExtractThinker's model centres on contracts between a document and the data you want from it. Both are Python, both target structured extraction, and the difference is mostly which mental model you find easier to hold in your head when a pipeline has five steps. Neither is objectively better. The honest test is to write the same extraction in both and see which one you can still read three months later.

LLMAIx and Langextract appear in the same search neighbourhood. The pattern across all of them is the same: the hard part is not the library, it is deciding what the correct output for a document actually is.

## Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-08-13, the same day v0.27.0 was released. That is recent enough that the project is being worked on, but the release cadence is the number to plan around: 0.25.1 in June 2026, 0.26.0 in late July, 0.27.0 in mid August. Roughly a minor release a month means the changelog is required reading, not optional. The repository ships a CHANGELOG.md at the top level, so the information is there; the cost is the time to read it and re-run your extraction tests after each bump.

The licence is Apache-2.0, declared in pyproject.toml as `license = { text = "Apache-2.0" }` and in the LICENSE file. For most teams that is permissive enough to use commercially without a conversation. Two things to check yourself rather than assume. First, the repository contains a LICENSES.md and a licenseal.review.toml, which suggests the project audits its dependency licences; that audit covers the project's own dependencies, not the terms of the LLM provider you point it at. Second, Apache-2.0 covers the framework, not the model output. Your obligations around the extracted data come from your provider contract and your own regulatory context, and nothing in this repository changes them. This is not legal advice; if the data is regulated, ask someone qualified.

## Conclusion

ContextGem fits Python teams that need structured fields plus source references from contracts, reports or other long documents, and are willing to accept a beta-stage API and per-document LLM cost. It is the wrong choice if you need a no-code interface, a vendor SLA, or deterministic output without a model in the loop. Before committing, install it with uv add contextgem, run the extraction example from the README against one real document, and count the LLM calls your aspect and concept definitions produce; that count is your cost model.

## FAQ

### What is ContextGem and what is it used for?

ContextGem is a free, open-source Python framework for extracting structured data and insights from documents with LLMs. It handles prompt generation, data modelling, reference mapping and pipeline orchestration so you describe what to extract rather than how.

### How do I install ContextGem?

The README recommends uv with the command uv add contextgem, and also documents pip with pip install -U contextgem. Python 3.10 through 3.14 are supported according to the project metadata.

### Can LLMs be used for document extraction?

ContextGem is built entirely on that premise: the README describes extracting structured data and insights from documents, including text and images, with automatic justifications and paragraph-level references. The framework's role is to make that extraction repeatable rather than to replace the model.

### What exactly is context engineering?

The project does not define the term. What ContextGem does in that space is generate extraction prompts dynamically and map outputs back to paragraph and sentence level references, which is the context it attaches to each extraction request.

## Sources

- [License: Apache-2.0](https://github.com/shcherbak-ai/contextgem/blob/main/LICENSE)
- [Project website](https://contextgem.dev/)
- [README](https://github.com/shcherbak-ai/contextgem/blob/main/README.md)
- [Releases](https://github.com/shcherbak-ai/contextgem/releases)
- [shcherbak-ai/contextgem on GitHub](https://github.com/shcherbak-ai/contextgem)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shcherbak-ai-contextgem
