Model or dataset
Open-Source-Legal/OpenContracts avatar
Open-Source-Legal/OpenContracts

OpenContracts: a self-hosted citation graph for document corpora

The open document intelligence platform for builders and hackers - DMS for the agentic world

1,493 stars194 forksPythonMIT

At a glance

What is it?
OpenContracts turns a folder of documents into a queryable citation graph with annotations, structured extraction, AI agents and an MCP server behind one API. Here is what the repository documents, and where it stops.
Who is it for?
Adopt OpenContracts if you have a corpus of citation-heavy documents, a Python team willing to run Django, Celery and Postgres, and a need for a graph you can query yourself rather than a SaaS dashboard. Do not adopt it if you need a finished contract lifecycle product with redlining, approvals and vendor support, or if nobody on the team will own a container stack.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem OpenContracts is built around

Most document tools stop at full-text search. You upload a folder, you get a search box, and the relationships between documents stay in the reader's head. OpenContracts takes the opposite position: the product is the graph. Point it at a repository of documents and it resolves references between them, so a filing is linked to the statute section it cites, and that statute is linked to everything citing it back.

The README frames the audience directly: builders and hackers, not legal operations buyers. The demo described in the README runs against a local install with no custom code, using 36 SEC filings wired to the Delaware General Corporation Law, the Securities Act and SEC rules. Law the library does not hold yet is not discarded. According to the README, it is tracked as a backlog with dashed nodes until you ingest it. That backlog behaviour is the part worth noticing, because it means an incomplete corpus still produces a usable graph.

This is a platform, not an application. The README states that everything the UI does runs on surfaces you can call yourself, and that the same graph is exposed through three surfaces: a GraphQL plus REST API, a Model Context Protocol server, and a React UI.

How the citation graph and extraction pipeline fit together

The architecture visible in the repository is a Django backend (opencontractserver/, manage.py, schema.graphql and schema.json at the root) with a React frontend in frontend/ and Celery handling asynchronous work. Long-running jobs are the norm here: parsing, embedding and thumbnailing are described in the README as swappable components, and the pipeline overview lives at docs/pipelines/pipeline_overview.md.

Structured extraction follows the same fan-out model. You define a fieldset, which the README describes as a set of columns where each column is a natural-language query, and run it across a corpus. Extraction fans out over Celery workers and lands in a spreadsheet-style grid, with human approve or reject on every cell. That review step matters: it means the output is a proposal until a person accepts it, which is a different posture from tools that write extracted values straight into a record.

Agents are scoped to a document or a corpus. The README gives this example:

python
agent = await agents.for_document(123, corpus=45)
async for chunk in agent.stream("Summarize the indemnification clauses"):
    print(chunk.content, end="")

Answers are described as grounded in the annotations and citations already in the corpus, not in a general model's memory. The deeper LLM framework documentation sits at docs/architecture/llms/README.md, and the README does not reproduce its contents.

Installing OpenContracts and running a first corpus

The repository ships several compose files rather than a single documented install path: local.yml, production.yml, test.yml, test.e2e-coverage.yml, test.authority-e2e.yml and local.e2e-coverage.yml at the root, plus a compose/ directory and an .envs/ directory for environment files. There is also a merge_production_dotenvs_in_dotenv.py script, which suggests the project expects multiple environment files to be combined for production. The README does not walk through a first install, so treat the compose files as the source of truth and read them before you run anything.

The general shape, based on those files, is a Docker Compose stack started from the local configuration. The repository provides local.yml for exactly this, and the README does not state which ports the stack binds, so check the compose file rather than assuming.

Once the stack is up, the README's own workflow is short: create a corpus, add documents, and click Set up. That single action installs the intelligence bundle, which the README says makes agents describe and summarize every document and starts resolving statutory citations into graph edges. The result is a navigable graph, and the README notes that citations are highlighted inline on the documents themselves.

To reach the graph from your own code rather than the UI, the API is the entry point. The MCP server is documented at docs/mcp/, with an anonymous endpoint at /mcp/ for public corpuses and an authenticated endpoint at /mcp/me/. Discovery files are served at /llms.txt and /.well-known/mcp.json. The tools the README lists are search_corpus, list_documents, get_document_text, list_annotations, list_relationships, list_threads and create_thread_message. Note the asymmetry: only create_thread_message writes anything, and the README describes annotation proposals as requiring authorization.

Where OpenContracts is the wrong tool

The README is candid that this is infrastructure for people who will build on it, and that framing sets the boundary. If you want a contract lifecycle product with redlining, approval routing, obligation tracking and a vendor to call when it breaks, OpenContracts is not that product. It gives you a graph and the surfaces to query it. The workflow layer is yours to write.

The operational cost is real. A Django application with Celery workers, a broker, Postgres and a React frontend is not a small footprint, and the repository's root directory carries production.yml, test.yml and several e2e coverage variants, which tells you the maintainers run a non-trivial matrix of environments. A team without anyone comfortable operating containers and a Python service will spend more time on the stack than on the documents.

The citation resolution is also domain-shaped. The README's worked example is statutory citations in SEC filings, and the intelligence bundle is described in terms of detecting and resolving statutory references. A corpus of, say, internal engineering design docs or support tickets has no statute graph to build. You could still use annotations, extraction and agents, but the headline feature would be idle.

Finally, the documentation has gaps. The README does not document rollback, backup or upgrade procedures, and it does not publish a migration guide between the v3 releases. The changelog.d/ directory and CHANGELOG.md exist, so release notes are maintained, but operational runbooks are not described in what is published.

How it compares with contract analysis libraries

The related searches around this project surface names like Accord Project, LexPredict and OpenCLM, and the difference in approach is worth stating plainly. Accord Project is a specification and tooling effort for smart legal contracts: it focuses on templates and executable contract logic, so the artifact is a contract that can be evaluated. OpenContracts does not execute contracts. It ingests documents you already have, resolves the citations between them, and exposes the resulting graph.

LexPredict's public work sits closer to analytics and prediction over legal text. That is a modelling exercise: train or apply a model, get a classification or a risk signal. OpenContracts is a storage and retrieval layer with agents on top. The persistent object is the annotated document and its edges, not a model's output.

OpenCLM and the broader contract lifecycle category target the process around a contract: drafting, negotiation, signature, renewal. OpenContracts starts after the document exists and asks what it references. If your problem is getting a signature, none of the graph machinery helps you.

The closest comparison is a general purpose RAG stack assembled from a vector database and a framework. OpenContracts does include embeddings and a vector store, but the README positions the citation graph as the product, and the MCP server exposes graph traversal tools like list_relationships alongside search_corpus. That is the distinction: retrieval over chunks versus traversal over typed edges.

Licence, maintenance and the cost of upgrading

OpenContracts is MIT licensed, and the repository carries a CLA.md contributor licence agreement, which matters if you plan to contribute rather than just consume. MIT is permissive: you can self-host, modify and redistribute, and the licence badge in the README points at the standard MIT text. This is not legal advice, and if you are embedding the software in a commercial product you should read LICENSE and CLA.md yourself.

The project is not archived, and the last push was on 2026-09-10. Releases are frequent enough to plan around: v3.1.0 on 2026-09-08, v3.0.0 on 2026-08-10, and a v3.0.0.b4 beta earlier in the year. That cadence is a cost as well as a signal. Three releases across the v3 line in roughly two months means schema and API surface can move, and the README does not publish a v3 upgrade guide. The schema.graphql and schema.json files at the root are the contract you should diff between versions before you upgrade a running instance.

The upgrade work itself is the usual Django shape: pull the new image or code, apply migrations, restart the web and worker processes. Because Celery handles parsing, embedding and extraction, an upgrade that changes a pipeline component can leave in-flight jobs in an inconsistent state. The repository does not document how it handles that, so a maintenance window with workers drained is the safer assumption.

Editorial conclusion

Adopt OpenContracts if you have a corpus of citation-heavy documents, a Python team willing to run Django, Celery and Postgres, and a need for a graph you can query yourself rather than a SaaS dashboard. Do not adopt it if you need a finished contract lifecycle product with redlining, approvals and vendor support, or if nobody on the team will own a container stack. Before committing, verify three things: that the compose files under compose/ and the production.yml entry point match your deployment target, that the parsing pipeline handles your file formats, and that the MCP tools listed in docs/mcp/ cover the operations your agents need.

Frequently asked questions

What is OpenContracts?

It is an MIT-licensed, self-hosted document intelligence platform. You point it at a repository of documents and it produces a citation graph with human annotation, structured extraction and AI agents, exposed through a GraphQL plus REST API, an MCP server and a React UI.

Does OpenContracts have an MCP server for AI agents?

Yes. The README lists an anonymous endpoint at /mcp/ for public corpuses and an authenticated endpoint at /mcp/me/, with discovery at /llms.txt and /.well-known/mcp.json. The documented tools include search_corpus, list_documents, get_document_text, list_annotations, list_relationships, list_threads and create_thread_message.

How do I install OpenContracts?

The repository ships Docker Compose files including local.yml and production.yml, plus a compose/ directory and an .envs/ directory. The README does not give step-by-step install instructions, so the compose files are the place to start.

What licence is OpenContracts released under?

MIT, per the licence badge and the LICENSE file at the repository root. The repository also includes a CLA.md contributor licence agreement.

Can OpenContracts run entirely on my own infrastructure?

The README describes it as self-hosted, and the repository provides Docker Compose configurations for local and production use. There is also a hosted demo at contracts.opensource.legal, but self-hosting is the documented deployment model.

Official sources

  1. License: MIT
  2. Open-Source-Legal/OpenContracts on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/open-source-legal-opencontracts.svg)](https://hysenlabs.com/projects/open-source-legal-opencontracts)