Model or dataset
yifanfeng97/Hyper-Extract avatar
yifanfeng97/Hyper-Extract

Hyper-Extract: turn documents into graphs and hypergraphs from the CLI

Hypergraph is more powerful. Transform unstructured text into structured knowledge with LLMs. Graphs, hypergraphs, and spatio-temporal extractions — with one command.

4,044 stars463 forksPythonNOASSERTION

At a glance

What is it?
Hyper-Extract is a Python CLI and library that sends unstructured text through an LLM and writes back a knowledge structure you can search, tag and export. It is early alpha software with a wide provider matrix, and the interesting part is the hypergraph and incremental-update model rather than the parsing itself.
Who is it for?
Adopt Hyper-Extract if you need hyperedges or spatio-temporal structures and are willing to run a 0.x alpha whose pyproject.toml still carries the Alpha classifier. Do not adopt it if you need a stable schema contract, an offline pipeline without an LLM key, or a documented rollback path for every failure mode.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Hyper-Extract targets: text in, structure out, without writing a pipeline

Most teams that want a knowledge graph from their own documents end up assembling the same four pieces: a document loader, a prompt that asks an LLM for triples, a normaliser that reconciles entity names, and a store that can be queried. Hyper-Extract packages that assembly behind one command. The README frames it as "Transform documents into structured knowledge with one command," and the CLI surface backs the claim: he parse reads a file, he search queries the result, he show renders it, he export obsidian writes a vault.

The intended users are named explicitly in the repository's own scenario list. A researcher feeds a paper and gets concepts, authors and citations. A financial analyst runs finance/earnings_graph over an earnings report and then asks for risk factors. A team that cannot send text to a hosted API runs Qwen3.5-9B and bge-m3 behind vLLM and keeps everything on-premise. The common thread is that none of them want to own the extraction code.

The differentiator is not parsing. It is the set of nine knowledge structures, which the README lists as ranging from raw chunk corpora and simple lists up to Graphs, Hypergraphs and Spatio-Temporal Graphs. A hyperedge connects more than two nodes at once, which matters when the fact you care about is a meeting, a co-authorship or a drug combination rather than a pairwise relation. That is the part worth evaluating, and it is also the part most competing tools do not offer.

How the extraction pipeline is put together

The dependency list in pyproject.toml is the clearest statement of architecture. langchain and langchain-openai handle model calls, faiss-cpu provides the vector index behind semantic search, pydantic appears in the keyword list for the data models, and typer plus rich build the CLI. Two less common packages, ontomem and ontosight, sit alongside them and are not described in the README excerpt, so their role has to be read from the code rather than the docs.

The flow the README implies is: a reader converts the input file into text, a template decides what to ask for, an LLM returns structured output, and the result is written into an output directory that holds both the structure and its embeddings. Templates are YAML and there are said to be more than 80 of them, grouped into Finance, Legal, Medical, TCM, Industry and General domains. A template name such as general/biography_graph or finance/earnings_graph is passed with the -t flag, so the extraction target is configuration rather than code.

Two mechanisms deserve attention. First, incremental evolution with provenance: he feed takes an updated document under the same --source, and the README states that old facts roll back automatically. he info ./output/ --sources then reports which documents contributed what, and he remove --document reverses an ingestion. Second, source tags and scoped search, added in the v0.8.1 and v0.8.2 releases, which let you attach a tag to a source and restrict a query to it. Both are the features that separate a demo from something you can keep running as your corpus changes.

Installing Hyper-Extract and running a first extraction

The README's quick start uses uv, and the package is also on PyPI as hyperextract, so pipx works as an alternative. Python 3.11 or newer is required according to both the README badge and the requires-python field in pyproject.toml. The first block installs uv and then the CLI as a standalone tool.

bash
curl -LsSf https://astral.sh/uv/install.sh | sh
uv tool install hyperextract

After that, the he command should be on your PATH. The next step is provider configuration, because nothing works without a model. The README offers a single-key path through OpenAI and a split path where the LLM and the embedder come from different vendors. DeepSeek is LLM-only and must be paired with an OpenAI-compatible embedder.

bash
he config init -p openai -k YOUR_OPENAI_API_KEY

If you would rather keep inference local, the README gives a vLLM configuration with two endpoints, one for the chat model and one for embeddings. The API key is the literal string dummy because the local server does not check it.

bash
he config llm -p vllm -u http://localhost:8000/v1 -k dummy -m Qwen/Qwen3.5-9B
he config embedder -p vllm -u http://localhost:8001/v1 -k dummy -m BAAI/bge-m3

With a provider in place, a first real run is one line. The README's own example parses a Markdown file with the biography template and writes the result to ./output/. The -l en flag sets the language.

bash
he parse examples/en/tesla.md -t general/biography_graph -o ./output/ -l en
he search ./output/ "What are Tesla's major achievements?"

The second command is what you should look at first. If search returns nothing useful, the extraction or the embedding step is misconfigured, and no amount of template tuning will fix it. From there, he show ./output/ opens a visual view, and he export obsidian ./output/ -o ./vault/ writes Markdown notes joined by [[wikilinks]].

Rich document ingestion and the extra dependency it hides

The v0.9.0 release is titled "Rich Document Ingestion & chunk_rag Baseline," and it addresses the most common complaint about extraction tools: they only read plain text. PDF, Word, PowerPoint, Excel, HTML and EPUB are supported, but not by default. The README is explicit that this requires an extra install, pip install "hyperextract[ingest]", and pyproject.toml shows why: the ingest extra pulls in markitdown with the pdf, docx, pptx and xlsx extras. The conversion backend is markitdown, and the dev dependency group notes that it is kept in dev so CI exercises it.

The practical consequence is that a default install and an ingest-enabled install behave differently on the same file. A PDF handed to a default install has no documented conversion path, and the README does not describe what error you get in that case. If your corpus is mostly PDFs, install the extra from the start rather than discovering the gap mid-pipeline.

The same release added chunk_rag, described as a zero-cost baseline. That is a useful control: it produces a chunk corpus rather than a graph, so you can compare what the graph structures actually add against plain retrieval over the same document before paying for the extra LLM calls a graph extraction requires.

Where Hyper-Extract is the wrong tool

The project is labelled Alpha. The classifier in pyproject.toml reads "Development Status :: 3 - Alpha," and the version numbering is still 0.x. That is a fair description of the risk: templates, CLI flags and output layout can change between minor releases, and the release list shows three versions shipped within a single week in September 2026. Pinning a version is not optional if you build on top of it.

Cost and latency are the second constraint. Every extraction calls an LLM, and every query calls an embedder. The README estimates DeepSeek at roughly $0.001 to $0.005 per page, which is a README figure rather than a measured one, but the shape of the problem holds: re-parsing a large corpus is a real bill, and he feed exists precisely so you do not have to. There is no documented offline mode. Without a provider configured, the tool has nothing to extract with.

The third constraint is the one the documentation is least clear about. The README states that he feed rolls old facts back automatically, but it does not document what happens when the update is ambiguous, when an entity disappears entirely, or how to undo a feed operation as opposed to a document removal. he remove --document is documented; rollback of an incremental feed is not. If your workflow needs an auditable, reversible history, treat that gap as unresolved until you test it.

Finally, the licence field is inconsistent. The README badge and pyproject.toml both say Apache-2.0, while the repository metadata reports NOASSERTION. The LICENSE file at the top level is the authoritative text and should be read directly rather than inferred from a badge.

Hyper-Extract versus GraphRAG and LightRAG

The README lists GraphRAG and LightRAG among the eleven-plus extraction engines the project ships, which makes them both alternatives and components. That is worth understanding before choosing. If you already run GraphRAG or LightRAG directly, Hyper-Extract's value is the layer above: a common CLI, a template system, provenance tracking and export, sitting over whichever engine you select.

The real difference in approach is the structure. GraphRAG and LightRAG build graphs of pairwise relations between entities. Hyper-Extract extends that to hypergraphs, where a single edge can join three or more nodes, and to spatio-temporal graphs, where a relation carries a place and a time. For a corpus of meeting notes, clinical case reports or multi-party contracts, a pairwise graph forces you to decompose a fact that is naturally n-ary, and you lose the fact that the participants acted together. A hypergraph keeps it intact.

The trade-off is maturity and ecosystem. GraphRAG and LightRAG have been used as standalone systems long enough that their output schemas and query patterns are widely discussed. Hyper-Extract is at 0.x and its nine structures mean the output shape depends heavily on which template you picked. If your downstream consumer expects one fixed schema, a single-engine pipeline is the safer choice. If you need n-ary relations and want the engine to be swappable, the abstraction is the point.

Maintenance, upgrades and licence

The repository is not archived, and the last push was on 2026-09-09. Releases v0.8.1, v0.8.2 and v0.9.0 all landed in early September 2026, which indicates a fast-moving codebase rather than a dormant one. Note that pyproject.toml declares version 0.10.1 while the newest release listed is v0.9.0, so the working tree is ahead of the tagged releases. If you install from PyPI, confirm which version you get before writing against its API.

Upgrade cost is dominated by two things. The release notes show feature-level changes between patch versions, not just fixes: v0.8.1 and v0.8.2 both concern scoped search and source tags, and v0.9.0 changed ingestion and added a baseline engine. That pace means reading the release notes before every bump. The dependency surface is also wide: langchain, langchain-community and langchain-openai are all direct dependencies, and LangChain's own release cadence will drive upgrades you did not ask for. The dev group pins pytest at 9.0.2 or newer, which suggests the test suite is expected to run against recent tooling.

On licensing, pyproject.toml declares Apache-2.0 and the README badge agrees, while the repository metadata reports NOASSERTION. Apache-2.0 permits commercial use and modification and includes a patent grant, but the inconsistency between the metadata and the declared field means you should read the LICENSE file and have your own counsel confirm the terms. Nothing here is legal advice. The optional extras pull in additional packages with their own licences, so an ingest-enabled install is not covered by the Apache-2.0 declaration alone.

Editorial conclusion

Adopt Hyper-Extract if you need hyperedges or spatio-temporal structures and are willing to run a 0.x alpha whose pyproject.toml still carries the Alpha classifier. Do not adopt it if you need a stable schema contract, an offline pipeline without an LLM key, or a documented rollback path for every failure mode. Before committing, run he parse on one real document with your chosen provider, then run he info ./output/ --sources and he remove --document to confirm that provenance and rollback behave the way your workflow needs. The package version in pyproject.toml is 0.10.1 while the newest listed release is v0.9.0, so check which one you actually install.

Frequently asked questions

What exactly does Hyper-Extract extract from a document?

It converts unstructured text into one of nine knowledge structures, chosen by the template you pass with -t. These range from raw chunk corpora and simple lists up to graphs, hypergraphs and spatio-temporal graphs.

What is the extract process in Hyper-Extract?

A reader converts the input file to text, a YAML template defines what to ask for, an LLM returns the structured result, and the output directory stores the structure plus embeddings for search. You then query it with he search.

Is Hyper-Extract the same as a Tableau extract?

No. Tableau extracts are a different product entirely. Hyper-Extract is a Python CLI and library named hyperextract that builds knowledge graphs and hypergraphs from text using LLMs.

What is Hyper-Extract?

It is a Python 3.11+ command line tool and library that turns documents into structured knowledge, with support for graphs, hypergraphs and spatio-temporal extractions. It installs from PyPI as hyperextract.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. yifanfeng97/Hyper-Extract on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yifanfeng97-hyper-extract.svg)](https://hysenlabs.com/projects/yifanfeng97-hyper-extract)