sift-kg lets the model design the schema, and its packaging still points at another repository
Turn any collection of documents into a knowledge graph. Extract entities and relationships via LLM, deduplicate with your approval. Map domains, find hidden connections, spot patterns across documents — knowledge that persists and compounds, for you and your AI agents. All from the CLI.
At a glance
- What is it?
- A CLI that turns a folder of documents into a NetworkX graph, with LLM schema discovery, human-approved entity merges and a browser viewer. The design decisions worth knowing are that the first extraction is not reproducible unless you keep the generated schema, that two of four claimed providers have a key slot and two do not, and that the project metadata names a different repository.
- Who is it for?
- sift-kg fits a corpus you own and want to interrogate rather than chat about, especially where provenance matters, since every entity and relation links back to the document and passage it came from. Four things to settle before the first run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 146 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The manifest points its four project URLs at another repository
The packaging metadata is wrong in a way that will send people to a different place. The project is `sift-kg`, published by one author and hosted at one GitHub account, but all four URLs under the project URLs table point at `github.com/civictable/sift-kg`: the homepage, the documentation link, the repository link and the issues link.
That is the set of fields a package index shows a user, so `pip show` and most index pages will render a homepage and a bug tracker that do not belong to this project. Anyone filing an issue from the package page lands somewhere else, and anyone following the documentation link gets a repository that is not the one they installed.
The rest of the manifest is more careful. The distribution is built with hatchling, the version reads 0.9.0, the licence is MIT, and the classifier says Development Status 3 Alpha, which is the honest label for a tool whose schema generation and entity resolution are both model-driven. Python 3.11 is the floor, and the classifiers list 3.11 and 3.12 only. The console entry point is a single command, `sift`, bound to the CLI app object, and the lint line length is set to 100.
The model designs the schema, and you have to keep the file it writes
The default is schema-free, and that is the most consequential choice in the tool.
A single LLM call samples the documents and designs an entity and relation schema tailored to that corpus. The result is written to a discovered domain file, and it is meant to be reused and edited rather than regenerated. The alternative is to pick a fixed schema from a bundled set, or to write your own YAML.
The bundled names are `schema-free`, `general`, `osint` and `academic`. The listing command prints what is available, a single run can take a domain flag such as `--domain-name osint`, and setting a domain in `sift.yaml` avoids passing it every time:
domain: academicA path to a custom domain YAML works in the same slot, and there is a separate environment variable for one as well.
The consequence for reproducibility is the part to sit with. If the schema is designed per corpus, then two runs over the same documents can produce different entity types unless the generated file is kept and reused. That file is therefore not a cache, it is an input, and it belongs in version control next to the documents if you want the graph to be rebuildable later.
Entity merges are proposed by the model and applied only after you approve
Deduplication is split across three commands, and the middle one is a person.
The first finds duplicate entities, the second is an interactive terminal interface where merges are approved or rejected, and the third applies the decisions. Nothing is merged automatically, and the stated position is that you control what gets merged.
That is the right default for entity resolution, where a wrong merge silently fuses two real things for the rest of the graph's life, and the stated position is that the graph is yours. The cost is that the step is blocking and interactive: a review session waits for a human, so the pipeline is init, extract, build, resolve, review, apply, and it stops at the fifth command until someone answers. In an unattended run there is no documented way to pre-approve a merge policy, so the graph either stays unresolved or you accept the duplicates.
Provenance is what makes review possible in the first place. Every entity and relation links back to the source document and the passage it came from, and the viewer exposes that as a source document filter, so a reviewer can check a proposed merge against the text that produced both sides.
It runs locally, unless you turn on the cloud OCR backend
The privacy claim and the option list need to be read together.
One feature says the run is local and the documents stay on your machine. The extraction engine is local, and OCR defaults to a local engine: Tesseract is the default, with EasyOCR and PaddleOCR as alternatives selected by a flag. Extraction runs through Kreuzberg, which handles a long list of document formats, and the graph itself is NetworkX in memory and then JSON.
But the same feature list offers Google Cloud Vision as an OCR fallback through a separate backend flag. Choosing it uploads the pages being read to a hosted service, and it is not installed by default: the Google client lives in an optional dependency group alongside a PDF library, so you add that extra before the option exists.
Model choice matters in the same way. Ollama is supported for local inference, and OpenAI, Anthropic and Mistral for hosted ones, all through LiteLLM. So the documents stay local only if the OCR backend and the model are both local, and neither half is forced.
Two provider keys have a slot, and two of the claimed providers do not
The example environment file is where the configuration surface is actually pinned down, and it is smaller than the feature list.
There are two key variables, one for OpenAI and one for Anthropic, plus a default model written in a provider-qualified form, with examples given as `openai/gpt-4o-mini`, `anthropic/claude-haiku` and `ollama/llama3.3`. That format is the LiteLLM convention, and it is why any compatible provider can be selected through the model string.
What is missing is a documented place to put a Mistral key, even though Mistral appears in the list of supported providers. Anyone using it is left constructing the variable name themselves. The same file defines the output directory, defaulting to `output`, and an optional domain path for a custom schema.
Cost is the other control surface, and it is a single flag: a maximum cost cap on LLM spending. That is worth pairing with the pipeline shape, because a run makes at least the extraction calls plus the schema discovery call, and the cap is the only ceiling offered rather than a per stage budget.
Any language in, one English graph out, and names get romanized
The multilingual promise is specific in a way that is worth checking against your corpus.
Extraction runs on documents in any language, and the output is a single graph in English. Proper names are left as they appear, while non-Latin scripts are romanized automatically, which is a job the `unidecode` dependency does.
The point to think about is the collision surface. Romanization is a many-to-one operation, and proper names that stay untranslated now sit in the same namespace as names that were transliterated. If two different names in two scripts reduce to the same Latin string, the graph cannot tell them apart, and the deduplication step then has to choose between merging them and leaving them as false twins. Nothing in the described flow says the pre-romanization form is kept alongside the transliterated one, and the export formats, GraphML, GEXF, CSV, SQLite and native JSON, are all downstream of the transform.
If your documents are in several non-Latin scripts, the safe move is to spot check the extracted names before trusting the graph, since that check is cheaper than discovering it after building a viewer full of near identical nodes.
Four JSON commands and a bundled skill make the graph an agent memory
The agent-facing surface is a set of commands that print JSON, plus a skill file.
sift topologysift query "topic"sift search "X" --jsonsift info --jsonThe first gives a structural overview, the second an entity neighbourhood subgraph, the third an entity lookup and the fourth project statistics. The design claim is durability: the graph persists across sessions and grows incrementally, so you extract new documents into the same output directory and rebuild, and deduplication is what keeps it coherent as it grows. The stated benefit is that an agent stops starting from zero after a context reset.
The teaching component ships in the repository at a skill file, `.agents/skills/sift-kg/SKILL.md`, and covers session orientation, entity exploration, reasoning across separate islands of knowledge, and grounded suggestion generation. That last item is the one to watch, because grounded suggestions from a graph are only as good as the provenance, which is the part the tool does commit to.
The human viewer is the other half: coloured community regions, hover previews, a focus mode that isolates a neighbourhood on double click, arrow key navigation, a trail breadcrumb of visited nodes, and pre-filter flags for neighbourhood, top nodes, community, source document and minimum confidence.
Editorial conclusion
sift-kg fits a corpus you own and want to interrogate rather than chat about, especially where provenance matters, since every entity and relation links back to the document and passage it came from. Four things to settle before the first run. Keep the generated domain file, because the schema is designed per corpus and an extraction is not reproducible without it. Decide which model runs, since a local Ollama model and a hosted one give very different answers for the same documents. Count the passes, because schema discovery is an extra call on top of extraction and `--max-cost` is the only ceiling offered. And pick a merge policy deliberately, since `sift review` is an interactive terminal step that will not run inside an unattended job. One caveat to check at the packaging level: the project URLs in the manifest name a different repository, so the issue tracker and the documentation link in your package metadata will not lead here.
Frequently asked questions
What does sift-kg do with a folder of documents?
It extracts text locally, has an LLM design an entity and relation schema from a sample of the corpus, extracts entities and relations, builds a NetworkX graph, proposes duplicate entity merges for your review, and generates a narrative summary. The graph can then be explored in a browser viewer or exported as GraphML, GEXF, CSV, SQLite or native JSON.
Do I have to use a hosted LLM to run sift-kg?
No. Ollama is supported for local inference, and the model is set as a provider-qualified string such as ollama/llama3.3. OpenAI, Anthropic and Mistral are also supported through LiteLLM. The example environment file has key variables for OpenAI and Anthropic only, so a Mistral key has no documented variable name.
Can sift-kg read scanned PDFs without sending them anywhere?
Yes, with the default settings. OCR defaults to Tesseract, with EasyOCR and PaddleOCR available as alternatives, all local. There is an optional Google Cloud Vision fallback selected by a backend flag, which does upload pages to a hosted service, and its client library is in an optional dependency group rather than the base install.
How do duplicate entities get merged in sift-kg?
In three steps: a command finds duplicate entities, an interactive terminal interface shows them for approval or rejection, and a third command applies your decisions. Nothing merges automatically. Because the review step waits for a person, the pipeline is not runnable unattended without a pre-agreed policy.
What language does the sift-kg knowledge graph use?
One unified English graph regardless of the input language. Proper names are kept as they appear, and non-Latin scripts are romanized automatically, which is why a name in another script and its transliteration share a namespace in the output.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/juanceresa-sift-kg)