# kg-gen: Extracting Knowledge Graphs from Plain Text with LLMs

> kg-gen is a Python package that turns any string or message array into entities, edges and relations using a model you choose. It is aimed at engineers who want graph-shaped output for RAG, synthetic data or text analysis, and who accept that the quality of the graph is the quality of the model behind it.

**stair-lab/kg-gen** — [NeurIPS '25] Knowledge Graph Generation from Any Text

- Repository: https://github.com/stair-lab/kg-gen
- Website: https://kg-gen.org
- Stars: 1,282 · Forks: 200
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stair-lab-kg-gen

## What kg-gen actually produces, and who needs that

Most text-to-graph tools ask you to define a schema first. kg-gen inverts that. You hand it a string, it hands back four things: a set of entities, a set of edge labels, a set of relations as (subject, edge, object) triples, and optionally cluster maps that fold synonyms together. The README's first example is three sentences about a family, and the documented output is entities={'Linda', 'Ben', 'Andrew', 'Josh'}, edges={'is brother of', 'is father of', 'is mother of'}, and the matching relation triples.

The intended audience is narrow but real. The README names four uses: building a graph to assist with RAG, creating graph synthetic data for model training and testing, structuring arbitrary text into a graph, and analyzing relationships between concepts in a source text. If your problem is one of those four, the shape of the output is already what you want. If your problem is populating a pre-existing ontology with fixed predicates, this is the wrong starting point, because the edge labels are generated by the model, not chosen by you.

The package is Python only, requires Python >=3.10 and <4.0, and ships as version 0.4.0 under the MIT license according to pyproject.toml. It is published on PyPI as kg-gen.

## The mechanism: LiteLLM for routing, DSPy for structure

Two dependencies do the heavy lifting. Model calls are routed through LiteLLM, which is why the model argument is a string in the form {model_provider}/{model_name}. The README lists openai/gpt-5, gemini/gemini-2.5-flash and ollama_chat/deepseek-r1:14b as examples, and points at the LiteLLM provider documentation for the full format. That single string is the whole provider abstraction: swapping from a hosted API to a local Ollama model is a string change plus, for local models, no API key.

Structured output generation uses DSPy. That is the part that turns a free-form model response into typed entities, edges and relations rather than prose you have to parse. The README does not document the prompt templates, but pyproject.toml declares package data for kg_gen.prompts (*.txt) and kg_gen.utils (template.html), so the prompts and the visualization template ship inside the wheel.

For text longer than a single request, generate() takes chunk_size and cluster. The documented example passes chunk_size=5000 and cluster=True over a passage about machine learning, and the documented output includes entity_clusters mapping a canonical name to its variants, for example 'artificial intelligence': {'AI', 'artificial intelligence'} and 'neural networks': {'neural networks', 'neural nets', 'NN'}. Edge clusters work the same way, folding 'is a type of' and 'is a kind of' into 'is type of'. That clustering step is what makes the output usable as a graph instead of a pile of near-duplicate nodes.

Graphs are also composable. kg.aggregate([graph_a, graph_b]) merges multiple graphs, and kg.cluster() can be run afterwards on the combined result. In the README's example, two texts that refer to the same person as both Joe and Joseph produce an entity_clusters entry mapping 'Joe' to {'Joe', 'Joseph'}. This is the only documented mechanism for resolving cross-document identity, and it is a clustering heuristic, not an entity resolution system.

## Installing kg-gen and running a first extraction

The published package installs with pip. The README gives this as the quick start, and there is no separate setup step beyond having a model credential available, either in the environment or passed explicitly.

```bash
pip install kg-gen
```

If you want the repository itself, with the tests and the dev extras, the README says to clone it and install with the dev extra, then run the basic test from the root directory. The test writes a visualization to tests/test_basic.html, which is the fastest way to see what the library is actually doing before you point it at your own corpus.

```bash
pip install -e '.[dev]'
python tests/test_basic.py
```

A minimal extraction looks like this. The model string follows the LiteLLM provider format, and the API key can be omitted if it is already in the environment or if you are pointing at a local model.

```python
from kg_gen import KGGen

kg = KGGen(
  model="openai/gpt-4o",
  temperature=0.0,
  api_key="YOUR_API_KEY"
)

text_input = "Linda is Josh's mother. Ben is Josh's brother. Andrew is Josh's father."
graph_1 = kg.generate(
  input_data=text_input,
  context="Family relationships"
)
```

What you should see is a set of entity names, a set of edge labels, and relation triples. The README's documented output for that exact input is entities={'Linda', 'Ben', 'Andrew', 'Josh'} with edges={'is brother of', 'is father of', 'is mother of'}. If you get back prose instead, the model you selected is not returning structured output reliably and you should try a different model string.

To render the graph, the README documents a static method taking the graph, an output path and an open_in_browser flag.

```python
KGGen.visualize(graph, output_path, open_in_browser=True)
```

The .env.example file in the repository shows the environment variables the project expects for configuration: LLM_MODEL, LLM_API_KEY, LLM_TEMPERATURE and RETRIEVAL_MODEL, plus provider-specific keys for OpenAI, Anthropic and Gemini used in testing. Note that these names differ from the keyword arguments in the Python API, so do not assume the env file configures the constructor.

## The MCP server and the kggen command

pyproject.toml registers a console script, kggen, pointing at kg_gen.cli:main. The README documents one use of it: starting an MCP server so that AI agents can call kg-gen for persistent memory. The documented sequence is to install the package and then run the server.

```bash
pip install kg-gen
kggen mcp
```

The README says this is usable with Claude Desktop, custom MCP clients, or other AI applications, and links to the mcp/ directory for documentation. The MCP dependencies (fastmcp>=2.10.6 and mcp>=1.12.1) are declared as a separate dependency group rather than in the base install, so the exact install command for the MCP path is something the mcp/ documentation would have to confirm. The README does not show it.

This is worth flagging as a documentation gap. The command is advertised in the quick start, but the base dependency list does not include the MCP packages, and the README does not state whether kggen mcp fails gracefully or with an import error when they are absent. Verify before you build an agent workflow on it.

## Where kg-gen is the wrong tool

The output is model-generated, so it inherits every failure mode of the model. The README's own example of a temperature-0.0 configuration suggests the authors expect determinism to matter, but temperature zero is not a determinism guarantee across providers, and nothing in the README describes caching, seeded sampling or output validation beyond the DSPy structured-output layer.

Chunking is the second soft spot. The documented example uses chunk_size=5000 characters. Nothing in the README describes what happens to a relation whose subject appears in one chunk and whose object appears in the next. Clustering is offered as the remedy, and it operates on the entity and edge strings after extraction, so a relation that was never extracted in the first place cannot be recovered by clustering. For documents where the important facts span paragraph boundaries, this is a real constraint and the README does not address it.

The third issue is ontology. Because edge labels are generated, two runs over different corpora will produce different predicate vocabularies. If you need to load the result into a store with a fixed schema, you are doing a mapping step that kg-gen does not perform for you. The README lists Neo4j as a dependency in pyproject.toml, but the README body does not document a Neo4j export path, so the presence of the driver should not be read as a supported integration.

Finally, cost. Every entity, edge and relation in the output is the product of model calls. For large corpora the token spend scales with input length plus the clustering work, and the README gives no figures on call counts or token usage per document.

## How kg-gen differs from iText2KG and schema-first pipelines

iText2KG is the closest comparison in the related searches, and the difference is architectural. Schema-first pipelines such as iText2KG-style incremental construction ask you to define the entity and relation types up front, then the model fills them in. That gives you a stable vocabulary and a predictable graph shape, at the cost of doing the ontology design yourself and redoing it when the corpus changes.

kg-gen goes the other way. You supply text and a context string, and the model decides what the entities and edge labels are. The context parameter in the documented example ("Family relationships") is the only lever the README shows for steering that vocabulary. The payoff is that you can point it at an unfamiliar corpus with no schema work; the cost is that you own the reconciliation problem afterwards, which is what the cluster and aggregate methods exist to partially solve.

The second difference is provider neutrality. Because routing goes through LiteLLM, the same code runs against a hosted model or a local Ollama model by changing a string. Schema-first frameworks in this space are more often tied to a specific provider's function-calling interface. If you need to run extraction on-premises for data reasons, that string swap is the whole migration.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-03-24. That is roughly six months before today, which puts it at the boundary where you should check the commit history yourself rather than assume a cadence. The release history shows three tagged releases between 2025-10-06 and 2025-11-23: WikiQA-evaluations, MINE-deduplication-scikitlearn-vs-faiss, and MINE-evaluations-expanded. Those are evaluation and dataset releases rather than library releases, so the tag stream is not a reliable signal of library change frequency.

pyproject.toml declares license = "MIT", which is permissive and imposes no copyleft obligation on your own code. The dependency list is the real licence consideration: it includes sentence-transformers, scikit-learn, networkx, neo4j, dspy-ai and openai, each with its own licence, and the default clustering path pulls in sentence-transformers, which brings model weights with their own terms. If you are shipping a product, audit the transitive set rather than the top-level MIT declaration. This is not legal advice; check with whoever handles licensing on your side.

The upgrade surface is broad. The pinned floors include pydantic>=2.0.0, scikit-learn>=1.7.2, sentence-transformers>=5.1.0 and dspy-ai>=3.0.4. dspy-ai in particular has moved quickly across major versions, and since DSPy is the layer producing structured output, a major bump there is the most likely source of a breaking change in what generate() returns. Pin your versions and test after any dspy-ai upgrade.

## Conclusion

Adopt kg-gen if you already have a model endpoint you trust and you want entity, edge and relation triples out of plain text without writing prompt plumbing yourself. Do not adopt it if you need a fixed ontology, a deterministic output, or a guarantee that the same input produces the same graph on every run. Before committing, run python tests/test_basic.py after pip install -e '.[dev]' and inspect tests/test_basic.html, then check whether the entity_clusters it produces match the vocabulary your downstream store expects.

## FAQ

### Is the knowledge graph still relevant?

kg-gen's README answers this indirectly by naming four concrete uses for the graphs it generates: assisting RAG, creating synthetic graph data for model training and testing, structuring arbitrary text, and analyzing relationships between concepts. Whether that matters to you depends on which of those four you are doing.

### Is knowledge graph part of AI?

In kg-gen's case the connection is direct: the package extracts knowledge graphs from plain text using language models routed through LiteLLM, with DSPy handling structured output generation. The graph is the output of an AI pipeline, not a hand-curated artifact.

### Is Neo4j a knowledge graph?

kg-gen lists neo4j>=5.0.0 as a dependency in pyproject.toml, but the README body does not document a Neo4j export path. Neo4j is a graph database, so it stores graphs rather than being one; the README does not describe how kg-gen output would be loaded into it.

## Sources

- [Issues](https://github.com/stair-lab/kg-gen/issues)
- [Project website](https://kg-gen.org)
- [README](https://github.com/stair-lab/kg-gen/blob/main/README.md)
- [Releases](https://github.com/stair-lab/kg-gen/releases)
- [stair-lab/kg-gen on GitHub](https://github.com/stair-lab/kg-gen)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stair-lab-kg-gen
