# AutoSchemaKG and the atlas-rag package: building knowledge graphs without a predefined schema

> AutoSchemaKG is an HKUST research framework that extracts triples with an LLM and then induces its own schema through conceptualization. The atlas-rag package is how you actually run it, and the pipeline is long, stateful and LLM-bound.

**HKUST-KnowComp/AutoSchemaKG** — This repository contains the implementation of AutoSchemaKG, a novel framework for automatic knowledge graph construction that combines schema generation via conceptualization.

- Repository: https://github.com/HKUST-KnowComp/AutoSchemaKG
- Website: https://hkust-knowcomp.github.io/AutoSchemaKG/
- Stars: 827 · Forks: 109
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/hkust-knowcomp-autoschemakg

## The problem AutoSchemaKG targets: schemas you did not have to write

Most knowledge graph pipelines start with a schema. Someone decides that Person, Organization and Event are the node types, that works_at and participated_in are the edge types, and that everything else is out of scope. That decision is usually made before anyone has read the corpus, and it is expensive to revise afterwards. AutoSchemaKG takes the opposite position. The repository describes a two-stage approach: first, knowledge graph triple extraction, where an LLM reads text and produces triples comprising entities and events; second, schema induction, where the framework generates the schema itself through conceptualization and, in the README's wording, creates semantic bridges between seemingly disparate information to enable zero-shot inferencing across domains. The audience is researchers and engineers who have text and want a graph, but do not want to commit to an ontology up front. The ATLAS family of graphs is the demonstration: ATLAS-Wiki from Wikipedia, ATLAS-Pes2o from academic papers and ATLAS-CC from Common Crawl, described as built without predefined schemas or manual intervention. Whether or not you ever reproduce a graph at that scale, the two-stage design is the part worth evaluating.

## How the extraction and conceptualization stages actually connect

The pipeline is sequential and file-backed, and the file layout is the architecture. Inside atlas_rag, kg_construction holds the extraction and conceptualization modules, llm_generator wraps the model, retriever and vectorstore handle the RAG side, and utils holds shared helpers. Extraction writes triples as JSON, which convert_json_to_csv turns into CSV. Concept generation then runs as a separate LLM pass over that intermediate data, writing a temporary concept CSV that create_concept_csv consolidates. Only after both passes does convert_to_graphml produce a networkx-readable graph. The important consequence is that the schema is derived from the triples, not the other way round. Concept generation sees what extraction produced and induces categories from it. That ordering is what makes the approach autonomous, and it is also why the two stages cannot be run in parallel or swapped: concept generation has no input until extraction has finished. The ProcessingConfig object carries the knobs that matter, including batch_size_triple for triple extraction and batch_size_concept for concept generation, plus max_new_tokens, max_workers and remove_doc_spaces. Those are separate batch sizes because they are separate passes with different prompt shapes, and tuning one does not tune the other.

## Installing atlas-rag and running a first extraction

The package is published on PyPI as atlas-rag, so installation is a single pip command. The README gives this as the install path and requires Python 3.9 or later per pyproject.toml.

```bash
pip install atlas-rag
```

If you intend to use NV-embed-v2 embeddings, there is a separate extra that pins transformers to a narrow range, >=4.42.4 and <=4.47.1. That pin is worth reading twice before you install it into an environment that already has a newer transformers, because the constraint is two-sided.

```bash
pip install atlas-rag[nvembed]
```

The README's own quick start builds a small graph from a sample document. It constructs an LLMGenerator around a client, here a local transformers pipeline, and a ProcessingConfig that points at example_data and selects input files by a filename_pattern. Note that the output_directory is set to import/<keyword> and that remove_doc_spaces is enabled to strip duplicated spaces from the document text.

```python
from atlas_rag.kg_construction.triple_extraction import KnowledgeGraphExtractor
from atlas_rag.kg_construction.triple_config import ProcessingConfig
from atlas_rag.llm_generator import LLMGenerator
from transformers import pipeline

model_name = "meta-llama/Llama-3.1-8B-Instruct"
client = pipeline("text-generation", model=model_name, device_map="auto")
keyword = 'Dulce'
output_directory = f'import/{keyword}'
triple_generator = LLMGenerator(client, model_name=model_name)
kg_extraction_config = ProcessingConfig(
    model_path=model_name,
    data_directory="example_data",
    filename_pattern=keyword,
    batch_size_triple=3,
    batch_size_concept=16,
    output_directory=f"{output_directory}",
    max_new_tokens=2048,
    max_workers=3,
    remove_doc_spaces=True,
)
kg_extractor = KnowledgeGraphExtractor(model=triple_generator, config=kg_extraction_config)
```

From there the run is four explicit calls in a fixed order. run_extraction and generate_concept_csv_temp both invoke the LLM, while convert_json_to_csv and create_concept_csv are conversions. The README comments mark which steps involve generation, which is a useful reminder that the cheap steps and the expensive steps are interleaved.

```python
kg_extractor.run_extraction()
kg_extractor.convert_json_to_csv()
kg_extractor.generate_concept_csv_temp(batch_size=64)
kg_extractor.create_concept_csv()
kg_extractor.convert_to_graphml()
```

What you should see after this is a directory under import/Dulce containing the triple JSON and CSV files, a concept CSV, and a graphml file. The README does not document what a failed run leaves behind, so if a generation step errors partway through, check the output directory before re-running rather than assuming the state is clean.

## Where the framework stops being the right tool

The failure modes here are mostly the failure modes of LLM extraction, and the framework does not hide them. Because both triple extraction and concept generation are model calls, the graph is only as consistent as the model is across batches. The same entity mentioned in two documents can come out with different surface forms, and the conceptualization pass is what is supposed to bridge that, but nothing in the README describes a deterministic canonicalization step or a fixed vocabulary the concepts must map into. If your downstream consumer needs stable node identifiers across runs, that stability is not something the pipeline guarantees by construction. Cost and latency follow directly from batch_size_triple, batch_size_concept and max_workers: raising max_workers buys throughput only if your LLM endpoint tolerates the concurrency. The repository does include EvaluateKGC, EvaluateFactuality and EvaluateGeneralTask directories, so quality measurement is clearly considered part of the project, but the README does not spell out thresholds at which a generated graph should be considered usable. There is also a dependency cost that has nothing to do with LLMs: neo4j, fastapi, uvicorn, sentence-transformers, transformers, datasets, accelerate and graphdatascience all come in as base dependencies of atlas-rag, so installing the package into an otherwise minimal Python environment pulls a substantial stack. If you only want triple extraction and nothing else, that footprint is worth weighing. And if your corpus is small enough that you can read it, hand-authoring the schema will be faster and more predictable than inducing one.

## AutoSchemaKG against fixed-schema extraction tools

The obvious alternative is a fixed-schema extractor: you define node and edge types, and the tool maps text onto them. The difference is not the model, it is where the schema comes from and when it exists. In a fixed-schema pipeline the ontology is an input, so extraction is a constrained labeling problem and the output shape is known before the first document is read. In AutoSchemaKG the ontology is an output of the second stage, which means the schema is only knowable after extraction completes, and a change to the corpus can change it. That is the trade: fixed-schema tools give you reproducibility and a stable contract at the cost of missing whatever categories you did not anticipate, while AutoSchemaKG lets categories emerge from the data at the cost of a schema that is a function of the run. The RELATED SEARCHES list around this project points at the same neighborhood of ideas, including KGGen, ODKE+, and work on graph memory for agents such as Zep and Graphiti. Those systems also build graphs from text, but the distinguishing claim in this repository is the schema induction stage and the ATLAS graphs built with it. If your requirement is a graph that fits an existing ontology you already publish, a fixed-schema extractor is the shorter path, and AutoSchemaKG is not trying to be that.

## Licence, maintenance and what an upgrade costs

The licence is MIT, stated both in the README's repository metadata and in pyproject.toml as license = {text = "MIT"}. That is permissive and imposes few obligations on downstream use, but it says nothing about the licence of the LLM you point the pipeline at, the embedding model you select, or the corpora you feed in. Those are separate questions and the repository does not answer them. On maintenance: the repository is not archived, and the last push was on 2026-04-29. The README's update log shows a pattern of incremental additions rather than a frozen release, including a 05/12 entry adding documentation for the atlas-rag package and example directory, a 05/07 entry adding batch generation and refactoring the codebase, and a 24/06 entry adding ToG, Chinese KG construction and separating the NV-embed-v2 transformers dependency. The published version in pyproject.toml is 0.0.5.post1, and no GitHub releases were retrieved, so pip is effectively the release channel. Upgrading means tracking that version number and reading the update log, because the refactor described in the 05/07 entry is the kind of change that moves internal module paths. The NV-embed-v2 pin is the sharpest upgrade hazard: transformers >=4.42.4,<=4.47.1 and sentence-transformers==2.7.0 will conflict with a newer transformers elsewhere in the same environment, and the repository's own decision to separate that dependency into an extra is a recognition of the problem rather than a fix for it.

## Conclusion

Adopt AutoSchemaKG if you have unstructured text, an LLM endpoint you are willing to pay for at scale, and a need for a graph whose schema you did not have to design by hand. Do not adopt it if you need a deterministic parser with a fixed output contract, or if you cannot run an LLM over your corpus repeatedly, because every extraction and every concept generation step is a model call. Before committing, verify three things in your own environment: that the LLMGenerator binds to the client you intend to use, that the NV-embed-v2 optional dependency resolves under your Python version, and that the output of convert_json_to_csv and create_concept_csv matches the column layout your Neo4j import scripts expect.

## FAQ

### How do I install AutoSchemaKG?

The package is published as atlas-rag and installs with pip install atlas-rag. The optional NV-embed-v2 support is a separate extra, pip install atlas-rag[nvembed], which pins transformers to a narrow version range.

### Does AutoSchemaKG need a predefined schema?

No. The repository describes it as fully autonomous knowledge graph construction without predefined schemas, using a two-stage approach of triple extraction followed by schema induction through conceptualization.

### Which Python versions does atlas-rag support?

pyproject.toml sets requires-python to >=3.9 and lists classifiers for Python 3.9 through 3.12.

### What are the ATLAS knowledge graphs?

ATLAS, short for Automated Triple Linking And Schema induction, is a family of knowledge graphs created through AutoSchemaKG. The README describes three variants, ATLAS-Wiki from Wikipedia, ATLAS-Pes2o from academic papers and ATLAS-CC from Common Crawl.

### What licence does AutoSchemaKG use?

The licence is MIT, declared in pyproject.toml as license = {text = "MIT"}. That covers the code, not the models, embeddings or corpora you run through it.

## Sources

- [HKUST-KnowComp/AutoSchemaKG on GitHub](https://github.com/HKUST-KnowComp/AutoSchemaKG)
- [Issues](https://github.com/HKUST-KnowComp/AutoSchemaKG/issues)
- [License: MIT](https://github.com/HKUST-KnowComp/AutoSchemaKG/blob/main/LICENSE)
- [Project website](https://hkust-knowcomp.github.io/AutoSchemaKG/)
- [README](https://github.com/HKUST-KnowComp/AutoSchemaKG/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hkust-knowcomp-autoschemakg
