GraphGen ships as graphg, and its package data is keyed to the wrong case
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
At a glance
- What is it?
- InternScience/GraphGen generates knowledge-graph-guided synthetic training data and fine-tunes with LLaMA-Factory or xtuner, but its install name is graphg, its badges still point at a renamed organisation, its rephrase comparison trains the baseline for two epochs against one, and it has no .dockerignore next to a COPY . . in the Dockerfile.
- Who is it for?
- GraphGen fits a team that has a source corpus, knows which knowledge a model is missing, and wants generated QA pairs rather than a scraped instruction set. Four things to check first.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 49 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The install name is graphg and the links still point at open-sciencelab
The repository is InternScience/GraphGen. Almost every link inside the README points somewhere else.
The badges, the issue links, and the project links all address github.com/open-sciencelab/GraphGen, and the setup metadata repeats that URL and lists the author as open-sciencelab with an address at pjlab.org.cn. That is the organisation the project lived under, and the repository now sits under a different one. The links still work, so this is a leftover rather than a break, but it means a reader following the badges lands on a path that is not the one they came from.
The package name is the second mismatch. The PyPI link in the README points at a project called graphg, and the distribution name in the packaging script is graphg. So the project is GraphGen, the repository is GraphGen, and the thing you install is graphg.
The releases follow a third convention. The two tags are v0.1.0.post20250930 and 20250422, a PEP 440 post-release carrying a date and a bare date used directly as a version. Two tags, two schemes, from a project whose version is read at build time out of a Python file.
The baseline trained for two epochs and both rephrase runs trained for one
The rephrase pipeline is the part of the README with numbers in it, and the numbers have a structural caveat.
The setup is a small model trained from scratch: Qwen3-0.6B on SlimPajama-6B. The baseline row is SlimPajama-6B trained for 2 epochs, averaging 24.24 across seven columns. The two rephrase rows are one epoch each. Executive-Summary Rephrase averages 25.56, marked as a gain of 1.32, and Cross-Domain Rephrase averages 25.14, marked as a gain of 0.9.
So the baseline saw the data twice and the rephrase variants saw it once, and the comparison is stated as a like-for-like data quantity because the point is that the rephrase adds variants without adding data. The claimed mechanism is that the gains come from how the same knowledge is expressed rather than from more of it.
That is a legitimate experiment design and also an unusual one to read as a table. A reader who does not notice the epoch column will take away that one epoch of rephrased data beats two epochs of the original by 1.32 points, when what the table shows is one epoch of rephrased data beating two epochs of the original. Running the baseline for one epoch would separate the two effects, and that row is not in the table.
Cross-Domain Rephrase is worse than the baseline on two of seven columns
Reading the individual columns tells a different story from the averages.
Cross-Domain Rephrase has the best ARC-E score in the table at 28.79 against the baseline's 25.55, and the best TruthfulQA-MC2 at 52.41 against 49.90. Those are the two wins the boldface marks.
It also loses two columns outright. ARC-C is 20.22 against the baseline's 21.08, and GSM8K is 0.00 against the baseline's 0.08. A rephrase method that scores exactly zero on grade school maths is a detail worth knowing before you adopt it for reasoning-heavy data.
Executive-Summary Rephrase is the more uniform of the two. Its worst column is ARC-C at 22.70, which is also the best ARC-C in the table, and its GSM8K score of 1.36 is the best of the three rows. It is ahead of the baseline on all seven columns.
The final SFT table follows the same pattern of reporting a loss. On SeedBench for the plant domain, the graph-generated data scores 65.9 against 51.5 for the Qwen2.5-7B-Instruct baseline. On CMMLU for common knowledge, it scores 73.6 against 75.8, so the baseline wins that one by 2.2 points.
The changelog names rocksdb and the requirements file installs rocksdict
The 2025.12.16 entries added two storage backends: a key-value store and a graph database. The same entry, and a companion one, also added a local inference backend and refactored the generation pipeline onto a distributed runtime.
The requirements file has a storage section with two entries. The key-value one is rocksdict, and the graph one is kuzu.
rocksdict is a Python binding for RocksDB, not RocksDB itself, so the changelog entry names the project rather than the package that actually gets installed. That is a small imprecision, and it is the kind that costs someone a search when they are trying to work out what they just pulled in.
The graph entry has no such ambiguity. Kuzu is an embedded graph database, and it is installed unconditionally alongside the key-value store rather than as an extra, so a small deployment gets both engines whether or not it uses either.
Counting the graph-shaped dependencies gives four: networkx, graspologic, igraph, and the rdflib triple store, on top of kuzu. There is also a separate cluster of three community detection libraries, leidenalg, igraph, and python-louvain, installed for a single feature that the changelog describes as community detection using the Leiden algorithm.
Package data is keyed on GraphGen while the package directory is graphgen
The packaging script reads the version by opening graphgen/_version.py and executing its contents, then pulling __version__ out of the resulting namespace. That works and it means the version lives in a Python file rather than in the packaging metadata.
The requirements reader does something more surprising. It reads requirements.txt line by line, skips lines starting with a comment character, and also skips any line containing the word textract. So one dependency is filtered out at packaging time by a substring match, with no comment explaining why in the file itself.
Then there is the case mismatch. The package data mapping is written with a capital G, keying on GraphGen and pointing at a configs glob, while the package directories come from find_packages over a directory named graphgen. On a case sensitive filesystem those are two different names, and the packaged wheel may not carry the configuration files at all. On macOS, where the filesystem is case insensitive by default, it works, which is exactly the kind of bug that survives.
The find_packages call also carries an exclusion for a bare models name, while the inference clients live under a nested path inside the graphgen package.
No .dockerignore sits next to a COPY . . in the Dockerfile
The container build is competent in most of its choices. It starts from python:3.10-slim, disables bytecode writing and unbuffered output, installs git and a compiler, creates a non-root user named appuser, installs the requirements in their own layer, creates the cache and log directories, switches to the unprivileged user, and runs the web interface with python webui/app.py on port 7860.
Two details undercut the tidiness.
There is no .dockerignore in the repository. The top level has a .gitignore, an .env.example, and a uv.lock, but nothing that tells Docker what to leave out. So the copy step takes the entire working tree into the image: caches, virtual environments, local logs, and any .env file sitting in the checkout, since the example file shows that is where configuration goes.
And the copy is a whole-tree copy placed after the dependency layer. The requirements are installed in their own step to get the layer cache, but then every source file is copied in one instruction, so any change to any file, or to the local state, invalidates the layer above the install. The layer caching works for the dependencies and not for the code, which is the half where it usually matters.
Six model variables default to empty strings
The Dockerfile declares six environment variables for the application's configuration and every one of them is set to an empty string: SYNTHESIZER_MODEL, SYNTHESIZER_BASE_URL, and SYNTHESIZER_API_KEY, then TRAINEE_MODEL, TRAINEE_BASE_URL, and TRAINEE_API_KEY.
The two roles are the design. The synthesizer is the model that writes the questions and answers. The trainee is the model whose knowledge is being measured to find the gaps. The framework's central claim, that it identifies what a model does not know using the expected calibration error metric, needs both: a generator and a model to test.
Shipping them empty means the container starts with nothing configured and expects you to supply both sides. For a local inference setup that is reasonable, since the changelog records local backends including a local inference engine and four client wrappers. For the default deployment it means an image that boots to an interface with no model behind either half of the pipeline.
The empty defaults also mean there is no built-in signal that configuration is missing, because the value the program sees is the same shape as a value that was deliberately set to nothing.
pyproject.toml holds formatter settings and nothing else
There is a pyproject.toml in this repository and it is not a build file.
It has two tables, one for black and one for isort. Black is set to a line length of 88 with Python files included, and isort follows the black profile at the same line length with a trailing comma, vertical hanging indent, grid wrap disabled, parentheses, and a newline before comments.
There is no build-system table, no project table, and no version. The version comes from graphgen/_version.py, the dependencies come from requirements.txt through the packaging script, and the metadata comes from the same script.
The comments in those tables are in Chinese, four of them, annotating each setting with why it was chosen: that 88 is the black default, that the isort profile is set to match black in one step, that the line length is kept consistent with black, and that the multi-line output style is black's preferred bracket wrapping.
There are two dependency mechanisms in play as well, since the repository carries both requirements.txt with requirements-dev.txt beside it and a uv.lock. That is workable, but a file named pyproject.toml that configures only the formatter is one more thing to check before you assume which install path the project uses.
Editorial conclusion
GraphGen fits a team that has a source corpus, knows which knowledge a model is missing, and wants generated QA pairs rather than a scraped instruction set. Four things to check first. That the rephrase numbers are what you need, because the baseline trained for two epochs and both rephrase runs trained for one, so the comparison is epochs against epochs as much as wording against wording. Which case your configs land in, since the package data key does not match the package directory name. What is in the container, since there is no ignore file and the whole working tree is copied. And which storage backend you want, since the changelog and the requirements file name different projects for the same role.
Frequently asked questions
What is GraphGen and how does it work?
It is a framework for knowledge-graph-guided synthetic data generation. It constructs a fine-grained knowledge graph from source text, identifies knowledge gaps in a model using the expected calibration error metric, prioritises QA pairs for long-tail knowledge, samples multi-hop neighbourhoods, and applies style control to diversify the output. Fine-tuning is then done with LLaMA-Factory or xtuner.
How do I install GraphGen from PyPI?
The package is published as graphg rather than under the GraphGen name, and the packaging script sets that as the distribution name. The version is read at build time by executing graphgen/_version.py, and the classifiers declare Python 3.10, 3.11, and 3.12.
What does the GraphGen rephrase pipeline claim, and how was it tested?
It adds LLM reformulation to generate diverse variants of the same corpus rather than repeating it. The setup trains Qwen3-0.6B from scratch on SlimPajama-6B, with the baseline trained for 2 epochs averaging 24.24, Executive-Summary Rephrase trained for 1 epoch averaging 25.56, and Cross-Domain Rephrase trained for 1 epoch averaging 25.14.
Which storage backends does GraphGen use?
The requirements file installs rocksdict for key-value storage and kuzu as the graph database. The update history describes adding RocksDB and Kuzu on 2025.12.16, along with a local inference backend and a refactor of the generation pipeline onto Ray.
What input formats can GraphGen generate data from?
Plain source text, PDF input added via MinerU, and HuggingFace Datasets added as an input source. It also generates VQA data with bash scripts/generate/generate_vqa.sh, supports benchmark synthesis for single choice, multiple choice, fill-in-the-blank, and true-or-false items, and can search NCBI, RNAcentral, UniProt, Google, Bing, and Wikipedia.
How is the GraphGen web interface configured?
The Dockerfile sets six environment variables for two model roles and leaves all six empty: SYNTHESIZER_MODEL, SYNTHESIZER_BASE_URL, SYNTHESIZER_API_KEY and TRAINEE_MODEL, TRAINEE_BASE_URL, TRAINEE_API_KEY. It exposes port 7860 and starts the app with python webui/app.py as a non-root user.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internscience-graphgen)