yoheinakajima/instagraph: a Flask knowledge graph demo with two disagreeing manifests
Converts text input or URL into knowledge graph and displays
At a glance
- What is it?
- The graph generation is one function in main.py. The interesting problems are the dependency manifests, the lint configuration nothing calls, and a roadmap with credit dates on it.
- Who is it for?
- InstaGraph is a readable, small Flask demonstration of LLM function calling that happens to have a large audience, and its rough edges are all visible in the repository rather than buried in a release process. Clone it for the prompting pattern and the driver split, pin nothing, and read main.py before anything else.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One Flask app, one function, and a waitlist link
InstaGraph takes text or a URL and returns a colored knowledge graph. The repository description says it converts text input or URL into knowledge graph and displays, and the README supplies the mechanism, OpenAI GPT-3.5 driven by function calling. The README even points at the right file, telling readers that the generation logic is the function call parameters taking up half of main.py. That is a helpful piece of honesty and a good place to start reading.
The file listing is short enough to hold in your head: main.py, models.py, drivers/, templates/, docker/, .env.example, Makefile, pyproject.toml, requirements.txt, poetry.lock, replit.nix. One process, a package for the optional graph database backends, a template directory for the rendered page. Nothing in that listing suggests a job queue, a user account system, or a test suite, which lines up with the author telling contributors in the README that they usually code on the weekends or at night, in pretty small chunks.
The popularity numbers are the outlier. 3555 stars, 287 forks, 26 open issues, not archived, last pushed 2026-10-04. Against that, the releases list is empty and the topics list is empty, so there is no version to pin and no package index entry. The README opens by sending non-coders to a waitlist at instagraph.ai, which suggests the hosted version of this idea matters as much as the code.
Two manifests, three names for the same package
The install path in the README is four steps: clone the repository, move into the directory, install the requirements, then rename .env.example to .env and fill in the OpenAI key. The file you are told to install is requirements.txt, and that file has no project name in it at all. The file with a name is pyproject.toml, and it says name = python-template, version 0.1.0, with an empty description and a placeholder author line. So a poetry install produces a distribution whose own metadata calls it python-template, while every page of documentation calls it InstaGraph.
They disagree on Python too. pyproject requires python = >=3.10.0,<3.11, a single minor version, while the README only says you will need Python and pip installed. On 3.11 or 3.12 the poetry path refuses to resolve and the README never warned you.
The sharpest disagreement is the HTML parser. pyproject lists bs4 = ^0.0.1 and requirements.txt lists beautifulsoup4==4.12.2. Those are two different distributions, and the bs4 name is the well known placeholder package that will not import on modern Python. One manifest describes a package that cannot do the job the other one pins correctly. The openai and flask entries do agree, as caret ranges in pyproject and exact pins in requirements.txt at 0.28.0 and 2.3.3. Instructor disagrees outright at ^0.2.6 against instructor==0.2.8. On top of that, requirements.txt mixes runtime and development concerns in one list: neo4j==5.12.0, gunicorn==21.2.0, and FalkorDB==1.0.1 sit next to unpinned black, isort, flake8 and mypy entries. If you need one of the graph drivers, install by manifest and edit the file; do not trust either file as a published artifact.
The lint you get is not the lint the project configured
pyproject.toml configures two tools that nothing ever runs. There is a tool.ruff section selecting the E, W, F, I, B, C4, ARG and SIM rule sets with a short ignore list, and a tool.pyright section pointing at Microsoft's own configuration documentation. The Makefile has two targets and both skip ruff and pyright entirely.
black . --check
isort . --check --profile black
flake8 . --ignore=E501,E203,W503So the check you actually get is black, isort, flake8 and mypy, while the configuration sitting in pyproject.toml describes two linters that never fire. Before you trust a green run, decide which source of truth you want and delete the other. The Makefile also carries four lines of inline commentary explaining why E501, E203 and W503 are suppressed, each pointing at upstream documentation, which is the fingerprint of a scaffold that was adapted rather than a setup that was designed for this code.
There is a small irony in that, given the earlier section. A repository whose package metadata is still python-template, whose linter configuration is unused boilerplate, and whose contributor docs are unusually personal is a good illustration of how much of any small project's readme and manifests are inherited from a template rather than written for the project. None of that is disqualifying for a demo, and all of it matters the moment you try to depend on it.
The run command, and the flag the README explains well
Local startup is one command with three flags:
git clone https://github.com/yoheinakajima/instagraph.git
cd instagraph
pip install -r requirements.txt
python main.py [--graph neo4j|falkordb] [--port port] [--debug]Then you open http://localhost:8080, paste text or a URL, and submit. The container route is equally plain, with a development compose file and a production image built and run in the background:
docker-compose -f docker-compose.yml up --build -dThe flag worth reading closely is debug. The README states that it enables the Werkzeug debugger and binds to 127.0.0.1 only, and gives the reason: the debugger allows arbitrary code execution, so it is never exposed on other interfaces. That is correct, and it is more care than most small Flask demos take over their own debug mode. Treat it as settled and do not spend time re-deciding it.
What the README leaves open is the shape of production. The production compose file is described as using gunicorn==21.2.0, and that pin is visible in requirements.txt, but the README does not say how many workers it launches, what port it binds, or whether the debug flag is off in the container. Those are three questions you will answer yourself from the compose files, so budget for it. A replit.nix at the repository root also tells you the app was set up to run inside a hosted notebook environment at some point, which is a cheap way to get a first render before you install anything locally.
The endpoint list contradicts its own method labels
The API section is where the README disagrees with itself in the smallest possible way. It numbers three routes and labels the first two with an HTTP method that does not match the description underneath. GET Response Data at /get_response_data, and GET Graph Data at /get_graph_data, each followed by a line that says the method is POST. The third route, /get_graph_history, is a plain GET with no parameters.
That contradiction matters more than a typo usually would, because the documented request body is a JSON object with a user_input key carrying the text, and a GET request cannot carry a body. So the labels are wrong and the methods are right, and anything you write against those names will need the methods corrected before it works.
The graph database choice is a command-line flag rather than a configuration key, with neo4j or falkordb as the two values, and the drivers/ directory is where both live. That is a clean design for an app with one optional persistence path. Two things stay unstated: what happens with no --graph flag at all, and whether the chosen driver is required before /get_graph_data returns anything. The configuration surface is also asymmetric, with three Neo4j values against one FalkorDB value, which reads like an adapter that grew one backend at a time rather than one that was generalized up front. The .env.example adds a sixth variable, USER_PLAN, that the README never mentions anywhere, so it is the first string to grep for if you are trying to understand what this app thinks a user is.
A roadmap with credit dates, and no releases
The contributing section is the most candid part of the project, and the most useful. The author writes that there are a lot of build a chart tools out there, that instead of user accounts and custom charts they would rather work toward the largest knowledge graph ever, and asks for help running the repository. Then comes a checklist. Four items are struck through and credited to one contributor with the date 9/13/23: storing the knowledge graph, pulling the graph back from storage, and the ability to expand a graph.
Three crossed-out items landing on one day is consistent with the rest of the tree, and with a /get_graph_history endpoint existing at all. The still-open items are the interesting ones: combining two graphs, combining two or more graphs from history, expanding a graph from specific nodes, and fuzzy matching of nodes for combining graphs using a vector match plus an LLM confirmation. That last idea is a description of a real engineering problem, and it is the point at which this project stops being a prompt demo.
Those dates sit oddly against the rest of the record. The last push is 2026-10-04, so the code is being touched, while the README still carries the 9/13/23 credits and no new ones. Releases remain empty and topics remain empty. The practical read is a continuously edited main branch with nothing to pin against, which is fine to clone and awkward to depend on, and it explains the waitlist page: the hosted path is where the versioning discipline would have to live.
So how do you judge this project without picking a winner between the readme and the repository? Both facts are true at once. The README describes an ambitious graph merging tool and the repository is a 3555-star Flask demo whose package is still named for its scaffold. If you want the prompting pattern and the driver split, fork it and rename the metadata first. If you want the merging features, you are starting from main.py rather than from a library.
Editorial conclusion
InstaGraph is a readable, small Flask demonstration of LLM function calling that happens to have a large audience, and its rough edges are all visible in the repository rather than buried in a release process. Clone it for the prompting pattern and the driver split, pin nothing, and read main.py before anything else.
Frequently asked questions
Which model does InstaGraph use?
The README says GPT-3.5 from OpenAI, and requirements.txt pins openai==0.28.0, which is the pre-1.0 generation of the OpenAI Python SDK. Expect to modernize that call if you fork it, since the pinned SDK predates the current client layout.
Does InstaGraph need a graph database to run?
No. The --graph flag selects neo4j or falkordb and is described as optional, so the app runs from main.py with an OpenAI key alone. What the README does not state is what the no-flag default is, or whether the graph data endpoint needs a driver configured before it returns anything.
Which Python versions work with InstaGraph?
pyproject.toml requires python >=3.10.0,<3.11, so the poetry path only resolves on 3.10. The README just says Python and pip are needed and names no version, and requirements.txt pins no interpreter at all. On anything newer than 3.10 you are installing past what the project metadata declares.
Is it safe to run InstaGraph with the debug flag?
On your own machine, yes, and the README says why: the flag enables the Werkzeug debugger, which allows arbitrary code execution, so the app binds to 127.0.0.1 only and is never exposed on another interface. That is the intended use. Putting it on a shared or routable host is the case the README is warning about.
Does InstaGraph remember graphs between runs?
Partly. The README checklist has storing and pulling the graph back from storage struck through and credited to a contributor on 9/13/23, and a /get_graph_history endpoint is documented, which points at persistence existing. The open question is scope, because the driver is selected per process by a command-line flag, so what survives depends on which backend you wired in.