Open-source project
MazzaWill/neo4j-python-pandas-py2neo-v3 avatar
MazzaWill/neo4j-python-pandas-py2neo-v3

Pinned to a 2016 certificate bundle, with a roadmap that shipped in ten minutes

Excel-to-Neo4j knowledge graph examples: legacy py2neo v3 plus modern Neo4j GraphRAG/vector search.

580 stars187 forksPythonMIT

At a glance

What is it?
This repository teaches building a knowledge graph from a spreadsheet, and it deliberately keeps the original stack broken: Python 3.6, pandas 0.23.4, py2neo 3, Neo4j 3.x, with every transitive dependency pinned to an exact version including a certificate bundle from 2016. Alongside the legacy path sits a modern GraphRAG example whose demo embeddings are deterministic placeholders, and a roadmap whose final two versions shipped ten minutes apart.
Who is it for?
Use the legacy half if you are learning, and use the modern half if you are building. The separation is deliberate and documented rather than accidental, and the modern example is the part that will run on a current interpreter.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 34 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The requirements pin the certificate bundle to 2016 and build tooling to 2018

The dependency file is a snapshot of one working environment, and it pins everything with an exact version rather than a compatible range. The direct dependencies are the interesting part: pandas at 0.23.4, numpy at 1.15.3, the py2neo library at version 3, the Neo4j driver at 1.6.2, and `xlrd` at 1.1.0 for reading the spreadsheet. The jupyter stack is present in full, which means the project was developed in notebooks.

The transitive pins are where the practical cost sits. `certifi` is fixed at a 2016 release, `wincertstore` at 0.2, `urllib3` at 1.24, `Click` at 7.0, `attrs` at 18.2.0, `tornado` at 5.1.1, `pyzmq` at 17.1.2, `ipykernel` at 5.1.0, `jupyter-console` at 6.0.0. Pinning a package manager's own dependencies is normal in a reproducibility lock, but pinning them in a flat requirements file that people are told to install with a single command is not: it means the install resolves the entire notebook stack to versions from 2017 and 2018.

Two consequences follow. On a current Python interpreter, the older messaging and notebook packages have no matching wheels, so the install attempts to build them from source and is likely to fail on a toolchain that no longer supports those versions. And a 2016 certificate bundle means the install depends on trusting certificate authorities as they were configured years ago, which is a real problem against any index that now uses a newer root.

The readme is honest about all of this. It calls the pins intentionally legacy and points at a tracking issue for dependency and security modernisation.

The repository name is the dependency pin, and it says so

The repository is called `neo4j-python-pandas-py2neo-v3`, and the readme records the same string as the project's legacy slug. The name encodes the driver, the dataframe library and the graph library major version, in that order, with the major version attached to the last one.

That is unusual, and it cuts both ways. It means the URL itself states which stack the examples target, so a search result or a bookmark tells you the vintage before you read a line. It also means the project cannot be modernised without either renaming or accumulating a contradiction, since a repository called py2neo v3 that recommends the official driver is asserting two things at once.

The readme resolves it by splitting the project into two paths and saying which is which in the opening paragraph. The legacy path exists for learners, notebooks and existing projects, built on py2neo v3 against Neo4j 3.x. The modern path is additive and lives in its own directory, using the current application stack. The repository-level scripts stay on the legacy path.

The stated original environment is specific enough to be checkable: Python 3.6.5, Windows 10, Neo4j 3.x and py2neo 3. Pinning an exact patch of Python is the kind of detail that usually signals a real machine someone once got it working on, rather than a documented support matrix.

The branch is called master rather than main, and there is no homepage recorded, so the repository is the whole of the project's presence.

A hard coded chdir of xxxx and a placeholder connection string

The quick start is honest in a way most tutorials are not. You install the requirements, you have a sample spreadsheet at the repository root, and then there are two edits to make before the main script will run.

The first is a working directory call in the source that contains the literal placeholder `xxxx`. The instruction is to replace it with your repository path, or to run the script from the repository root so the relative form works. A placeholder that is four letters long, sitting in a directory change, is the kind of thing that survives in a tutorial for years because everyone overwrites it and nobody removes it.

The second is the database connection. The graph construction class contains a placeholder `Graph(...)` call that you replace with your server URL and credentials. In the py2neo v3 API that constructor is where the endpoint and authentication live, so this is the single point every example depends on.

The readme also flags that the current roadmap includes replacing those hard coded paths with configurable arguments, which means the project knows about them and has not done it yet.

The rest of the layout is small and legible. One script reads the spreadsheet, extracts node and relationship rows and writes them to the database. A class module wraps the node and relationship creation. A third script goes the other way, pulling relationship data back out of the graph and converting it into matrices for downstream machine learning experiments. The extraction functions that produce the node and relationship dataframes are named in the walkthrough as the pair you would modify to map a different spreadsheet.

A committed bytecode cache directory sits at the repository root alongside these files.

The modern example uses deterministic local embeddings so a demo needs no key

The additive example is the part worth running, and one design decision in it deserves attention.

It keeps the same invoice dataset concept and swaps the stack: the official Neo4j Python driver instead of py2neo, vector indexes from Neo4j 5 and later, GraphRAG style semantic retrieval, and optionally the production embeddings from the `neo4j-graphrag` package. Alongside those sits local deterministic embeddings, described as being for keyless demos and continuous integration.

That is the right call and it changes what the demo proves. Deterministic local vectors mean the example runs with no API key, no network and no billing, so it works in a CI job and in a fresh clone. It also means the retrieval quality you observe in the demo is not evidence of retrieval quality in production. Nothing in the readme claims otherwise, and the word deterministic is doing honest work: the same input yields the same vector every run, which is exactly what you want for a test and exactly what you do not want from a semantic model.

So the two embedding sources are not two quality levels of the same thing. One is a reproducible placeholder, the other is a real model. If you are evaluating this example, you need the production path, and the difference is worth knowing before you conclude that a vector index plus a graph retrieval step performs well or badly on your data.

The entry point is a module run with an input path, a row limit and a bare `dry-run` token:

bash
python -m examples.modern_invoice_graphrag.app \
  --input examples/modern_invoice_graphrag/sample_invoice_rows.csv \
  --limit 2 \
  dry-run

The dry run token is positional rather than a named flag, and it comes after the options, which is worth knowing if you script it.

The roadmap describes two versions that shipped ten minutes apart

The readme carries a roadmap with three version ranges, written as future work.

The first, version 0.2, covers restoring maintenance: triaging historical issues, documenting the known working legacy environment, adding setup notes for Windows, local paths and credentials, and publishing maintenance restart releases. The second, 0.3, covers reproducible examples: the modern GraphRAG and vector search invoice example, replacing hard coded paths with arguments, smoke tests for Excel extraction and relationship dataframe generation, and better issue templates. The third, 0.4, covers modern compatibility: evaluating newer Python and pandas support, documenting the migration path from py2neo v3, and adding continuous integration checks for the legacy environment where practical.

The release history says all of it has happened. v0.3.0 was published at 04:20 on 2026-06-01 and is titled as the modern GraphRAG example. v0.3.1 followed six minutes later, titled as project positioning polish. v0.4.0 followed ten minutes after that, titled as the Neo4j knowledge graph skill. So two of the three roadmap ranges are published, and both of them shipped in the same ten minute window, one as the example, one as the skill, one as a rename of the release titles rather than a body of work.

The version 0.2 range has no release at all, so the maintenance restart releases it promised are absent while the later versions exist. The roadmap text in the readme was never updated to reflect any of this.

What is genuinely done is visible in the tree. A skills directory holds an agent skill for designing a Neo4j graph from spreadsheet data, generating safe Cypher, and choosing between the legacy library and the current driver, and it includes a reusable script for profiling tabular data before you model it. A tests directory exists, and the example directory holds the modern example with its own sample rows.

The maintenance policy is to keep the baseline broken by modernity

The project states its own status in a section titled maintenance status, and the statement is a policy rather than a boast. It says the project is maintained again as of a specific month, that the goal is to keep the original py2neo and Neo4j 3.x example usable for learners, notebooks and legacy projects while adding a current driver, vector index and GraphRAG example, and that the repository is intentionally maintained as a legacy educational example where modern versions of Python, pandas, Neo4j and py2neo may require code changes.

The reason given for keeping the old baseline intact is that modernisation work is tracked separately so the legacy baseline stays clear. That is a defensible choice and a rare one. Most projects in this position quietly upgrade, which destroys the thing learners were reading, or quietly abandon, which strands it. This one names the tradeoff and moves the modernisation to its own place.

The governance around it is thorough for an educational repository: a licence, contributing guide, security policy, support policy, code of conduct and a pull request template, all listed in the readme rather than left to be discovered. Issue triage is described as happening in batches, and the template asks for Python version, Neo4j version, py2neo version, operating system, the command run and the full error, with a pointer to the included sample spreadsheet for data questions.

The numbers fit the shape. 580 stars against 187 forks is a high fork ratio, which is what a tutorial repository looks like when people fork to try rather than to depend, and three open issues against that many forks suggests the archive is serving its purpose.

The last commit is dated 2026-09-01, and the repository is not archived.

Editorial conclusion

Use the legacy half if you are learning, and use the modern half if you are building. The separation is deliberate and documented rather than accidental, and the modern example is the part that will run on a current interpreter. Three things to check before you start. Check whether the legacy requirements install at all on your Python, because exact pins on transitive dependencies and build tooling from the jupyter stack will not resolve on a modern interpreter, and the readme tells you as much without pretending otherwise. Check which embeddings you are evaluating, since the demo path uses deterministic local placeholders to avoid needing a key, and placeholder vectors will not tell you anything about retrieval quality. And check the roadmap section against the releases before planning, because two of its three version ranges are already published.

Frequently asked questions

What does this repository demonstrate?

Reading invoice style spreadsheet data with pandas, extracting node and relationship rows from it, creating Neo4j nodes and relationships through the legacy py2neo library, and converting relationship data back into matrices for machine learning experiments. A second example does the same domain with the current driver and vector search.

Will the pinned dependencies install on a current Python?

Probably not. The requirements pin exact versions down to transitive dependencies, including a 2016 certificate bundle and jupyter stack packages from 2017 and 2018, so a modern interpreter will try to build them from source. The readme says the pins are intentionally legacy and points at a tracking issue for modernisation.

What has to change before the legacy example runs?

Two things. A working directory call in the source contains the literal placeholder `xxxx` that must be replaced with your repository path, and the placeholder `Graph(...)` connection in the data class must be replaced with your server URL and credentials.

Which embeddings does the modern example use?

Two sources. The production path uses the embeddings package for Neo4j GraphRAG, and the demo and continuous integration path uses local deterministic embeddings so no key is needed. Deterministic vectors are reproducible but say nothing about retrieval quality.

Is the roadmap in the readme up to date?

No. It describes three version ranges as pending. The 0.3 and 0.4 releases were both published on 2026-06-01 within ten minutes of each other, and no 0.2 release exists, while the roadmap text was never updated.

Official sources

  1. Issues
  2. License: MIT
  3. MazzaWill/neo4j-python-pandas-py2neo-v3 on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mazzawill-neo4j-python-pandas-py2neo-v3.svg)](https://hysenlabs.com/projects/mazzawill-neo4j-python-pandas-py2neo-v3)