# YiGraph: the licence badge links to a file the repository does not contain, and the requirements pin a dozen CUDA libraries exactly

> YiGraph is a Python agent system for graph data analytics built on an analytics-augmented generation design, where a language model plans the analysis and a library of graph algorithms does the computing. The design is careful about not letting a model write and run its own code. The packaging is less considered: a requirements file with thirteen exact CUDA version pins, a graph database client four years older than the PyTorch beside it, and no package metadata at all.

**iDC-NEU/YiGraph** — YiGraph is an LLM-driven agent for autonomous Graph Data Analytics based on Analytics-Augmented Generation.  易图（YiGraph）是一套基于 AAG（分析增强生成）框架构建的图分析智能体系统，致力于挖掘数据之间的关联关系，释放数据价值。

- Repository: https://github.com/iDC-NEU/YiGraph
- Stars: 1,202 · Forks: 96
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/idc-neu-yigraph

## The licence badge points at a file the repository does not contain

The first block of the readme is a row of badges, and one of them is a link to a licence file at the repository root. The repository root does not contain one. The top level listing holds two ignore and configuration files, a submodule file, the English and Chinese readmes, and six directories: the engine directory, a configuration directory, a documentation site, a figures directory, a requirements file and a web directory. There is no licence file anywhere in that list, and the platform reports no licence for the repository at all. So a reader who clicks the badge, as the layout invites them to, lands on nothing, and a reader who checks the repository metadata finds the same answer. That is not a licence question the project has answered with the wrong document, it is a licence question the project has not answered in the tree. Anyone planning to build on this code needs the terms from the authors directly, because nothing in the repository states them.

## Every CUDA library is pinned to one exact version

The requirements file is unusually explicit, and the deepest part of it is a block of thirteen NVIDIA runtime packages, each pinned to an exact build number rather than a range:

```text
nvidia-cudnn-cu12==9.5.1.17
triton==3.3.1
```

Alongside them sits the framework they serve:

```text
torch==2.7.1
torch-geometric==2.6.1
```

Two consequences follow. A CPU-only installation still resolves that whole block, because a requirements file has no notion of hardware, so a machine that never calls a GPU downloads several gigabytes it will not use. And the pins are a set rather than independent choices: they describe one CUDA generation as it was assembled for one framework release, so upgrading anything means re-deriving the entire block rather than bumping one line. The file opens with a comment stating that Python 3.11 or newer is required, and links to the Python downloads page rather than declaring a floor in packaging metadata, which is consistent with the rest of this project having no package manifest.

## The graph database client is four years older than the framework beside it

Read the pins as a set and they span a wide range of ages. The numeric and validation libraries are current:

```text
numpy==2.3.1
pandas==2.2.3
networkx==3.5
```

The graph query client is not:

```text
neo4j==4.4.12
```

That is a client from the 4.4 line, pinned exactly, sitting next to a framework release from the middle of the current generation. Exact pinning is usually the safer choice, but pinning to an old branch means the client cannot pick up fixes from later releases in the same major version, and it sets the ceiling on which server versions the project can be pointed at without a dependency change. There is another old-thing dependency in the file, and it explains itself: the build section carries setuptools with a comment saying it is needed at runtime because a legacy component still calls into the packaging resource API. So the dependency list documents its own archaeology, which is more than most do, and a reader can see exactly which parts of the stack are load-bearing on something old.

## The algorithm table accounts for 85 of a claimed 200 across 9 of 21 categories

The feature claim is a library of more than 200 graph algorithms across 21 major categories, and the readme says outright that the table shows only a subset of representative categories. Counting the visible rows gives the size of the gap. Basics has 10, path 13, centrality 14, connectivity and components 13, clustering and community 17, tree and spanning tree 3, flow and cut 5, matching and coloring 6, and cliques and cores 4. That is 85 algorithms in 9 categories, so a little under half the claimed count across a little under half the claimed categories. Each row names its algorithms and links to a tutorial page in the documentation site, so the full list exists somewhere. The last visible row is cut off mid cell, after three algorithm names and a fourth that trails off, which means even the subset table does not finish. Nobody should conclude the library is smaller than advertised on this evidence, but nobody can size it from the readme either.

## The design refuses arbitrary code execution, and that refusal is the architecture

This is the part of the project worth understanding before any of the packaging details. The stated approach is analytics-augmented generation, and the claim is narrow and specific: the system does not let the model write a piece of uncontrollable code and run it. Instead the system centres on verifiable algorithm modules that are invoked and combined, which the readme frames as making each analysis step reproducible, meaning the same input yields stable output, and traceable, meaning you know which algorithms ran and in what order. The division of labour is explicit: the language model understands intent, breaks the question into steps, and organises the final output, while the computing is done by the modules and the model then interprets and summarises what they returned. That is a defensible architecture for the scenarios the project targets, which are all compliance or risk shaped. It is also a claim made by the project about itself, with no independent verification in the visible text.

## No package metadata exists, so the only documented setup is a requirements file

There is no packaging manifest in this repository. The top level holds a requirements file and nothing that declares a version, a package name, an entry point or a console command. The readme's visible text contains no installation command and no invocation either; it opens with what the project is, lists the industries it targets, describes four core features and then shows the algorithm table. Everything actionable is one level down, in the documentation site directory, which holds the per category tutorial files the table links to and which is published at a GitHub Pages address. Two details in the layout explain the shape. The submodule file at the root means at least one component is pulled in from elsewhere rather than developed here, and the readme badge row points at a documentation site, a figures directory and two WeChat images, which suggests the documentation is treated as a separate deliverable from the code. In practice this is a source tree you run in place, not something you install.

## Five target scenarios and not one number to judge them by

The applicable scenarios list is specific and entirely about risk. Financial anti-money laundering and suspicious transaction analysis, described as building transaction networks to find abnormal fund paths and suspicious loops. E-commerce risk control, which the readme frames as combining accounts, devices and addresses to surface organised fraud and what it calls wool party behaviour, a term for coordinated abuse of promotions. Enterprise association investigation through ownership and transaction structures. Event analysis for a park or city, unifying access control and trajectory data. And supply chain risk, tracing transmission paths between enterprises. Every one of these is defensive analysis that a regulated organisation might legitimately want, and the project's pitch is that the output is a traceable report rather than prose. What is missing is any evidence at all in the visible text: no benchmark, no dataset, no accuracy figure, no run time, and no description of what happens when the graph library has no algorithm for the question the user asked.

## WeChat groups and a pages address are the only doors into this project

The badge row at the top of the file is a map of how this project is reached, and it is unusual. Alongside the licence and Python links there is a documentation site address, two WeChat images, a notes platform image, and an X account. So the public entry point is a documentation site hosted on the project's own pages domain, and the community channels are WeChat groups plus a social account, with no issue tracker mentioned in the visible text at all. That has a practical consequence for anyone outside that ecosystem: the answer to most questions will live in a chat group rather than in a searchable thread. Inside the repository, the documentation site directory is the real reference, and its structure is visible from the table links alone, one tutorial file per algorithm category with names covering basics, paths, centrality, connectivity and components, clustering and community, trees, flow and cut, matching and coloring, and cliques and cores. Nothing else about that site is described in the readme.

## Conclusion

YiGraph fits a team doing relationship analysis on data it already owns, where the requirement is a report a reviewer can retrace rather than a chat transcript. The architecture is the reason to look: computation is delegated to named algorithm modules, and the language model plans and summarises rather than executes, which is the correct instinct for compliance work. Four things to verify first. Whether the results are reproducible on your data, because the claim of stable output for the same input is the project's own and the visible text contains no benchmark, no dataset and no accuracy number at all. Whether the dependency set installs on your machine, because the requirements file is pinned to one CUDA generation and one old graph database client. What the licence actually is, because the platform reports none and the badge in the readme points at a file that is not in the tree. And what you are reading, because the algorithm table shows nine of twenty-one claimed categories and its last visible row is cut off mid cell, so the library's real size has to be measured against the documentation site rather than the readme.

## FAQ

### What does the AAG framework do in YiGraph?

It treats analytical computation as a core capability, invoking graph algorithms and graph systems at key stages to produce verifiable calculations that the model then interprets and summarises. The project states it does not let the model write and run uncontrolled code, and centres on verifiable algorithm modules instead.

### How many graph algorithms does YiGraph include?

The readme claims more than 200 across 21 major categories and says its table shows only a subset. The visible rows total 85 algorithms across nine categories, and the last row is cut off partway through its cell.

### What does installing YiGraph require?

A requirements file that states Python 3.11 or newer, and which pins the numeric and validation libraries, a graph processing stack including NetworkX, a Neo4j client at version 4.4.12, PyTorch with its geometric extension, and thirteen NVIDIA CUDA runtime packages plus a compiler runtime, each at an exact version.

### What license is YiGraph released under?

The platform reports no license, and there is no license file in the repository's top level listing, even though the readme's first badge links to a license path. The terms have to be confirmed with the authors directly.

### How does YiGraph turn a question into an analysis?

It works out what the question is trying to solve, then decomposes it into steps covering which data fields and relationships are needed, what graph to build, which methods and parameters to use, and how results should be interpreted and presented. The steps are executed through algorithm modules rather than generated code.

### Where is the YiGraph documentation?

In a documentation site directory inside the repository, with one tutorial file per algorithm category linked from the table in the readme, and published at a GitHub Pages address. Community discussion runs through WeChat groups and an X account, and the visible text mentions no issue tracker.

## Sources

- [iDC-NEU/YiGraph on GitHub](https://github.com/iDC-NEU/YiGraph)
- [Issues](https://github.com/iDC-NEU/YiGraph/issues)
- [README](https://github.com/iDC-NEU/YiGraph/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/idc-neu-yigraph
