PaperGuru reports two PaperBench means and ships no implementation
Lifecycle-Aware Memory for long-horizon LLM agents — 66.05% on PaperBench, 94.66% on SurveyBench, 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB
At a glance
- What is it?
- A memory architecture for long-horizon agents, published as a paper, two benchmark result folders and a set of figures. The design argument is well specified, the numbers are close to a human-expert bar on one benchmark, and the repository contains nothing you can install.
- Who is it for?
- PaperGuru is worth reading as a design document, because the four axioms and the head versus content split are a clear statement of what a long-term memory layer has to do that a vector store does not. It is not something you can adopt today from this repository: there is no implementation, no releases, and the last push is dated 8 June 2026.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 116 days ago.
- What is it written in?
- Mainly TeX, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two PaperBench means, one digit apart, two different paper sets
The project's headline figure appears twice with different scopes. The results table reports a mean reproduction score of 65.95% on a 20-paper shared set, and the surrounding text reports a per-paper mean of 66.05% across all 23 papers. The repository description quotes the second one.
| **PaperBench** (OpenAI, 2025) | Mean reproduction (20-paper shared set) | **65.95%** | 35.74% | **+30.21%** |
| **PaperBench** | Papers above 41% human ML-PhD bar | **20 / 23** | 4 / 23 | **+16 papers** |The distinction matters because only one of them is the number that can be compared with the 35.74% baseline. A mean over 20 shared papers is measured against other systems on the same 20; a mean over all 23 includes papers the baselines did not run. Reading the two as the same measurement is the easiest way to overstate the result.
The second row is the more informative figure anyway. PaperBench scores a runnable code tree with a leaf-judge model against a hand-written rubric, and the official human-expert reference is 41% given a 48-hour budget from a machine-learning PhD. Going from 4 of 23 papers over that bar to 20 of 23 is a statement about distribution rather than about a mean. SurveyBench results follow the same pattern, with a content score of 94.66% against 80.60% and a composite richness score of 43.76% against 20.36%.
Memory is split into a bounded routing surface and an unbounded body
The central design move is a split. Each artifact gets a compact chunk head, one per artifact, forming a bounded routing surface. The raw text lives separately as chunk contents and is reached only on demand. A single capital chunk indexes every head and supports capital-first routing over a temporal artifact graph.
This split is what makes the third axiom affordable. The axiom is bounded query cost under unbounded archive growth: the archive grows every day and routing cost must not grow with it. As long as routing only ever touches heads, the query cost is capped by the head surface rather than by the archive. The head is a table of contents, not the book.
The four axioms are stated as numbered claims in the paper's third section. Versioned content is the first: a statement that was once correct has to be able to become stale after revision, deprecation or retraction, and the memory layer has to know about that transition rather than keep serving the old truth. Structural multi-hop relevance is the second, and it is defined as the right evidence being two citations away rather than one cosine-similarity hop away. Provenance-grounded composition is the fourth: every claim in the output traces back to a verifiable artefact in memory.
Historical-causality edges are the part a plain knowledge graph never grows
The temporal artifact graph unifies two classes of edge, and the distinction between them is where the design earns its keep:
| **Structural edges** | `cites`, `benchmarked-on`, `introduced-by`, `implements` |
| **Historical-causality edges** | `discussed-in`, `deprecated-by`, `retracted-by`, `superseded-by` |Structural edges are the ones a conventional graph wrapper already handles. `cites` and `implements` say how artefacts relate to each other in content terms. The second class is about a relationship in time: a claim that was discussed in one paper is deprecated by a later one, retracted, or superseded. Those four edge types are what let the graph express the first axiom. Without them, a corrected fact and the fact it corrected are siblings, and the retrieval step has no way to prefer the newer one.
The cost is that the extractor now has to produce eight distinct relationship types rather than two or three, and a `deprecated-by` edge is only useful if it is right. Nothing in the repository describes how any of these edges are extracted, how a retraction is detected, or what happens when a superseding paper is itself superseded. The eight names are a specification of the target state, not a description of a mechanism that is currently running.
Reason takes 45 percent of the run and contains the only loop
The pipeline has four stages, and the project publishes an indicative share of wall-clock time for each, for a typical 200K-token survey run:
| **01 · SEARCH** | Topic query, candidate archive | Ranked artifact heads | ~15% |
| **02 · EXTRACT** | Heads + chunk contents | Evidence cards (text + provenance) | ~20% |
| **03 · REASON** | Evidence cards | Draft segments (Compose → Critique → Mutate loop) | ~45% |
| **04 · VERIFY** | Draft segments | Cited, provenance-checked output | ~20% |Two things are worth pulling out. First, the stage that consumes the most time is the only one with a loop inside it: compose, then critique, then mutate. Retrieval and verification are single passes; reasoning iterates, which is why it lands near half the budget. Any optimisation aimed at the search stage touches the smaller share of the work.
Second, the shares are labelled indicative and stated to vary by task. They are a rough shape, not a profile, and the figures given are for a survey-sized run rather than for the paper-reproduction task that produces the PaperBench number. A reader comparing cost across systems should not treat the 15 percent as a measured constant.
Evidence cards are the only structure the rest of the system touches
Query-time context is assembled by a route first, expand second, distill last pipeline, and what it yields is an evidence card: text plus provenance. The README is explicit that this is the single data structure the rest of the system operates on, which is a strong architectural claim. Search does not pass documents forward, and reason does not see raw archive text; both ends meet at cards.
That boundary is what enforces the second and fourth axioms together. A card cannot be built from one high-similarity neighbour, because expansion happens after routing and distillation runs over what came back, so a claim that needs two citations to make sense has to be assembled from both before it reaches the model. And a card carries provenance by construction, so the verify stage has something to check against rather than a free-text draft with citations appended afterwards.
The claimed benefit is that this satisfies all four axioms at once without per-task tuning. The claim about competitor designs is that memory-tier schemes and forgetting-curve schemes each cover one or two of the axioms and never all four. That comparison is asserted rather than measured here; no ablation isolating one axiom at a time appears in the repository.
The repository is a paper, two result folders and no code to run
The top-level contents are a licence file, a gitignore, the English and Chinese READMEs, an assets directory, a paper directory, and two benchmark directories named PaperBench and SurveyBench. There is no source directory, no package manifest, no dependency file, and no build configuration.
The primary language recorded for the repository is TeX, which tells you where the content lives: the paper directory holds the manuscript, and the link in the README header points at a PDF called PaperGuru-CCM.pdf. The two benchmark folders appear to hold the evaluated artefacts rather than the system that produced them.
So there is nothing here to install and no command to run, and the README does not offer a quick start. Its table of contents promises sections covering what is in the repository and how to reproduce the figures, which is the right place for that information, but neither can be followed from the tree alone: reproducing a 65.95% mean requires the implementation, and the implementation is not published here. Anyone planning to build on this is building from the paper.
No releases, no open issues, and a last push on 8 June 2026
The repository has no published releases at all. Its last push is dated 8 June 2026, roughly four months before the reference point used to judge how current it is. It carries 1325 stars, 198 forks and zero open issues.
Zero open issues is the number that deserves a second look. For a repository carrying a claim of being the first system built on a four-axiom formalisation, plus two benchmark tables and a peer-review track record, the absence of any open issue is more consistent with a results archive than with a project taking feedback. It is not evidence of anything wrong, and the closed-issue count is not visible, so the honest reading is simply that there is no visible public support channel.
The claims themselves should be read as claims. The README states 10 peer-reviewed acceptances across five venues since Q4 2025, and the header describes state-of-the-art results on both benchmarks. Nothing in the repository is a citation list for the five venues, so those counts cannot be checked from here.
The licence field carries no identifier while a LICENSE file sits at the top level
The repository's licence metadata resolves to no recognised identifier, which is what GitHub records when it cannot classify a licence file automatically. At the same time, a LICENSE file is present at the top level, and the README's table of contents ends with a licence section. So the terms exist somewhere in the repository and are not machine-readable from the outside.
For a repository that is mostly paper, figures and benchmark artefacts, that gap is the one worth closing first. Anyone wanting to reuse the figures, the survey results or the manuscript has to open the file and read it, and the absence of an SPDX identifier means tooling will not warn them first. Copying a figure out of a paper PDF is the kind of reuse most people do by accident.
The bilingual README is worth noting too, with a Chinese version linked from the header. For a project presenting itself as a research artefact rather than a tool, that is a sensible choice, and it does not change any of the limits above.
Editorial conclusion
PaperGuru is worth reading as a design document, because the four axioms and the head versus content split are a clear statement of what a long-term memory layer has to do that a vector store does not. It is not something you can adopt today from this repository: there is no implementation, no releases, and the last push is dated 8 June 2026. Before quoting a figure, check which of the two PaperBench means you are citing, and read the licence file, since the repository's licence field carries no recognised identifier.
Frequently asked questions
What score does PaperGuru get on PaperBench?
Two figures are given for different scopes. Mean reproduction on a 20-paper shared set is 65.95% against a 35.74% baseline, and a per-paper mean across all 23 papers is 66.05%. On the count of papers clearing the 41% human ML-PhD bar, it is 20 of 23 against a baseline of 4 of 23.
Can I install PaperGuru from this repository?
No. The top-level contents are a LICENSE file, two READMEs, an assets directory, a paper directory and the PaperBench and SurveyBench folders, with no source directory, package manifest or dependency file. There is no install command and no published release.
What are the four axioms behind PaperGuru's memory design?
Versioned content, so a statement can become stale after revision or retraction. Structural multi-hop relevance, where the right evidence is two citations away rather than one similarity hop away. Bounded query cost under unbounded archive growth. And provenance-grounded composition, where every output claim traces to a verifiable artefact.
Which stage of the PaperGuru pipeline uses the most compute?
The reason stage, at about 45 percent of wall-clock time for a typical 200K-token survey run, and it is also the only stage containing a loop, a Compose, Critique and Mutate cycle. Search takes about 15 percent, and extract and verify about 20 percent each. The shares are labelled indicative and vary by task.
What licence does PaperGuru use?
The repository's licence field carries no recognised identifier, though a LICENSE file is present at the top level and the README's table of contents includes a licence section. The terms have to be read from the file rather than inferred from repository metadata.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/autotrustai-paperguru-benchmark)