Academic Commercialization Agent: a scorecard that documents its own gaps
Evidence-constrained commercialization assessment with deterministic retrieval, six LLM stages, auditable scoring and checkpoint recovery.
At a glance
- What is it?
- This agent turns a paper or topic into a source-linked commercialization assessment, with deterministic retrieval in front of six model stages and a deterministic weighted total at the end. What makes it unusual is the evidence section: thirty baseline runs reported as five separate checks, three disclosure documents about what the market score cannot yet compare, and a root directory full of one-off audit scripts.
- Who is it for?
- This project fits a research group that wants a written assessment with citations it can audit, and a reader who trusts a tool that publishes the checks it fails rather than a single accuracy number. It does not fit anyone who needs diligence advice, because the README rules out technical, legal, regulatory, investment and freedom-to-operate conclusions before you read a word.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Deterministic retrieval, six stages, one weighted total
The architecture is the clearest statement of intent, and the README prints it as a diagram:
Topic / PDF + optional Decision Context
│
Deterministic retrieval → validation → frozen source registry
│
┌─────────┼─────────┐
Academic Patent Market
└─────────┼─────────┘
Writer
│
Reviewer
│
Scorer → deterministic weighted total
│
Shared run artifacts + terminal truth
│
FastAPI / browser / CLI / recoveryTwo structural choices stand out. Retrieval is deterministic and the validated sources are frozen into a registry before any model reasons, so the evidence set is fixed rather than sampled per stage. And the score is a deterministic weighted total, which means the arithmetic is auditable even when the inputs are not.
The three evidence specialists run in parallel and the writer, reviewer and scorer follow in sequence, with one caveat the README puts in bold type: a stage is not a promise of exactly one model request. So six stages means six decision points, not six API calls.
The boundary table that follows makes the same point in seven rows, from evidence provenance tiers through to cost states named complete, lower-bound and unavailable.
The baseline is thirty runs, reported as five separate checks
The measured results section is the most unusual part of any project README, and it starts by refusing to be a single number. The frozen baseline is ten topics times three live repetitions, and the table reports five checks that are explicitly described as different checks rather than a combined accuracy score.
End-to-end completion is 30 out of 30. Weighted formula correctness is 30 out of 30, which is the deterministic part behaving. Complete report structure is 30 out of 30. Unsupported numeric lines are zero across all thirty reports. TRL calibration, the one that is not perfect, is 26 out of 30, with seven of ten topics meeting their expected range in all three runs.
Then the caveats, stated rather than buried. The expected TRL ranges were adjusted after early observations, so this is not independent held-out validation. The uncited-numeric proxy does not measure all hallucinations. And a valid citation ID does not establish that a source entails a claim.
A vendor claiming 100% would have had three rows of this table available and chosen one.
The ablation says four nodes beat six
The topology ablation is the most actionable number in the repository, and it points against the project's own default shape.
Across a 90-cell ablation, the four-node arm used 54.89% fewer median tokens and 47.03% lower median cost than the six-node arm. The conclusion the README draws is the restrained one: six nodes were not established as universally necessary.
That matters because the pipeline described in the architecture is built around the parallel specialists, and the cost of a parallel stage is paid on every run. Nearly half the median cost of the pipeline is attributable to structure that the ablation cannot justify.
The repository also carries the tooling for this kind of study as first-class scripts rather than notebooks: an ablation script, an ablation check, a benchmark script with fixtures and an identity module, and a checkpoint fault audit. The presence of a `benchmark_identity` module suggests the team worried about results drifting between runs, which is the failure mode that quietly invalidates an ablation.
The utility study and the pilot both failed their own rule
Two pieces of user evidence are reported, and in both cases the registered success rule did not pass.
The five-reviewer utility study collected 20 eligible judgments, and the registered success rule failed despite a six-to-four preference for the full workflow in every round. So reviewers preferred the thing and the measurement did not register that preference as success.
The target-user pilot ran with two people. Both retained a defer answer and a maybe-to-reuse answer, and neither checked external sources. The README states plainly that this does not establish product adoption.
Reporting a failed registered rule next to a unanimous preference is unusual, and it is arguably the most credible thing in the evidence ledger: a metric that never fails is a metric nobody is measuring.
The recovery evidence follows the same pattern. Thirty of thirty offline fault-injection children completed, and one production child reused four committed nodes, with the explicit note that this is not an exactly-once or general cost-saving guarantee.
Recovery works by immutable children, not by resuming
Interrupted runs are recovered by branching rather than by continuing, and the mechanism is described precisely enough to check.
A recovered run becomes an immutable child built from the longest validated checkpoint prefix of the interrupted run, with fresh credentials. Checkpoints are content-addressed, recovery children are immutable, and the terminal records are write-once.
The user-facing version of the same rule is a single sentence: recover an interrupted run as an immutable child using its longest validated checkpoint prefix and fresh credentials.
That is a different promise from resume. It means an interrupted analysis is never silently completed with a mixture of old and new state, at the cost of doing work again from the last validated point. For a paid model pipeline, that trade is usually the right one, and the wording here does not pretend it is free.
The runtime boundary row lists the same guarantees in one line: subprocess isolation, content-addressed checkpoints, immutable recovery children and write-once terminal records.
Three disclosure documents about the score cap
The market score is the part the project keeps qualifying, and it does so with three linked documents rather than a footnote.
The first discloses unverified estimate comparability, and the framing is careful: it makes the historical cap's limitation visible without fixing or validating the underlying market-scoring policy. The second shows pre-cap and post-cap scores with the actual deduction, and states that historical missing receipts cannot be reconstructed. The third is a 20-source inspection that still establishes no fully comparable real pair.
The conclusion the README draws is that the scoring policy is not yet validated.
So a reader who looks only at the scorecard sees a number whose comparability the project itself will not defend. That is an uncomfortable design, and it is the honest one: a weighted total is only as meaningful as the distribution it was calibrated against, and there is not yet one.
The saved-source lookup path is the same story. It makes at most one bounded model selection from saved identifiers and titles, returns exact saved text, generates no research answer and performs no new search.
Turning a warning into an error is how the tests catch API leaks
The test configuration contains one of the better engineering comments in any Python repository, and it explains itself.
Warnings are configured as errors. The stated reason is that the paths that call a real provider warn and fall back rather than raising, so a test that quietly spent money would otherwise pass. Promoting the warning to an error is what makes the leak fail.
The configuration lives in the file rather than on the command line for a shell-specific reason worth repeating: the filter list contains a backtick, which PowerShell reads as an escape character and bash reads as command substitution, so the same command cannot survive both shells.
The same reasoning appears in the test collection boundaries, where output directories and assistant worktrees are excluded from recursion because a byte-frozen repository snapshot would shadow the canonical tests once a study has run.
Both comments describe tests protecting evidence rather than coverage. A test suite that only counts assertions would have shipped the leak.
The container pins the interpreter, the font and the init process
The deployment files are commented at the level of decisions, and three of those comments are worth reading before you build anything.
The interpreter is pinned to a specific Python version through a build argument rather than tracking latest, so a rebuilt image resolves the same interpreter the lockfile and the CI matrix were tested against. The dependency stage is split from the source stage because the framework and its transitive dependencies take minutes to resolve, and without the split editing one Python file would reinstall all of them.
The font is the most specific. Noto CJK is the obvious choice and does not work, because Debian ships it as OpenType with CFF outlines while the PDF library's TrueType reader rejects that with a postscript outlines error, after which the renderer degrades to a fallback face and reports come out as blank boxes. A different CJK font was chosen because it is TrueType and covers Simplified, Traditional, Japanese and Korean glyphs.
Finally, the compose file deliberately omits the usual init setting because the image already runs tini as its entrypoint, and setting both produced two init processes in a chain that reaps nothing the first one would not. The memory limit is labelled a ceiling that keeps one container from starving its host rather than a measured requirement.
Editorial conclusion
This project fits a research group that wants a written assessment with citations it can audit, and a reader who trusts a tool that publishes the checks it fails rather than a single accuracy number. It does not fit anyone who needs diligence advice, because the README rules out technical, legal, regulatory, investment and freedom-to-operate conclusions before you read a word. Before you rely on a score, read the three disclosure documents the repository links about estimate comparability and the score cap, and treat the five baseline checks as five separate facts rather than a combined accuracy figure, because the author says so explicitly.
Frequently asked questions
What does the Academic Commercialization Assessment Agent produce?
A source-linked commercialization assessment draft with an auditable scorecard and risk notes, from a research topic or an attached paper PDF. You choose the report language and scoring profile, optionally add Decision Context, follow progress, inspect citations and reliability warnings, and export Markdown or PDF or share a run link.
How does the agent gather evidence?
Retrieval is deterministic: source-native clients plus web search, URL and DOI checks, provenance tiers, deduplication and registered source IDs, producing a frozen source registry before any model reasons. Six model stages then work over that validated evidence, with three specialists running in parallel and a deterministic weighted total at the end.
What do the measured results actually show?
A frozen baseline of ten topics times three repetitions: 30/30 end-to-end completion, 30/30 weighted formula correctness, 30/30 complete report structure, zero unsupported numeric lines across 30 reports, and 26/30 on TRL calibration. The README says these are different checks rather than a combined accuracy score, and that TRL ranges were adjusted after early observations.
Can I recover an analysis that was interrupted?
Yes, as an immutable child run built from the longest validated checkpoint prefix with fresh credentials. Checkpoints are content-addressed and terminal records are write-once, so an interrupted run is never silently completed with mixed state. Offline fault injection completed 30 of 30 such recoveries.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/shuxiachai-academic-commercialization-agent)