Muuuun/luxas: npm install silently patches a vendored dependency, the cost table contradicts the prose about which model draws the figures, and no human in the loop is contradicted by the gallery note
An autonomous research colleague — from a question to a compiled manuscript, while you sleep.
At a glance
- What is it?
- A TypeScript harness that runs multi-agent scientific research from a topic file to a compiled PDF, with nine example reports published on a hosted gallery. It builds on four tarballs vendored from another agent framework, applies two patch scripts on install, and routes each agent role to a different model family depending on a profile. Its own prose describes one vision model for figures while its cost table attributes figures to a different one.
- Who is it for?
- The reasoning in this repository is better than the engineering scaffolding around it, and the two are worth separating. The independent-author principle is stated where most projects would bury it: a figure is drawn by one model and judged by another, and the cheap profile exists to cut cost while specifically excluding the auditor from the cut, because an auditor running the same model that produced the work is not an auditor.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
npm install runs two patch scripts against a vendored dependency
The install hook in the package manifest is two shell commands joined by an and. The first lifts read limits in the vendored coding agent. The second removes a tool-retry guard from the vendored agent core. Both run on every install, before you have a project, from scripts checked into the repository rather than applied by you. Lifting a read limit is a defensible change for a tool whose whole job is to read papers. Removing a retry guard is a different kind of decision, because the guard existing means somebody concluded that a failing tool call should be retried, and taking it out changes the failure semantics for every tool the agent calls. Nothing in the visible documentation explains which direction the guard worked in, so it is not possible to tell from the readme whether this prevents an unbounded retry loop or removes protection against a tool failing quietly. Two consequences follow regardless. On Windows without a POSIX shell the install fails outright, since it invokes bash, and the documented setup covers only macOS and Linux. And the patches have to be re-applied by hand after any upstream release, which is the ongoing cost of vendoring.
The vendored packages are published under a different name from the upstream
All four of the framework dependencies are file references to tarballs committed under a vendor directory, all at the same version. The readme names the upstream project and links it, crediting its author for the agent loop, tool lifecycle and hook primitives the harness is assembled from. The dependency names in the manifest use a different organisation scope entirely, one that appears nowhere else in the document. So a reader who wants to check what they are installing follows the link, arrives at the upstream repository, and finds a differently named set of packages. Version pinning is genuinely tight, since all four tarballs are the same build, which is the right way to vendor a framework whose internals you depend on. The cost is that the chain from the readme's link to the bytes on disk is broken by the scope change, and nothing in the documentation repairs it. The vendor directory also makes the pinned build auditable, which is worth more than it costs, provided you know to look there.
The cheap profile deliberately leaves the auditor on the expensive model
Model routing is the part of this design that is worth reading closely, because the cheap path is defined by what it refuses to cheapen. A profile redirect takes every agent that declared a Claude tier and moves it to a different family, applied in one place in the source. Provider-specific picks for math and reasoning bypass the redirect, and that is stated as deliberate. Vision-capable agents need a separate assignment, because the cheaper text models cannot see images, so the profile also sets a vision model for the three agents that draw, write and typeset figures. Then comes the exception. The figure auditor is left on its original tier, and the reason is a single sentence that generalises well beyond this project: it is the ship-or-no-ship eye, and an auditor running the same model that drew the figure is not an independent one. The same principle appears elsewhere as sibling agents writing implementation and tests blind to each other. Cost reduction is applied everywhere except the one place where cheapness would destroy the property being bought.
The cost table and the prose disagree about which model draws the figures
The model-switching section says the dual profile sets the vision assignment to one specific model, and adds that this model drew the best figure of three tested at a twentieth of the cost of the tier it replaces. The cost table below it describes the same profile as deepseek text plus deepseek vision, and its notes column attributes the figures to that provider's multimodal model. Those are two different vendors rendering the figures in the same preset. The table also notes that the preset loses ephemeral prompt caching, which is consistent with the system-prompt design listing cache-controlled blocks that the cheaper path cannot use. The cost figures themselves are labelled anecdotal, with a pointer to a per-project usage log for real numbers, which is an unusual and welcome bit of honesty about a number everyone else would present flatly. The vision contradiction is not labelled anywhere. Neither is the three-model figure comparison: the claim that one model drew the best figure of three is the only quality evidence offered for the cheapest profile, and no images or criteria are published with it.
No human in the loop is contradicted by the project's own gallery note
The summary paragraph promises a multi-hour run with no human in the loop, and the section describing example reports immediately qualifies it. Of the nine published runs, the note says a few required restarts or iterations through a pushback file when the reviewer and the reasoning component genuinely disagreed. So there is a human in the loop, there is a named mechanism for steering a run, and the project's own evidence that it gets used. The second half of that sentence is the better engineering statement: the harness is built around those crashes rather than against them, which is an honest description of what a long-running agent system actually is. The same paragraph makes another claim worth taking at face value for once. The finish conditions are described as deterministic gates that no prompt can talk past, meaning they are enforced in code rather than in instructions, which is the only way to stop an agent from arguing itself past its own stopping rule. The gallery note is the more trustworthy document of the two, because it reports failures.
An anti-detect browser reaches paywalled venues with no described scope
The literature stage names three scholarly indexes and then adds a fourth source with a different mechanism: paywalled venues, reached through an anti-detect browser. An anti-detect browser exists for the purpose of not being identified as automated traffic, so this is a deliberate measure aimed at publisher access controls rather than an incidental library choice. Unlike the other three sources, which are open indexes queried by topic, the target set for this one is decided by whatever the search returns, which means the system chooses which publisher sites to approach on its own. The visible documentation says nothing about request volume, pacing, which hosts are contacted, whether a publisher's terms are checked, or whether any list of permitted hosts exists. Two distinctions are worth keeping. This is not a delivered tool for reading content behind a paywall; it is one acquisition stage inside a research pipeline, and nothing in the repository is built around that capability. And the same caution applies to the rest of the pipeline, since a long run also spends a paid model API and writes to your filesystem. Check what a given venue permits before you choose a topic that will lead the crawler there.
The finish gates are also the test suite, and no test framework is installed
The package manifest's test script does not run a test framework, because there is no test framework in the development dependencies. It invokes a TypeScript script through a loader, and that script takes a flag to list what it would run and another flag on a sibling script to run in live mode. So the project's definition of passing is a bespoke gate runner with an inventory mode and a networked variant, and the same word covers both its own regression checks and the finish conditions of a research run. Using one mechanism for both is a genuinely good idea and worth crediting: the exit criteria a run must satisfy are the same kind of artefact as the criteria a change must satisfy, so they get tested by the same runner. The risk is the usual one. A gate runner written by the same team, in the same repository, with no framework underneath it, has no independent check on whether the gates are the right gates. The development dependencies are a TypeScript loader, the compiler, and three type packages, and one of those type packages is in the runtime list.
The runtime needs LaTeX, poppler, tmux and Python in the system interpreter
The document is unusually honest about this, opening the quick start by saying that the package install alone is not enough because agents shell out to a typesetting system, Python and a terminal multiplexer. The Debian line is the same set expressed as distribution packages:
sudo apt install texlive-latex-extra texlive-fonts-recommended poppler-utils tmux python3-matplotlib python3-numpyThe macOS line offers a full LaTeX cask or a smaller one at roughly a hundred and fifty megabytes, plus poppler, tmux and a specific Python version, and then installs two Python packages directly into the interpreter. So the real footprint is a Node project plus a LaTeX distribution plus a PDF toolkit plus a terminal multiplexer plus Python libraries, and the Python libraries land wherever pip puts them unless you have set up an environment yourself. tmux is the requirement most likely to surprise, since it means the harness depends on a terminal multiplexer being present to manage its long-running work, which rules out minimal containers and any machine without one. The tree is consistent with the design: two shell entry points at the root, one of them a daemon, plus a monitoring directory, and a native database binding and a native terminal binding in the dependency list.
Editorial conclusion
The reasoning in this repository is better than the engineering scaffolding around it, and the two are worth separating. The independent-author principle is stated where most projects would bury it: a figure is drawn by one model and judged by another, and the cheap profile exists to cut cost while specifically excluding the auditor from the cut, because an auditor running the same model that produced the work is not an auditor. Finish-gates that no prompt can talk past, and a harness described as built around crashes rather than against them, are the right instincts. Four things to weigh before running it. The install applies two patch scripts to a vendored dependency, one of which disables a retry guard, which means every fresh install silently changes library behaviour and those patches have to be re-checked after any upstream bump. The vendored packages are published under a different scope from the upstream project the readme links, so the provenance is not verifiable from the page. The cost table and the prose disagree about which model renders figures, and the one figure-quality claim has no comparison published. And the anti-detect browser for paywalled venues has no described scope, so check what a given venue's terms permit before you point a topic at it. For research you own, on venues you are entitled to read, this is an interesting harness. For unattended runs against a paid model API, budget it first: the default profile is quoted at twenty to eighty dollars a run.
Frequently asked questions
What does luxas do?
It is a multi-agent harness for autonomous scientific research. You give it a topic in a markdown file and it crawls the literature, reads what it finds, designs and runs experiments with implementation and tests written by separate agents that cannot see each other's work, produces figures from the raw results, writes a LaTeX report, subjects it to adversarial review on content, figures and layout, and emits a compiled PDF with real citations.
What does installing luxas actually do?
Besides the normal dependency install, the package manifest runs two checked-in patch scripts automatically: one that lifts read limits in the vendored coding agent and one that removes a tool-retry guard from the vendored agent core. Because the hooks invoke bash, installation requires a POSIX shell, and the documented setup commands cover only macOS and Linux. System dependencies are a LaTeX distribution, poppler, tmux and Python with matplotlib and numpy.
How much does a luxas research run cost?
The document labels its figures anecdotal and points at a per-project usage log for real numbers. It quotes twenty to eighty dollars per full run for the default profile, which uses Claude throughout and is described as the only profile with prompt caching, and two to ten dollars for the dual profile, which redirects most agents to a cheaper family and is noted as losing ephemeral caching.
Can I run luxas without a human in the loop?
The summary claims no human in the loop, but the example-report note says a few of the nine published runs needed restarts or iterations through a pushback file when the reviewer and the reasoning component disagreed. Finish conditions are described as deterministic gates enforced outside the prompt, so a run stops on its own, and the document says the harness is built around those crashes rather than against them.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/muuuun-luxas)