Agents-A1 argues you can scale what a model does for forty-five thousand tokens instead of how big it is
Scaling the Horizon, Not the Parameters
At a glance
- What is it?
- A mixture-of-experts model with about three billion parameters active, trained on agent trajectories averaging forty-five thousand tokens, and evaluated on a table it publishes in full including the benchmarks where it loses to the largest systems.
- Who is it for?
- Read the published table rather than the badge, because the table is the honest part of this release. This model wins clearly in its own size class and loses to the largest systems on several of the benchmarks where the difference matters most for long tasks, including the long-horizon search set that is the centre of its own thesis.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 81 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The claim is about trajectory length, not parameter count
The title is the argument. The project describes itself as scaling the horizon rather than the parameters, and reaching what it calls trillion-parameter-level performance with a mixture-of-experts model in the thirty-five billion class with about three billion active per token.
Read that as a hypothesis rather than a result, because that is what it is. The claim is that a meaningful share of the capability attributed to very large models comes from how many steps a model is allowed to take, and that you can buy some of it back with better training data rather than more weights. Whether the numbers support that is a separate question, and the evaluation section is more interesting than the framing because it lets you check.
The horizon is defined operationally. The project built an infrastructure it calls long-horizon knowledge-action, connecting external knowledge, actions, observations and verifier outcomes, and used it to produce agentic trajectories with an average length of forty-five thousand tokens. Forty-five thousand tokens is a long conversation with a tool-using model: dozens of tool calls, several retrieval rounds, and intermediate results the model has to reason about rather than recite.
That is the substantive bet. Everything else in the repository is either the method for producing such trajectories or the evidence about what they buy.
Three stages, and the third one is the interesting one
Training is a three-stage recipe, and the stages tell a coherent story about what is hard about this problem.
Stage one is full-domain supervised fine-tuning across everything, described as the step that aligns the base model with broad agentic behaviour. That is the step that turns a base model into something that uses tools at all, and doing it across all domains at once rather than per domain is the choice that makes a single deployable model possible later.
Stage two trains a separate teacher model per domain, each capturing specialised expertise. This is the expensive part and the part you cannot skip: you cannot get expert-level performance in six unrelated domains from one teacher without either training six teachers or losing the specialisation.
Stage three is what makes the whole thing tractable. It is a multi-teacher, domain-routed distillation that runs on policy, with what the project calls salient vocabulary alignment to improve transfer efficiency across domains. Routing means the student is trained against the teacher for the domain it is currently in, so the specialised teacher is used where it is strong rather than averaged everywhere. Running on-policy means the distillation samples come from the student's own distribution rather than the teacher's, which is the difference between distilling what the model would say and distilling what it would actually reach.
Salient vocabulary alignment is the least explained piece and the most suggestive: if the teachers use different words for the same concepts, the student is being asked to reconcile token distributions that do not agree, and aligning the vocabulary where the disagreement concentrates is a plausible fix for exactly the transfer loss this setup would otherwise suffer.
The benchmark table publishes its losses
The results section is a full table rather than a highlight reel, and it compares against seven other models split into two columns: three of roughly the same size class, and four considerably larger ones including two at the top of the current range.
The two badges used throughout the table mean overall best result and best among models in the same size class. Reading them honestly, the pattern is consistent. On long-horizon search, the model is the best of the three comparables on all four benchmarks, and the best overall on one of them, while sitting below the largest systems on two others.
On engineering tasks the gap to the frontier is wider: the model leads the comparable class and lands mid-pack among the larger ones, and on at least one of those rows the largest system is far ahead. That is the honest shape of the result, and the fact that the readme shows the losing rows rather than cropping them is the single most useful thing about it.
The scores the readme calls out as overall best are on a mix of specialised benchmarks: instruction following on two, a scientific olympiad set, a scientific research set, and a browser-based agent benchmark it leads among comparables while a larger system edges it. Averages across a table like this are close to meaningless because the benchmarks weight very different abilities, which is a point worth making before anyone quotes a single number from it.
Six domains, and one model standing in for six specialists
The unifying claim is that six heterogeneous domains end up in one deployable student model, and the evaluation is organised along a matching set of directions: long-horizon search, engineering tasks, scientific research, instruction following, general agentic tasks, and scientific agentic tasks.
That list is a reasonable design for the thesis. If you believe the horizon is what matters, then your benchmarks should include tasks where the horizon is long: search that requires many retrieval rounds, coding that requires many tool calls, research that requires following a citation chain. If your benchmarks were all single-turn, you would be testing instruction following and calling it agentic ability.
Six domains is also where the multi-teacher design pays or fails. Two domains are enough to demonstrate the idea. Six is where the domain-routed routing has to decide, at training time, which teacher owns which input, and where a mis-routed sample silently degrades the student rather than failing loudly.
The listed capabilities are the standard four for this class: decomposing complex tasks into executable sub-steps and adapting based on intermediate results, native function calling and tool integration, long-context coherence and recall, and precise following of multi-constraint instructions. Nothing there is novel. All four are things the horizon claim depends on, and instruction following in particular is a natural casualty of very long trajectories, which makes its presence in the results a more meaningful signal than its presence in the feature list.
What is in the repository, and what is not
The repository is small and its contents define what kind of artefact this is: a licence, the readme, assets, documentation, an evaluation directory and a scripts directory.
There is no training code. The three-stage recipe is described in prose in a technical report, and the pipeline that produced forty-five thousand token trajectories is described in prose as well, but neither ships as something you can run. What ships is evaluation code for selected domains.
That makes this a weights release rather than a method release, and it is an honest one rather than a disappointing one: teams publish what they intend to maintain, and re-running a multi-teacher distillation across six domains is not a thing you can commit to keeping working. The cost is that the central claim is not independently checkable. You can download the student and run the benchmarks; you cannot download the teacher and check that it was better than the student in the first place.
The assets are distributed through the usual model hubs, on both the primary hub and a regional mirror, and there is a technical report on a preprint server and a project page. The model itself is offered at a larger size and at a much smaller one, with the smaller release described as requested by people building local assistants, which tells you something about who the authors expect to use it.
Quantised variants exist because the interesting deployment is on a laptop
Quantised variants were released about a week after the main model, and the readme thanks a community quantisation group for producing versions at multiple scales, with an explicit invitation to run it on a Mac.
That invitation is worth taking seriously. A mixture-of-experts model with a small active parameter count is an unusually good fit for local deployment: the memory footprint is dominated by the weights you have to hold rather than the ones you compute over, so quantisation buys you more than it would on a dense model of the same total size.
The news timeline also shows what happened next. The larger model came first, with evaluation code and the report. Quantised variants followed. Then a note that the smaller model was coming because the community asked for it, and it was released a few days later.
That sequence is a small but real signal about priorities. The distribution strategy is not one checkpoint on one hub; it is a family, produced in descending order of enthusiasm and with community quantisation treated as a first-class channel rather than a nicety. For anyone deciding whether to spend an afternoon on this, the small model is the right place to start, precisely because it is the one nobody optimised for you.
Editorial conclusion
Read the published table rather than the badge, because the table is the honest part of this release. This model wins clearly in its own size class and loses to the largest systems on several of the benchmarks where the difference matters most for long tasks, including the long-horizon search set that is the centre of its own thesis. That is a coherent position: if your budget is a serving cluster rather than a research one, the comparison you care about is against the other models at your size, and on that comparison the numbers are strong. What the release does not give you is training code or a recipe you can rerun, which makes this a set of weights to evaluate rather than a method to reproduce.
Frequently asked questions
What is Agents-A1?
A mixture-of-experts agentic model in the roughly 35B class with about 3B active per token, released with weights, evaluation code for selected domains, and a technical report. Its central claim is that scaling the agent horizon, the number of steps a model takes, buys capability usually attributed to parameter count.
How was Agents-A1 trained?
In three stages. First, supervised fine-tuning across all domains to align the base model with broad agentic behaviour. Second, a separate teacher model trained per domain for specialised expertise. Third, a multi-teacher, domain-routed distillation that runs on policy with salient vocabulary alignment to transfer that expertise into one deployable student covering six heterogeneous domains.
What does on-policy distillation mean here?
The distillation samples come from the student's own distribution rather than the teacher's. Distilling from the teacher's outputs teaches the student what the teacher would say; distilling on policy teaches it what it would actually reach given the states it finds itself in. Domain routing then picks which teacher supervises which input.
Does Agents-A1 beat frontier models?
Not across the board, and the readme publishes the rows where it does not. It is the strongest of the models in its own size class on most of the benchmarks listed, and the best overall on several, while sitting below the largest systems on a number of others including in the long-horizon search set that is central to its own thesis.
Can I run Agents-A1 locally?
Quantised variants were released shortly after the main model at several scales, with community-produced versions for Apple hardware, and the readme explicitly invites you to run it on a Mac. A smaller variant was then released after the community asked for one, described as making it easier to build a local assistant.
Is the training code available?
No. The repository contains the readme, assets, documentation, evaluation code for selected domains and scripts, but not the training pipeline. The three-stage recipe and the trajectory-generation infrastructure are described in the technical report rather than shipped as runnable code, so this is a weights release to evaluate rather than a method to reproduce.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internscience-agents-a1)