Model or dataset
XYZ-AI-Lab/AxisAgentic avatar
XYZ-AI-Lab/AxisAgentic

The claim is about the exported data, not the model, and the results section argues against its own chart

AxisAgentic: An Extensible Runtime and Trajectory-Collection Framework for Long-Horizon Agents.

1,120 stars143 forksPythonApache-2.0

At a glance

What is it?
A runtime for long-horizon agents that writes an append-only trace of everything the model saw, so the same record serves recovery, replay, benchmark evaluation and training-data export. The exported examples are supposed to exclude rolled-back actions and hidden history. Type checking is configured off in three places, and the results carry an explicit warning that they are not a controlled ranking.
Who is it for?
The interesting part of this framework is not the agent loop, it is the trail. Recording what the model actually saw at each stage, including what was rolled back and what was compacted away, and then requiring the training export to respect the same visibility rules, is a claim about data integrity that most agent frameworks do not make at all.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 70 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The correctness claim is about the exported data, not the model

Read the capability list and this looks like a general agent runtime. Read the middle section and it becomes something narrower and more interesting.

Every task writes an append-only trace of runtime events, and the trace is described as the common source for replay, evaluation, and trajectory collection. The same record therefore serves three jobs: recovering a failed run, scoring a run, and producing training data from it.

The load-bearing sentence is about export. Runtime markers record rollback, context compaction, and discard-all events, and replaying them reconstructs what the model saw at a given stage. Trace inspection and the supervised fine-tuning export then use the same rules, which the document says means supervised examples exclude hidden history and rolled-back actions.

That is a data-integrity claim, and it is the kind almost nobody makes. An agent run contains steps the model never acted on and context it could not see. Export all of it and you train on a history that never happened. This framework says its exporter cannot, because the exporter reads the same markers the replay does.

The scope is stated too, and it is sensible. Recipe exporters emit training formats along with the source trace, task status, and optional metadata, and the external training pipeline owns final correctness filters, loss masks, and optimization. So the framework supplies the replay and export boundary, and the training side owns everything after it. The claim is that inference and training use the same interaction history, not that the resulting model is good.

Three type checkers, all configured to be quiet

The tooling configuration is the most surprising file in the repository, and it is worth reading as a whole rather than line by line.

The linter is configured to select every rule that exists, with a line length of a hundred and fifty characters and a target of a recent Python version. Then comes an ignore list. It removes the docstring rules for packages, modules, classes, methods and functions, plus the initialiser, which is normal. It also removes the rule against dynamically typed expressions, the rule about deserialising untrusted data through pickle, the rule against catching blind exceptions, the rule about f-strings in logging calls, the rule about commented-out code, and the rule about implicit string concatenation.

So: maximum coverage, with the noisy and the occasionally-inconvenient turned off. That is a coherent position. A hundred and fifty characters is long, and a limit of twenty for cyclomatic complexity, twenty branches, fifteen returns and fifteen arguments is generous for a runtime with an orchestrator in it.

The type checkers are where the position stops being coherent. Three are configured. One of them disables attribute-defined errors, type errors in arguments, and type errors in has-type checks, and ignores missing imports entirely. Another suppresses six separate report categories covering general type issues, unknown variable types, unknown argument types, unknown member types, and abstract usage.

Two tools pointed at the same code with different suppressions is a maintenance cost rather than a virtue. The likely reason is defensible: this is a framework that consumes model output, and strict attribute and argument checking fights code that handles whatever a model returned. But the effect is that the type layer contributes little, and a reader who sees every linter rule enabled may reasonably assume the same of the type checkers.

Packaging lives in one file and tool configuration lives in another

There are two configuration files at the root and they do completely different jobs, which is not the usual arrangement.

The project file contains no package definition at all. It holds the build requirement, the linter settings, and the two type-checker settings. The actual package is defined in a setup file, which carries a licence header comment, a copyright line, and the full call: the distribution name, a version of zero point one point zero, the discovered packages for two source trees, package data, requirements, extras, the minimum Python version, a description, a licence, an author, a URL, and a project URL table pointing at the documentation and a technical report.

Three details in that file are informative.

The distribution name is not the repository name. The repository is called one thing and the installed package is called something else, which is a normal and slightly annoying decision that anyone automating a dependency will notice.

The package data mapping is per recipe. Configuration files under a config directory are shipped with the main package and with model asset files, and each of two named recipes ships its own configuration directory. That means a recipe is not just code: it is code plus a directory of configuration that has to travel with it, and the packaging file is where that is declared.

And the runtime requirement list is short: an async file library, an HTTP client, a schema validator, a YAML parser, a data validation library, and the model client. Among those, the JSON repair library is the one to notice. It is not a general-purpose utility; it exists to fix malformed JSON, which tells you the runtime parses structured output from models routinely and has decided to repair rather than fail. That is a design decision with a name, and it is more informative than it looks.

The environment is built by a script you have to source, and there is a dry run before anything starts

Setup is six commands, and two of them are unusual.

Clone the repository, change into it, create a virtual environment with an explicit interpreter version, activate it, run a setup script, then activate a second script that the first one created, then copy an example environment file into a directory inside the project. The requirement is Python three point twelve or newer and an OpenAI-compatible model endpoint.

The second activation is the odd part. The project creates its own activation script under a directory inside the repository and tells you to source it, so the environment is composed of a virtual environment you made and a shell fragment the project made. That is how a configuration that the Python installer does not read gets into your shell.

The environment example file explains the split in two lines. Runtime settings, and it names them, belong in YAML configuration: dataset paths, output directories, concurrency, token limits, timeouts, judge settings, and tool limits. Provider default model names may live in the environment file and can be overridden by YAML. So the precedence is: the environment file supplies keys and default model names, YAML supplies everything operational, and YAML wins.

Then there is the dry run, which is the part to use:

bash
cp recipe/web_search/configs/default.yaml my-search-run.yaml
python -m recipe.web_search.runners.run_eval_config \
  --config my-search-run.yaml \
  --dry-run

The document says this validates the recipe without starting a run. For a framework whose central claim is about what a run records, being able to validate the configuration without producing a run is a reasonable thing to offer, and it is the first command to reach for.

Seven provider blocks, and the two auxiliary model roles default to a loopback port

The environment example is organised as seven provider blocks, and reading them in order tells you what a single run talks to.

First the main model: a name, a base address defaulting to a public API, and a key. Then a web search provider with its own key and its own base address for a specific third-party search service. Then a web scrape provider, again with a key and a specific base address. Then a hosted code sandbox, which has a key and nothing else. That is four external services before any model is asked to summarise anything.

Then three more, and these are the interesting ones. A summarisation model role, a context-compression model role, and an evaluation judge. The first two are marked optional and both default to a base address on a local port with a chat-completions path. The judge is also optional and is stated to default to the main provider.

So the asymmetry is explicit in the file: point this at a local model server and your summariser and your compactor run there, while your judge runs on whatever paid endpoint you configured. That is a sensible default for cost on the two hot paths and an expensive one for the third, since a judge is called once per example and a summariser is called many times per task. A deployment that has not thought about it will pay for judging and not for summarisation.

Worth noting as well that the sandbox role is a third-party hosted interpreter with a single key, and that a run which scrapes and searches will send query text to at least two services other than the model endpoint.

The results section argues against its own chart

The benchmark section is where a project is most likely to overstate, so it is worth reading what it actually claims.

A technical report covers two sizes of the flagship search system across seven agentic benchmarks, and the document reproduces comparisons for six of them. The metrics are then itemised per benchmark, and the itemisation is where the caution lives: four benchmarks use a language model as the judge, one uses a macro-averaged score, and one uses an item-level score at four attempts. The set also contains an English and a Chinese variant of the same benchmark counted as two, so seven benchmarks is closer to six distinct questions.

Then the caveat, and it is the right one: some baseline values come from public reports gathered with different harnesses, different tools, different judges, and different evaluation dates, and the figure should be read as a benchmark-level comparison rather than a controlled universal ranking.

A project could easily have printed the chart and stopped. Printing the caveat, in those terms, tells you the authors know exactly which comparison they are making and which one they are not. It also means the numbers cannot be used to settle a ranking argument, and a reader who is evaluating this framework should treat them as evidence that the system runs and produces reasonable answers, not as evidence that it is ahead of anything.

There is a separate document on evaluation and reproducibility, which the footnote points to for the metric definitions, and that is where a serious reader should go next.

Recipes are directories with their own configuration, and the flagship system is not in the repository

The extensibility claim is structural, and the repository layout is how you can see it.

The source is two trees: one for the runtime and one for recipes. The root also carries a configuration directory that is separate from both, which suggests shared defaults that are not owned by any single recipe. Recipe exporters and recipe policies are both named as replaceable alongside model clients, tools, orchestrators, datasets, evaluators, and reward functions, so the swap points are a documented list rather than a convention.

Two recipes are named as the current references, one for web search and one for a wider search task, and both are wired into packaging with their own configuration directories. A recipe therefore ships as code plus configuration, and the two reference recipes are the templates a new one would copy. The document also says the same extension points can support domain-specific, general-purpose, and coding agents, which is the claim, and having two search-shaped examples is the evidence offered for it.

What is not in the repository matters as much. There are no model weights, stated explicitly. And the flagship system, a search recipe that combines search and scraping with context management, recovery, evaluation, and state-faithful export, is described as built with the framework but lives outside it, with its results in a technical report. So what you can evaluate is the runtime and two recipes.

The surrounding files are the conventional ones for a project that wants to be cited: a changelog, a citation file, a licence, a notice file for third-party attribution, a code of conduct, a contributing guide, a security policy, and a pre-commit configuration.

Editorial conclusion

The interesting part of this framework is not the agent loop, it is the trail. Recording what the model actually saw at each stage, including what was rolled back and what was compacted away, and then requiring the training export to respect the same visibility rules, is a claim about data integrity that most agent frameworks do not make at all. If you are generating training data from agent runs, that is the feature to evaluate, and the dry-run path plus the export inspection is how you would check it. Three cautions. The framework does not include model weights and the flagship search system is not in the repository, so what you are evaluating is the runtime. The two auxiliary model roles, summarisation and context compression, default to a local endpoint while the judge defaults to the main paid provider, which is a cost and data-path asymmetry worth deciding deliberately. And the benchmark figure is described in the document itself as a comparison rather than a controlled ranking, with baselines gathered under different harnesses and judges, so the numbers establish that the system works rather than where it ranks.

Frequently asked questions

What is AxisAgentic?

An extensible runtime for long-horizon AI agents that also collects the trajectories produced while they run. It works with OpenAI-compatible endpoints and pluggable local model clients, and handles multi-turn execution, tool orchestration, context management, recovery and benchmark evaluation. Each trace preserves the state visible to the model, so one record can serve recovery, replay, evaluation or training-data export. The repository does not include model weights.

How does AxisAgentic keep rolled-back actions out of its training data?

It records runtime markers for rollback, context compaction and discard-all events in an append-only trace, and replaying those markers reconstructs what the model saw at a given stage. Trace inspection and the supervised fine-tuning export use the same rules, so exported examples exclude hidden history and rolled-back actions. The external training pipeline still owns final correctness filters, loss masks and optimization.

What do I need to run AxisAgentic?

Python 3.12 or newer and an OpenAI-compatible model endpoint. Setup clones the repository, creates a virtual environment with that interpreter, runs a setup script, sources a second script that script created, and copies an example environment file into a directory inside the project. Runtime settings such as paths, concurrency, token limits and timeouts belong in YAML rather than the environment file.

Can I validate a recipe without running an agent?

Yes. Copy a recipe configuration to your own filename, run the recipe's evaluator module against it, and pass a dry-run flag. The document describes this as validating the recipe without starting a run, and it is the first command worth reaching for when a configuration is misbehaving.

How trustworthy are the benchmark results reported for AxisAgentic?

The document itself is cautious. Two sizes of the flagship search system are reported across seven benchmarks with the figure reproducing six, using four different metric definitions including an English and a Chinese variant of the same benchmark. It states that some baseline values come from public reports with different harnesses, tools, judges and evaluation dates, and that the figure is a benchmark-level comparison rather than a controlled universal ranking.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. XYZ-AI-Lab/AxisAgentic on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xyz-ai-lab-axisagentic.svg)](https://hysenlabs.com/projects/xyz-ai-lab-axisagentic)