Model or dataset
aisa-group/PostTrainBench avatar
aisa-group/PostTrainBench

PostTrainBench says Harbor support exists and, further down, that it is planned

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

581 stars67 forksPythonMIT

At a glance

What is it?
A benchmark that gives a CLI coding agent ten hours on one H100 and a base model, then scores the post-trained model on six tasks. Its manifest is a stub with a placeholder description, its documented job submission lines end in an ellipsis, and its two statements about Harbor contradict each other.
Who is it for?
Use it as a research harness for agentic post-training research, not as a turnkey benchmark you can run on a laptop: the documented path is an HTCondor cluster, the manifest is a stub, and the one cloud path is described both as working and as unbuilt. Before running anything, decide whether you can accept the credential handling, since two scaffolds require copying a subscription auth file or a long-lived OAuth token into a directory the scheduler copies into jobs.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The page says Harbor works, then says Harbor is planned

The prominent callout near the top instructs readers to run the benchmark on cloud GPUs through Harbor, and says that while the reference pipeline targets the authors' HPC cluster on HTCondor, `src/harbor_adapter/` runs the full benchmark on Modal with the same prompt, judges and evaluation and no cluster needed.

Further down, in the environment section, the same page says that only the HTCondor job scheduler is supported and that Harbor support is planned.

Those cannot both describe the current state of the tree. The directory table lists `src/harbor_adapter/` as the component that runs the benchmark on cloud GPUs through Harbor on Modal, with its own README, so the code appears to be there. What the configuration table says is narrower: `POST_TRAIN_BENCH_JOB_SCHEDULER` accepts `htcondor` or `htcondor_mpi-is`, and neither option is Harbor or Modal.

So the honest reading is that a cloud adapter exists in the tree, the shipped scheduler variable has no value for it, and the documentation contradicts itself about whether that adapter is supported today. Anyone planning around the cloud path has to read `src/harbor_adapter/README.md` and find out for themselves.

The manifest is a stub with a placeholder description

`pyproject.toml` for this benchmark is four meaningful lines. The project is named `posttrainbench` at version 0.1.0, the description is the literal placeholder `Add your description here`, and it requires Python 3.10 or newer.

The dependency list has two entries: `modal>=1.3.0.post1` and `python-socks>=3.0.0`. Note what that implies. The reference execution path is an HTCondor cluster with Apptainer containers, yet the only substantial dependency is Modal's client library, which is what the cloud adapter needs. There is no PyTorch, no transformers, no datasets and no evaluation library, because the work is not done by a Python application: it is done by shell scripts that build a container, download a Hugging Face cache and submit jobs to a cluster scheduler.

The rest of the repository matches that shape. Alongside `src/` and `agents/` there are `containers/`, `dev_utils/`, `scripts/`, an `example.env`, a `.python-version` and a `uv.lock`, so the Python environment is uv-managed even though the entry points are bash. The repository publishes no releases, the last push is dated 2026-10-01, and it is MIT licensed.

Two scaffolds require a credential file on disk

Most agents authenticate with API keys from the environment, and the page explains those are passed into the container automatically by `run_task.sh`. Two scaffolds are different, because the models they drive are only reachable through a subscription rather than an API key, with GPT-5.3-Codex through ChatGPT Pro given as the example.

For the Codex scaffold, the setup is: enable device code login in the ChatGPT security settings, run `codex login --device-auth`, then copy the resulting credentials into the agent directory:

bash
cp ~/.codex/auth.json agents/codex_non_api/auth.json && chmod 600 agents/codex_non_api/auth.json

For the Claude Code scaffold, the route is a long-lived OAuth token with roughly a year of validity, generated by `claude setup-token` and written to a file:

bash
echo "sk-ant-..." > agents/claude_non_api/oauth_token

So both require placing a personal subscription credential inside the repository tree, in a path the scheduler copies into the job directory. The page does address the obvious concern: `auth.json` and `oauth_token` are gitignored, and `run_task.sh` copies them conditionally, only for the agents that need them. What it does not discuss is the lifetime of the Claude token, which is about a year, or what a scheduled job can do with it.

solve.sh unsets your API keys so the subscription path is used

The reason for the credential juggling is stated in each agent's `solve.sh`. The Codex script unsets the API keys and sets `forced_login_method = "chatgpt"` so the CLI uses the copied credentials instead. The Claude script reads the token from its file, exports it as `CLAUDE_CODE_OAUTH_TOKEN`, and unsets `ANTHROPIC_API_KEY` to avoid auth conflicts.

That is a benchmark-integrity measure as much as a convenience. Without it, a subscription-based agent would silently fall back to a metered API key present in the environment, and a score attributed to one billing path would have come from another.

The job submission lines, meanwhile, are documented with an ellipsis:

bash
condor_submit_bid 50 -a "agent=codex_non_api" -a "agent_config=gpt-5.3-codex" ...

The Claude example is the same shape with `agent=claude_non_api` and `agent_config=claude-opus-4-6`. So the attribute style is visible, a 50-slot bid is visible, and the `agent_config` value names the model under test, but neither command is complete enough to run as printed.

The agent gets the evaluation script and unrestricted internet

The full agent prompt is published in the page, and it is permissive. The agent is told to post-train a small model to excel at the named benchmark through systematic research and experimentation, that it has complete freedom in its approach including data sources and training methods, that it can run multiple iterations, and that internet access is unrestricted. It is told it can query the benchmark through the `evaluate.py` script, and it must store its best trained model in a `final_model` folder.

The repository also removes one guess from the loop on purpose. Each evaluation folder under `src/eval/tasks/` may contain a `task_context/` directory holding information about exactly how the evaluation is performed, described as existing so the agent does not have to guess.

Both choices are defensible for measuring research ability, and both narrow what the number means. An agent with direct access to the scoring script and the network can iterate against the metric rather than against the task, and an agent handed the evaluation details is not being scored on discovering them. The score is a measure of post-training skill under those conditions, which is a narrower claim than post-training skill in general.

Four scaffolds are named, a fifth agent appears in a key description

The scaffolds section says agents run through one of four CLI scaffolds: Claude Code, Codex CLI, Gemini CLI and OpenCode. The environment table then describes `OPENCODE_API_KEY` as used by the `opencode` agent and `ZAI_API_KEY` as used by the `opencode` and `glm5` agents.

So a fifth agent, `glm5`, is named in the configuration documentation and in the directory listing through its variant names, but it is absent from the list of four scaffolds. Either it is a variant of one of them rather than a separate scaffold, or the count is stale. The variant names are the clue: the submission lines pass `agent=codex_non_api` and `agent=claude_non_api`, so the naming convention is agent plus authentication mode, which makes `glm5` look like a distinct agent rather than a variant of OpenCode.

The four API keys in the table also imply a specific set of vendors to have accounts with before anything runs: OpenAI, Anthropic, Google Gemini, OpenCode and Z.AI.

What each scaffold is measured against is chosen per run through `agent_config`, with values such as `gpt-5.3-codex` and `claude-opus-4-6` shown in the examples.

Six benchmarks inside a ten-hour H100 budget

The task envelope is small on purpose. The agent receives an evaluation script and ten hours on an H100, and the measurement is the benchmark score of the model it produced. The page argues that this setup naturally evaluates an agent's ability to conduct AI research and development, which is a claim about the task's shape rather than about any particular run.

The six benchmarks span five named areas, reasoning, knowledge, math, health and code: AIME 2025 for competition mathematics, Arena Hard Writing adapted from ArenaHard v2, GPQA for graduate-level science, GSM8K for grade school arithmetic, HealthBench Easy for medical knowledge and reasoning, and HumanEval for code generation. Math appears twice by construction, since one is a competition set and one is grade school.

The rest of the layout supports comparing runs. `results/` holds evaluation results with baseline runs prefixed `baseline_`, and `src/baselines/` holds the scripts that compute those baseline scores, so an agent's number has something to be measured against. `logs/` holds HTCondor scheduler output per job with `.err`, `.out` and `.log` files and is gitignored. Everything routes through three environment variables that name the container, the results directory and the containers directory, with the container name defaulting to `standard` and the prompt variant defaulting to `prompt`.

Editorial conclusion

Use it as a research harness for agentic post-training research, not as a turnkey benchmark you can run on a laptop: the documented path is an HTCondor cluster, the manifest is a stub, and the one cloud path is described both as working and as unbuilt. Before running anything, decide whether you can accept the credential handling, since two scaffolds require copying a subscription auth file or a long-lived OAuth token into a directory the scheduler copies into jobs. And read the published agent prompt first, since unrestricted internet access and a direct evaluation script shape what the resulting score means.

Frequently asked questions

What does PostTrainBench measure?

Whether a CLI coding agent can post-train a base LLM. The agent gets an evaluation script and 10 hours on an H100, and the score is the benchmark score of the model it produces. The framing is that this evaluates an agent's ability to conduct AI research and development.

Can PostTrainBench run without an HPC cluster?

The page says both things. A callout says the src/harbor_adapter directory runs the full benchmark on Modal with no cluster, and a later line says only HTCondor is supported with Harbor support planned. The scheduler variable accepts htcondor or htcondor_mpi-is, neither of which is Harbor.

Which benchmarks does PostTrainBench use?

Six: AIME 2025, Arena Hard Writing adapted from ArenaHard v2, GPQA, GSM8K, HealthBench Easy and HumanEval, spanning reasoning, knowledge, math, health and code. Baseline scores for comparison come from src/baselines, with baseline runs prefixed baseline_ in the results directory.

How do the subscription-based agents authenticate?

Codex uses device code login and then copies ~/.codex/auth.json into agents/codex_non_api with mode 600. Claude Code uses a long-lived OAuth token from `claude setup-token` written to agents/claude_non_api/oauth_token. Both files are gitignored and copied into job directories only for the agents that need them.

What can the agent do inside a PostTrainBench run?

The published prompt grants complete freedom of approach including data sources and training methods, allows multiple iterations, states that internet access is unrestricted, and lets the agent query the benchmark through evaluate.py. Evaluation folders may also include a task_context directory so the agent does not have to guess how scoring works.

Which agent scaffolds does PostTrainBench support?

Four are named: Claude Code, Codex CLI, Gemini CLI and OpenCode. The environment table also mentions a glm5 agent alongside opencode for the Z.AI key, so the count may be out of date. Jobs are submitted to HTCondor with condor_submit_bid and an agent_config attribute naming the model.

Official sources

  1. aisa-group/PostTrainBench on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aisa-group-posttrainbench.svg)](https://hysenlabs.com/projects/aisa-group-posttrainbench)