SkillZero: in-context agentic RL for skill internalization
Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"
At a glance
- What is it?
- SkillZero is the official research code for SKILL0, a framework that trains agents to internalize skills during reinforcement learning instead of carrying them in the prompt. It ships training scripts for ALFWorld and Search-QA, and it is a research release, not a product.
- Who is it for?
- SkillZero is for researchers who already run GPU RL training and want to reproduce the SKILL0 recipe on ALFWorld or Search-QA. It is not for anyone who wants a packaged agent library, because the install spans two conda environments and a local retrieval server.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 50 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SkillZero solves, and who the repository is actually for
The SKILL0 paper, linked from the repository, frames the problem as skill internalization. An agent that depends on skills written into its context pays for them on every step, and the prompt grows with the skill library. SkillZero is the code that accompanies the paper, and its aim is to move that skill knowledge into the policy through in-context agentic reinforcement learning rather than keeping it in the prompt.
The README states that SKILL0 achieves improvements over the standard RL baseline on ALFWorld and Search-QA, and points to a metrics figure under docs/skillzero/. Those two environments are the scope. ALFWorld is a text-based household task suite, and Search-QA is a retrieval-backed question answering setup.
The audience is narrow on purpose. This is a training repository built on a vendored copy of verl, the Volcano Engine Reinforcement Learning library, which is visible in pyproject.toml under the name verl and in the top-level verl/ directory. If you are not training a policy, most of the tree is irrelevant to you. The README also lists a series of follow-up releases from the same group, including SDAR, SKILL1, SEED, OPID, SkillRise and AgentOPSD, each pointing at a separate repository. That list is useful context: SkillZero is one point in a fast-moving line of work, not the final word.
How the training loop is put together
The repository is a fork of a verl-style RL stack. pyproject.toml declares the project name as verl and reads its version from verl/version/version, and setup.py carries the Bytedance copyright header and the Apache-2.0 notice. The dependency list in setup.py names ray, tensordict, transformers, hydra-core and wandb, which is the shape of a Ray-based distributed trainer configured through Hydra.
The examples/ directory shows the trainer variants that ship with the stack: grpo_trainer, ppo_trainer, rloo_trainer, gspo_trainer, dapo_trainer, sft, and the gigpo variants. SkillZero's own recipe sits alongside those, and the scripts/ directory holds the entry points. The README says all scripts live under scripts/ and assume the repository root as the working directory, because they change into it automatically.
The environment side is split into packages under agent_system/environments/env_package/. ALFWorld installs from PyPI. Search has its own third_party package plus a separate retrieval service, which is why the install section needs two conda environments. The README documents a retriever environment with pyserini and faiss-gpu, and a launch script at examples/search/retriever/retrieval_launch.sh. That split is the most important architectural fact about running Search training: the policy trainer and the retriever are separate processes with separate Python versions, 3.12 for the trainer and 3.10 for the retriever.
Installing SkillZero and running the first training script
The README gives a conda-based install. The trainer environment is Python 3.12, and two packages are pinned: vllm at 0.10.0 and flash-attn at 2.7.4.post1. flash-attn is installed with --no-build-isolation and --no-cache-dir, which is the usual way to get it to compile against an existing torch install.
conda create -n skillzero python=3.12 -y
conda activate skillzero
pip install vllm==0.10.0
pip install flash-attn==2.7.4.post1 --no-build-isolation --no-cache-dir
pip install -e .The final command installs the repository itself in editable mode. If you use Weights and Biases logging, the README notes that many scripts pass trainer.logger=['console','wandb'], and it exports the key rather than storing it in a config file.
export WANDB_API_KEY=your_key_hereFor ALFWorld, three pip installs and one data download are listed. The downloader writes PDDL and game files plus a pre-trained MaskRCNN detector into ~/.cache/alfworld/.
pip3 install gymnasium==0.29.1
pip3 install stable-baselines3==2.6.0
pip3 install alfworld
alfworld-download -fSearch needs more. The third_party package installs in editable mode, gym is pinned to 0.26.2, and a preprocessing script produces the dataset under ~/data/searchR1_processed_direct. Note that the README writes this path with a tilde, so it resolves relative to the home directory of whoever runs it.
cd ./agent_system/environments/env_package/search/third_party
pip install -e .
pip install gym==0.26.2
cd repo_root/
python examples/data_preprocess/preprocess_search_r1_dataset.pyThe retriever is a second environment. It pins numpy 1.26.4, torch 2.6.0 with CUDA 12.4 wheels, and faiss-gpu 1.8.0 from the pytorch and nvidia channels.
conda create -n retriever python=3.10 -y
conda activate retriever
conda install numpy==1.26.4
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install transformers datasets pyserini huggingface_hub
conda install faiss-gpu==1.8.0 -c pytorch -c nvidia -y
pip install uvicorn fastapiThe index download follows, and the README splits the parts and decompresses the wiki dump in place. Then the retrieval server starts from the launch script.
conda activate retriever
local_dir=~/data/searchR1
python examples/search/searchr1_download.py --local_dir $local_dir
cat $local_dir/part_* > $local_dir/e5_Flat.index
gzip -d $local_dir/wiki-18.jsonl.gz
bash examples/search/retriever/retrieval_launch.sh > retrieval_server.logThe README explains the redirect: output to the terminal was observed to cause spikes in server response times. That is a concrete operational note, and it is worth following rather than treating the log file as cosmetic. After that, one more preprocessing command generates the validation parquet.
python -m examples.data_preprocess.generate_search_r1_valTraining itself is a script call from the repository root. The README gives the ALFWorld script and a Search script, and the Search line appears without a .sh suffix in the README text, so check the actual filename under scripts/ before running it.
bash scripts/train_alfworld_skillzero_3b.sh
bash scripts/train_search_skillzero_3bCheckpoint merging is covered separately. The README points to scripts/model_merger.py for FSDP and Megatron merge examples, with paths under ./checkpoints/.
Where the setup breaks, and where SkillZero is the wrong tool
The install is the weak point, and the README is honest about the shape of it without warning you about the consequences. Two conda environments with different Python versions and different torch versions have to coexist on one machine, and the retriever needs a GPU for faiss-gpu. If you only have one GPU, the retriever and the vLLM rollout engine are competing for it, and nothing in the README describes a CPU-only retriever path.
Version drift is the second failure mode. The trainer pins vllm==0.10.0 and flash-attn==2.7.4.post1, while requirements.txt in the repository lists transformers==4.51.1 and a commented-out vllm==0.8.4, and setup.py allows transformers up to 4.57.3 and vllm from 0.8.5 to 0.11.0. Those three files do not agree. The README install sequence is the one to follow, because it is the sequence the authors wrote for the paper code, but a reader who runs pip install -r requirements.txt instead will get a different and probably incompatible stack.
There is no packaging story. No recent releases were retrieved, and the install is from source with pip install -e . against a tree that vendors verl. There is no CLI, no published container image in the repository, and no documented way to run inference on a trained checkpoint other than the merger script. If your goal is to use a skill-augmented agent rather than train one, this repository gives you nothing to call.
The README also does not document rollback, checkpoint resume semantics, or how to recover a partially completed training run. The model_merger.py pointer is the only post-training tooling mentioned. Treat any long run as something you need to instrument yourself.
How SkillZero differs from the other skill-augmented RL repositories
The README's news section is effectively a comparison list, and it is the most useful part of the page for positioning. SkillZero is skill internalization through in-context agentic RL. SDAR, released 2026-05-15, is described as self-distilled agentic reinforcement learning for the same goal. SKILL1, released 2026-05-07, takes a different route: it evolves skill-augmented agents in one unified policy. OPID, released 2026-06-25, is about skill evolving beyond internalization, and SEED, released 2026-07-17, adds self-evolving on-policy distillation beyond skill internalization. SkillRise, released 2026-07-29, adds cross-task skill evolution, and AgentOPSD, released 2026-08-06, adds recursive credit update for SDAR.
Read that sequence as a research program rather than a set of interchangeable tools. If your question is whether an agent can learn without the skill text in context, SkillZero and SDAR are the closest matches. If your question is how skills should change across tasks, SkillRise and OPID are the later answers from the same group. The practical difference for an adopter is which repository's training scripts and environments you are willing to set up, because each one carries its own dependency pins and its own environment packages. SkillZero's own scope stops at ALFWorld and Search-QA, and the README does not claim support for other environments.
Maintenance, licence and what an upgrade costs you
The repository is not archived, and the last push was on 2026-08-12. That is recent, but the README's news list shows the group's attention moving to newer repositories through August 2026, so SkillZero is best read as a paper artifact that receives occasional updates rather than a project with a support commitment. There are no retrieved releases, so there is no versioned artifact to pin against. Your reference point is a commit hash on main.
The licence is Apache-2.0, and the LICENSE file is at the repository root. setup.py carries the Apache header from Bytedance, and pyproject.toml points its license field at the LICENSE file. Because the tree vendors verl, the Notice.txt file at the root is worth reading before you redistribute anything; that is where third-party attributions usually live. This is not legal advice, and if you plan to ship a derivative, read LICENSE and Notice.txt yourself.
The upgrade cost is the real constraint. Dependencies are pinned at specific versions in the README, in requirements.txt and in setup.py, and those pins disagree. Moving to a newer vLLM or transformers means re-validating the flash-attn build, the retriever's torch 2.6.0 install, and the faiss-gpu channel install, which is a multi-hour job on a fresh machine. Budget for that before you plan any version bump. If you need a stable artifact, pin the commit and keep the two conda environments frozen.
Editorial conclusion
SkillZero is for researchers who already run GPU RL training and want to reproduce the SKILL0 recipe on ALFWorld or Search-QA. It is not for anyone who wants a packaged agent library, because the install spans two conda environments and a local retrieval server. Verify the retriever server responds before launching training, and check that your GPU stack matches the pinned vllm and flash-attn versions.
Frequently asked questions
What is SkillZero training?
SkillZero is the official code for SKILL0, an in-context reinforcement learning framework for skill internalization. Training runs through scripts under scripts/, such as scripts/train_alfworld_skillzero_3b.sh, and the README reports improvements over the standard RL baseline on ALFWorld and Search-QA.
What environments does SkillZero support?
The README documents two: ALFWorld, installed with pip along with gymnasium and stable-baselines3, and Search, which needs a third_party package plus a separate retriever environment running a flat e5 server. No other environments are claimed.
How do I install SkillZero?
Create a conda environment with Python 3.12, install vllm==0.10.0 and flash-attn==2.7.4.post1, then run pip install -e . from the repository root. Search training also needs a second conda environment with Python 3.10 for the retriever.
What is the SkillZero licence?
The repository is Apache-2.0, with the LICENSE file at the root and the Apache header in setup.py. Because the tree vendors verl, the Notice.txt file at the root should be read before redistributing anything.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zju-real-skillzero)