# SKILL0 trains for skill internalization, and its own setup script still calls itself verl

> An in-context reinforcement learning framework that reports gains over a standard RL baseline on ALFWorld and Search-QA, with two environments, a pinned inference stack and a hand-assembled retrieval index. The packaging metadata is inherited from the framework it builds on.

**ZJU-REAL/SkillZero** — Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"

- Repository: https://github.com/ZJU-REAL/SkillZero
- Website: https://arxiv.org/abs/2604.02268
- Stars: 379 · Forks: 20
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zju-real-skillzero

## The target is skill internalization, tested on two environments

SKILL0 is described as an in-context reinforcement learning framework built for skill internalization, and the claim attached to it is narrow: substantial improvements over the standard reinforcement learning baseline on ALFWorld and Search-QA.

Two environments, two different shapes of problem. ALFWorld is a text-and-actions household environment that also needs a perception stack, since the setup downloads PDDL and game files along with a pre-trained MaskRCNN detector into a cache directory. Search-QA is a retrieval problem that needs a document index and a serving process.

That pairing is the reason the setup is as long as it is. One environment exercises tool use inside a simulator, the other exercises retrieval against a corpus, and a method that claims to internalise skills has to be measured on both.

No task-level scores are published in the write-up. The claim is a comparison against a baseline on the two benchmarks, without the numbers behind it, which is worth noting before planning around a reproduction.

## The install pins vLLM and flash-attn to exact versions

The environment is created on a specific Python version and the inference stack is pinned rather than ranged.

```bash
conda create -n skillzero python=3.12 -y
conda activate skillzero

pip install vllm==0.10.0
pip install flash-attn==2.7.4.post1 --no-build-isolation --no-cache-dir
pip install -e .
```

The two flags on the attention build are not incidental. Disabling build isolation and the cache is what makes a pinned source build of that package workable in a fresh environment.

Logging is opt-in and the note about it is specific: the training scripts pass a logger configuration including both the console and Weights and Biases in many cases, so if you use that service you need a key in the environment before the run rather than discovering the failure mid-training.

That is the whole Python-side setup. Everything that follows is either the ALFWorld environment or the Search environment, and the two are installed differently enough that the instructions treat them as separate sections.

## Search runs in a second conda environment

The ALFWorld side is a short list. A pinned gymnasium, a pinned stable-baselines3 and the environment package itself, followed by a download command that fetches the planning domain files, the game files and the detector weights into a cache directory.

The Search side is where the setup stops being a pip install.

A third-party directory inside the search environment package is installed in editable mode alongside a pinned gym version. Then the dataset is preprocessed by a script run from the repository root, writing to a fixed location under the home directory.

The retriever itself gets its own environment, created on a different Python version than the training one, which is the detail that catches people out. It pins numpy, installs torch, torchvision and torchaudio from a CUDA 12.4 index, adds the transformers, datasets, pyserini and hub libraries, installs a GPU version of FAISS from the PyTorch and NVIDIA channels, and finishes with the web server pair.

So a full Search reproduction touches two Python versions, two package managers and a CUDA index, which is the real cost of the benchmark rather than anything specific to the method.

## The index is assembled by hand from downloaded parts

The retrieval corpus is not a single download, and the instructions show the assembly.

```bash
conda activate retriever

local_dir=~/data/searchR1
python examples/search/searchr1_download.py --local_dir $local_dir
cat $local_dir/part_* > $local_dir/e5_Flat.index
gzip -d $local_dir/wiki-18.jsonl.gz
```

Three steps. A script writes into a chosen local directory. The downloaded parts are concatenated into one flat index file by wildcard. The compressed corpus listing is decompressed in place.

Building the index by concatenating parts rather than by an indexer means the step is cheap and reproducible, but it also means the file has to be reassembled on every machine and there is no checksum in these instructions to confirm the parts arrived intact.

After the index exists, a separate module generates the validation parquet that the training pipeline reads, which is the last Search-specific preparation step before training starts.

## The retriever log is redirected because the terminal caused latency spikes

One instruction in the Search setup carries an explanation rather than a command, and it is the kind of note that only appears when something was measured.

The retrieval server is launched from a shell script under the retriever environment, with its output redirected into a log file instead of the terminal. The reason given is that printing to the terminal was observed to cause spikes in server response times.

That is a small operational fact with a large implication. The retrieval server sits in the request path during training, so a slowdown caused by logging is a slowdown in the training loop, and the fix was to stop writing to a terminal that was evidently slower than a file.

It also means the server is a long-running process you start before training and leave alone, which is the practical reason the instructions treat it as a separate step from the index build.

The last piece of Search preparation is generating the validation parquet, after which the training scripts become the entry point.

## Training scripts change directory themselves

The training entry points are two shell scripts under a scripts directory, one per benchmark.

```bash
bash scripts/train_alfworld_skillzero_3b.sh
bash scripts/train_search_skillzero_3b
```

Both file names carry the same suffix, which reads as a model size designation, and both assume the repository root as the working directory. The note says the scripts change into it automatically, so they can be invoked from anywhere.

That assumption has a consequence for anyone customising a run. If a script depends on relative paths it sets up itself, editing it to point at a different data root means checking what it does first, because the working directory it establishes is part of the contract.

Checkpoint merging is a separate script rather than a flag on training. It contains FSDP and Megatron merge examples operating on paths under a checkpoints directory, which is the expected shape once distributed training is producing sharded checkpoints.

The optional logging detail appears again here: the trainer logger configuration including Weights and Biases is passed in many of these scripts, so a run started without a key configured will fail at the point where it tries to log.

## The packaging metadata is inherited from veRL

The most surprising thing in the repository is not in the documentation.

The project manifest names the package verl and describes it as reinforcement learning for large language models from an engine vendor, with the version read from a file inside a vendored verl directory. The fallback installation script carries a copyright header from a different company entirely and installs a dependency set built for that upstream framework: ray, tensordict, transformers, peft, datasets, and optional groups for tests, GPU kernels, maths verification, vLLM, SGLang and TRL.

So this repository is a derivative of that framework, vendored whole, with the skill internalisation work layered on top. The acknowledgements say as much, listing that framework and its agent variant alongside the environments and the earlier skill reinforcement learning work.

Two consequences are worth knowing. A dependency installed by file and one installed by package are not the same here, because the file version pins a transformer release the package version caps differently, and a tensor library range that differs between the two. And the linter configuration is inherited too, including a line length set to three hundred with a note in the file that it should be reduced.

## Conclusion

SKILL0 fits a research group reproducing the skill internalization line of work, since the environment setup for both benchmarks is written down step by step rather than abstracted. It does not fit anyone wanting a library to import: this is a training repository with pinned inference builds and a retrieval server to stand up first. Two things to check before starting. The Search path needs a second conda environment and an index assembled by concatenating downloaded parts, so budget time for it. And the dependency files disagree with each other on transformer and tensor library versions, which means installing by file and installing by package gives different environments.

## FAQ

### What does SKILL0 actually do?

It is an in-context reinforcement learning framework aimed at skill internalization, reported to achieve substantial improvements over the standard RL baseline on ALFWorld and Search-QA.

### Which Python and inference versions does SkillZero need?

A conda environment on Python 3.12, with vLLM pinned to 0.10.0 and flash-attn pinned to 2.7.4.post1 installed with build isolation and cache disabled, followed by an editable install of the package.

### Why does the SkillZero search setup need its own conda environment?

The retriever has its own dependency set, so it is created on a different Python version with pinned numpy, torch from a CUDA 12.4 index, a GPU FAISS build and a web server pair, separate from the training environment.

### How is the SkillZero retrieval index built?

A download script writes index parts into a local directory, the parts are concatenated with a wildcard into a single flat e5 index file, and the compressed corpus listing is decompressed in the same place.

### Why does SkillZero redirect the retriever output to a file?

Because printing the retrieval server's output to the terminal was observed to cause spikes in its response times. The server sits in the request path during training, so a logging slowdown becomes a training slowdown.

### Which projects does SkillZero build on?

AgentOCR, verl-agent, veRL, ALFWorld, SkillRL and Search-R1. The veRL framework is vendored into the repository, and the packaging metadata still carries its name and dependency set.

## Sources

- [Issues](https://github.com/ZJU-REAL/SkillZero/issues)
- [License: Apache-2.0](https://github.com/ZJU-REAL/SkillZero/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2604.02268)
- [README](https://github.com/ZJU-REAL/SkillZero/blob/main/README.md)
- [ZJU-REAL/SkillZero on GitHub](https://github.com/ZJU-REAL/SkillZero)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zju-real-skillzero
