# imitation-learning: eight algorithm names, one of which is a configuration value

> A research library that reimplements six imitation algorithms on a shared soft actor-critic base, with the variant of each algorithm living in a config key rather than in a separate codebase. Its Results section is empty and its release links point above the repository root.

**Kaixhin/imitation-learning** — Imitation learning algorithms

- Repository: https://github.com/Kaixhin/imitation-learning
- Stars: 574 · Forks: 44
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kaixhin-imitation-learning

## Eight selectable algorithms, two of which are baselines

The library's scope is stated in one line: imitation learning algorithms, with soft actor-critic as the base reinforcement learning algorithm underneath all of them. Six methods are named, AdRIL, DRIL, GAIL, GMMIL, PWIL and RED, and each is qualified by the variant that was ported, since DRIL is the dropout version, PWIL is the version without the fill step, and GAIL is also known as DAC or SAM when the base algorithm is off-policy. Then there are eight names in the actual command interface, not six, because behaviour cloning and plain SAC are selectable through the same key as the baselines the six are compared against. Environments are equally fixed: four MuJoCo locomotion tasks. The run command is one line, and the algorithm and environment are the only two things you have to name.

```sh
python train.py algorithm=GAIL env=hopper
```

## Four of the six implementations are ports, and the page credits each one

The acknowledgements section names five upstream repositories, and reading them tells you the provenance of nearly everything here. One credit is for a soft actor-critic implementation with a GAIL discriminator attached. The second is the author's own reference implementation of RED. The third is the Google research tree that holds DAC, which is the off-policy face of GAIL. The fourth is a project called pillbox, which supplies GMMIL. The fifth is the Google research tree holding PWIL. So five of the eight selectable names are ports or baselines rather than original implementations, and the page credits them by repository rather than by author. That is worth knowing before you read the library as a single author's contribution, and it is also why the variant names matter, since each port carries one particular implementation's choices into the comparison.

## The defaults are the pragmatic part of the title

A library that calls itself a pragmatic look should be judged on its defaults, and they are unusual in a good way. Behaviour-cloning pre-training defaults to zero iterations, so nothing is warmed up unless you ask for it. The absorbing-state indicator defaults to true, while state-only imitation defaults to false, which is the conventional setting for the on-policy half of these methods. Mixing agent and expert data defaults to none, so a run is pure unless you opt into a mix of batched or prefilled variants. The behaviour-cloning auxiliary loss defaults to false everywhere except DRIL, where it is the point. Read together, that is a library designed so that the simplest command line gives you the standard formulation, and everything unusual has to be asked for by name.

## One integer turns AdRIL into a different published algorithm

The AdRIL options are the clearest illustration of how this library is built. There is a balanced sampling switch that alternates expert and agent data batches instead of mixing them in one batch, and there is a discriminator update frequency that takes any non-negative integer. The documentation then says to set that frequency to zero for SQIL, which is a different published method sharing the same discriminator machinery. That is the whole architecture in one sentence: algorithm identity lives in a configuration key, not in a separate implementation. GAIL is the busiest case, with switches for reward shaping, subtracting the log of the policy term, a choice among three reward functions, a gradient penalty, spectral normalisation, an entropy bonus, and a choice of three discriminator loss functions with their own extra knobs. The same pattern holds for PWIL and RED, which differ only in the scale parameters they expose.

## Sweeps and Bayesian optimisation write to two different directory shapes

Results are written under an outputs directory and the naming encodes what produced them, down to a final subfolder holding the current date and time. A plain run lands in a directory named for the algorithm and environment, a parallel run across all environments lands under an all-infix variant, and a sweep with the multi-run flag lands under a suffix with a subdirectory per job number. The Bayesian optimiser is a separate path with its own suffix, and the plotting script is pointed at a concrete directory from one of those runs. Nothing is wrong with any of it, but a reader comparing two experiments has to know which of the two multi-run mechanisms produced a given directory, and the documentation introduces them in different sections with different suffixes rather than comparing them in one place.

## Tuned hyperparameters need two flags, and one of them is easy to forget

The hyperparameters live in a base config file and a per-algorithm file, which is the expected Hydra layout. The tuned sets from the paper are then selectable by name, in the form of an option combining the algorithm with a number of expert trajectories, so a five-trajectory AdRIL configuration has its own name. The documentation adds a warning that the algorithm key must also be specified to load the other algorithm-specific hyperparameters, which is a genuine footgun: forgetting it does not error, it silently loads the defaults and the run looks legitimate. The worked example passes both flags together, along with the environment. The number in the name is the amount of expert data the tuning assumed, so a tuned set applied to a different data budget is a configuration nobody validated.

## torch 1.13.1, gym 0.23.1 and a D4RL commit hash

The dependency file is a snapshot, pinned with equals signs throughout, and it dates the project precisely. The PyTorch pin is 1.13.1, the gym pin is 0.23.1, and D4RL is not pinned to a release at all but to a commit hash in a git URL, which is the only honest way to pin that project and also means the install reaches the network at install time. Around them sit the hyperparameter-optimisation stack at matching old versions, with the Ax platform, BoTorch and GPyTorch, plus a Hydra sweeper plugin, and the evaluation stack with a reliable-benchmarks library and Plotly. The page separately lists PyTorch, Gym, D4RL and Hydra as the notable requirements and adds Ax and its sweeper plugin only for optimisation, which matches the file.

## The Results section is empty and the release links escape the repository

Two documentation gaps are worth knowing before you rely on this for anything. The page has a Results heading with nothing under it, so the numbers live in the paper and in a figures directory rather than in the repository, and the only quantitative handle on the tuned configurations is their existence. And the two release references are written as relative paths that climb above the repository root, so they are only meaningful on the site that published them rather than on a code host. The text attached to them is still the useful part: version 1.0 of the library carried the on-policy algorithms and version 2.0 carried the off-policy ones, which is the single most useful fact about how the project evolved. The licence file is also named with an extension rather than as a plain licence file.

## Conclusion

The design choice here is worth stealing regardless of whether you use it: one training loop, one set of environment and algorithm interfaces, and every published variant expressed as a hyperparameter, which makes side-by-side comparison cheap and removes the usual excuse that two implementations differ in more than the method. Two cautions before you run it. The dependency set is pinned to a torch generation from late 2022 with a D4RL commit rather than a release, so expect a dated environment. And the tuning results are not in the repository, so the tuned hyperparameter sets are the only handle on what the authors considered good, which means you should verify them yourself rather than trusting them by name.

## FAQ

### Which imitation learning algorithms does this repository implement?

Six, all built on soft actor-critic: AdRIL, DRIL in its dropout version, GAIL (also known as DAC or SAM with an off-policy base), GMMIL, PWIL in its nofill version, and RED. Behaviour cloning and plain SAC are selectable through the same command key as baselines, making eight names in total.

### How do I run one of the algorithms?

With `python train.py algorithm=GAIL env=hopper`. The algorithm is one of AdRIL, BC, DRIL, GAIL, GMMIL, PWIL, RED or SAC, and the environment is one of ant, halfcheetah, hopper or walker2d, from the Gym MuJoCo locomotion tasks. Results are written under outputs/ with a datetime subfolder.

### What general configuration options does the library expose?

Behaviour-cloning pre-training with a default of zero iterations, state-only imitation defaulting to false, an absorbing-state indicator defaulting to true, mixing agent and expert data defaulting to none, and a behaviour-cloning auxiliary loss defaulting to false except for DRIL.

### Can a configuration value select a different algorithm?

Yes. The AdRIL discriminator update frequency takes any non-negative integer, and the documentation says to set it to zero for SQIL. GAIL similarly switches between three reward functions and three discriminator loss functions from its own keys.

### How are hyperparameters swept or optimised?

Pass the multi-run flag and a comma-separated list of values, for example `python train.py -m algorithm=PWIL env=walker2d reinforcement.discount=0.97,0.98,0.99`. Bayesian optimisation across all environments runs through train_all.py with the same flag, using Ax and its Hydra sweeper plugin, and results can be plotted with a script pointed at an output path.

### Which dependencies does this imitation learning library pin?

Exact versions throughout, including torch 1.13.1, gym 0.23.1, numpy 1.23.5 and transformers-adjacent tooling from the same era, with D4RL pinned to a specific git commit rather than a release. Ax, BoTorch and GPyTorch are included for hyperparameter optimisation, and a reliable-benchmarks library with Plotly for evaluation.

## Sources

- [Issues](https://github.com/Kaixhin/imitation-learning/issues)
- [Kaixhin/imitation-learning on GitHub](https://github.com/Kaixhin/imitation-learning)
- [License: MIT](https://github.com/Kaixhin/imitation-learning/blob/master/LICENSE)
- [README](https://github.com/Kaixhin/imitation-learning/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kaixhin-imitation-learning
