schedule_free: training runs that do not need a decay schedule
Schedule-Free Optimization in PyTorch
At a glance
- What is it?
- A small PyTorch package that swaps an optimizer's momentum for interpolation plus averaging, giving you an evaluation point that is decoupled from the gradient iterate.
- Who is it for?
- schedulefree is a genuinely different optimizer rather than a scheduler that happens to be off by default, and the distinction matters most when your training run length is uncertain or changes between experiments. The cost is real integration work: paired train and eval modes, a BatchNorm refresh before evaluation, manual cache updates if your framework caches parameters in fp16, and learning rates an order of magnitude above the classical range.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 70 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The idea is to stop evaluating at the point you take gradients
Every standard optimizer couples two things that most training runs want to separate: the point at which gradients are evaluated, and the point at which you report test loss. Schedule-free learning splits them. The gradient descent update is written in the README as three sequences, where `x` is the sequence evaluations should occur at, and `z` and `y` are the primary iterate and the gradient evaluation location:
The three updates are a convex combination of the previous `z` and the current `x`, a plain gradient step on `z` using the loss evaluated at `y`, and an averaging of `z` into `x`. What the README emphasises is that Schedule-Free learning does not require a decreasing learning rate schedule, yet typically out-performs, or at worst matches, SOTA schedules such as cosine-decay and linear decay.
The framing in the README is that Schedule-Free learning replaces the momentum of an underlying optimizer with a combination of interpolation and averaging. So this is not a scheduler wrapper that you bolt onto AdamW. It changes what the optimizer accumulates. A separate line in the README places the method on a spectrum between primal averaging at one end and Polyak-Ruppert averaging at the other.
Six classes, two of them contributed by someone else
The README lists what ships in the package. There are `SGDScheduleFree` and `SGDScheduleFreeReference`, `AdamWScheduleFree` and `AdamWScheduleFreeReference`, plus `RAdamScheduleFree`, which the README marks as a community contribution from nhamanasu. There is also an experimental `ScheduleFreeWrapper` intended to wrap other optimizers.
The `Reference` variants use a simplified implementation but more memory. That is a real trade, not a stylistic detail: the optimised versions compute the third sequence on the fly, which is why the README can claim the method has the same memory requirements as the base optimizer, described as parameter buffer plus momentum. If you are working on memory-constrained hardware, the reference implementation is the safe choice and the optimised one is the space-saving one.
There are also `ScheduleFreeClosure` versions for code that already uses PyTorch optimizer step closures. That matters more than it sounds, because of the next section.
Installation is a single package name:
pip install schedulefreeThe project metadata in `pyproject.toml` names it `schedulefree` at version 1.4.1, licensed as the Apache Software License, with `requires-python` set to `>=3.4` and two dependencies, torch and typing_extensions. An Optax implementation is documented as living in that library for JAX users.
You have to tell the optimizer when you are evaluating
This is the practical cost of the method and the README is direct about it. Because the optimizer uses two different points for gradient calls and test or validation loss, it is necessary to switch the parameter buffer between the two during training. You call `optimizer.train()` where you call `model.train()`, and `optimizer.eval()` where you call `model.eval()`. The optimizer should also be placed in eval mode when storing checkpoints.
If your code supports PyTorch optimizer step closures, the closure forms avoid needing those calls at all. That is the cleaner integration if you have it, and the reason a paired API exists rather than a single class.
The wrapper version handles the same idea for arbitrary optimizers:
base_optimizer = torch.optim.RMSprop(model.parameters(), lr=0.0025)
optimizer = ScheduleFreeWrapper(
base_optimizer, momentum=0.9, weight_decay_at_y=0.1)Note `weight_decay_at_y` in that constructor. The README explains that if you set weight decay on the base optimizer, it computes weight decay at `z`, and the wrapper offers the option of computing it at `y`, which the project says seems to give better results in its experiments. There is also a `ScheduleFreeWrapperReference` that uses more memory but is more numerically stable, recommended for early experimentation or research work.
BatchNorm breaks evaluation unless you refresh it
The first entry in the caveats section is the one most likely to produce a quietly wrong number. If your model uses BatchNorm, additional modifications are required for test and validation evaluations to work correctly. Right before eval, the README suggests a short forward pass in train mode with the optimizer in eval mode:
model.train()
optimizer.eval()
with torch.no_grad():
for batch in itertools.islice(train_loader, 50):
model(batch)
model.eval()The reason given is that BatchNorm keeps a `training_mean` and `training_var` cache which is updated on every forward pass in train mode. Without the refresh, those statistics describe `y` while the weights being evaluated are `x`. Using PreciseBN is named as an alternative that avoids the issue.
A related caveat covers half-precision. Many codebases cache parameters in fp16, and the cached versions have to be updated manually so that evaluation uses the correct `x` sequence rather than `y`. The README notes that some GradScalers do this. Nothing in the package fixes it for you.
Learning rates go up, and tuning does not go away
The caveat list is the most informative part of the README and worth reading in full rather than skimming. For SGD, a learning rate 10x to 50x larger than classical rates is suggested as a good starting point. For AdamW, learning rates in the range 1x to 10x larger than with schedule-based approaches seem to work. A reader moving from a tuned scheduled run will need to retune rather than transplant a value.
`beta` is also flagged as more sensitive than you would expect from standard momentum. The default of `0.9` works on most problems, but the README says it may be necessary to raise it to `0.95` or `0.98`, particularly for very long training runs. That is a direct statement that a single default does not cover the long-horizon case.
Two more lines deserve attention. Using learning rate warmup is recommended, supported through the `warmup_steps` parameter. And the method does require tuning: it will not necessarily out-perform a schedule approach without also tuning regularization and learning rate parameters. That last sentence is a useful corrective to the headline claim of faster training without schedules.
Also worth stating plainly: there is no need to use a learning rate scheduler, but the code is compatible with one. Keeping a scheduler wired up costs nothing and gives you an escape hatch.
Version history explains some API surprises
The releases section of the README documents two changes that will otherwise look like inconsistencies in your code. Version 1.4 added the RAdam implementation credited to nhamanasu. Version 1.3 changed the behaviour of weight decay during learning rate warmup in order to improve stability and to be more consistent with the behaviour of standard AdamW in PyTorch. The previous implementation is still available under the name `AdamWScheduleFreePaper`, which is a helpful naming convention because it tells you exactly which behaviour you are getting.
Those two lines are the entire release history the README carries, and the repository has no tagged releases of its own beyond that account. The current version is in `pyproject.toml`, which reads 1.4.1. The last push was on 2026-07-28.
Examples live in an `examples` directory, and the README names one: image classification on MNIST using convolutional networks, modified from the PyTorch examples repository, with a note that more examples are to be added. A single MNIST script is a thin demonstration for a package whose main appeal is at convergence scale, so the paper is where you will find the evidence. The preprint is `The Road Less Scheduled`, arXiv 2405.15682, authored by Aaron Defazio, Xingyu Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled and Ashok Cutkosky. The README calls the approach faster training without schedules and stresses that you do not need to specify the stopping time or steps in advance.
Editorial conclusion
schedulefree is a genuinely different optimizer rather than a scheduler that happens to be off by default, and the distinction matters most when your training run length is uncertain or changes between experiments. The cost is real integration work: paired train and eval modes, a BatchNorm refresh before evaluation, manual cache updates if your framework caches parameters in fp16, and learning rates an order of magnitude above the classical range. It is a package for a researcher or an infrastructure team willing to own those details, not a drop-in swap for a scheduled run. Start with `AdamWScheduleFree` at the default `0.9`, keep a scheduler available in case you want it, and read the arXiv preprint at 2405.15682 before committing a long run.
Frequently asked questions
What is schedule-free learning?
It is an optimizer modification that replaces the momentum of the underlying optimizer with interpolation and averaging. Instead of one iterate, the method maintains a gradient evaluation point `y`, a primary iterate `z`, and an averaged evaluation point `x`. Test and validation loss should be computed at `x`, which is why the optimizer has separate train and eval modes.
Do I still need a learning rate scheduler?
No, and that is the main point of the package. The README says the code is nevertheless compatible with a scheduler, so you can leave one wired up if you want the option. It also recommends learning rate warmup, which is handled through the `warmup_steps` parameter rather than through a separate scheduler object.
Why does my BatchNorm evaluation look wrong?
Because BatchNorm caches running statistics during forward passes in train mode, and those statistics end up describing the `y` sequence while the weights being evaluated are `x`. The README suggests running a short pass with `model.train()` and `optimizer.eval()` just before eval, or using PreciseBN.
What learning rate should I use with the AdamW variant?
Larger than with schedule-based approaches. The README suggests a range of 1x to 10x the usual rate for AdamW, and 10x to 50x for SGD. It also notes the method still needs tuning of regularization and learning rate parameters, so the higher rate is a starting point rather than a settled value.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-schedule-free)