# HELM entered maintenance mode in June 2026, and its package name is not its binary name

> Stanford's evaluation framework for foundation models, spanning language, vision-language, text-to-image, audio, table reasoning and medical evaluation across ten papers and a web interface for inspecting individual prompts. The page announces maintenance mode, releases stopped two months before that, and the three console commands it installs share their prefix with an unrelated tool in a different ecosystem.

**stanford-crfm/helm** — Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.

- Repository: https://github.com/stanford-crfm/helm
- Website: https://crfm.stanford.edu/helm
- Stars: 2,933 · Forks: 421
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stanford-crfm-helm

## HELM entered maintenance mode on June 1, 2026

The status of the project is the first substantive thing on its page, and it is not ambiguous.

An italic note, sitting directly under the badges and above the description, states that HELM entered maintenance mode on June 1, 2026, and links to a maintenance mode policy on the documentation site. The framework's own maintainers have therefore settled the question of whether it is still being extended: it is not.

The dates around it give the shape of the freeze. The three listed releases stop at version 0.5.16, published 2026-04-30, roughly a month before the announcement. The last push to the default branch is 2026-09-01, so there have been commits since maintenance mode was declared but no release since April.

The file itself opens oddly for a project of this standing. Its very first line is a developer's note to self, written as a markdown comment with an empty target, about the fact that an image tag requires its source to be a URL when you want to specify a size. So the page begins with an internal note and then, immediately, a status announcement.

What this means in practice is a split between the record and the tooling. The leaderboards, the ten papers and their cross-links describe work that will keep, and the web interface for inspecting individual prompts and responses will keep running. What will not arrive is support for the next generation of models.

## The package is crfm-helm and the commands are helm-something

One tool, three different names, and a collision risk that follows from the third one.

The distribution name is hyphenated and prefixed. The PyPI badge points at a package called crfm-helm, and the install command installs crfm-helm:

```sh
pip install crfm-helm
```

The console commands are not prefixed at all. Running a benchmark, summarising the results and starting the web server are three separate entry points named for the tool with no vendor prefix.

So the repository name, the index name and the executable prefix are three different strings for the same program. The project is helm on the forge, crfm-helm on the package index, and helm-something on your path.

That last form is worth thinking about before you install. The Kubernetes package manager has been called helm for years, and it is a common thing to have on a machine. Four more commands in that namespace, from an unrelated project, is not a collision in the strict sense since the names differ, but it is the same prefix in the same search path, and anyone reading a shell history or a Dockerfile will have to work out which project a given command came from.

The divergence is most likely why the distribution needed renaming at all, since helm was already taken as an index name.

## The quick start is machine-managed, and its example model is from 2019

The quick start sits between a matching pair of HTML comment markers, one to open the block and one to close it.

That is the signature of a section that tooling generates or verifies, which means the commands on the front page cannot silently drift away from the documentation. That is a better arrangement than most projects manage for their own front page, and it is visible in the raw file rather than merely asserted.

The example run is a single line combining four things: a selector pairing a benchmark subject with a model, a suite name to key the run under, and a cap on evaluation instances.

```sh
# Run benchmark
helm-run --run-entries mmlu:subject=philosophy,model=openai/gpt2 --suite my-suite --max-eval-instances 10

# Summarize benchmark results
helm-summarize --suite my-suite

# Start a web server to display benchmark results
helm-server --suite my-suite
```

Then two more commands sharing that same suite name, one to summarise and one to serve. So the shape of a session is run, then summarise, then inspect, with the suite name as the handle all three share. The web interface then appears on a local port.

Two details are worth noting. The instance cap in the example is ten, which confirms the intent: this is a smoke test to prove the pipeline runs, not a measurement. And the model in the example is a second-generation general text model released in 2019.

That is the right pedagogical choice for a quick start, and it is also a choice with a consequence. A reader who copies the line learns the selector syntax and the suite workflow without learning anything about realistic scale, which is the part that is hard.

## Three flagship leaderboards, and rather more than three linked

The leaderboard section names three as the current flagship boards: capabilities, safety, and the holistic evaluation of vision-language models.

The papers section then links a leaderboard for nearly every entry it lists. There is one for the original language evaluation, one for vision-language, one for text-to-image, one for table reasoning and robustness, one for medical tasks, one for audio-language, and one for each of the two named flagship boards.

So the number of leaderboards a reader can reach from this page is roughly double the number declared as flagship, and the word doing that work is flagship.

The page also says leaderboards are maintained for a range of domains, naming medicine and finance, and for a range of aspects, naming multi-linguality, world knowledge and regulation compliance. It does not enumerate them. That list is on the website instead, so the page gives categories where it could have given addresses.

The papers section is the better index of what the framework actually covers, and it is worth reading as a map. Ten entries, each cross-linked to its paper and, where one exists, to its leaderboard and its documentation. They span language, vision-language, text-to-image, audio-language, table reasoning, medical and enterprise evaluation, with the medical one published in a Nature journal rather than a preprint server.

That range is the concrete form the projects name takes here. It is not one benchmark applied widely, it is a method applied across modalities and domains, and the page's claim that the framework can reproduce the published results from those papers is the claim that makes the leaderboards comparable.

## The base dependency list pulls in a deep learning framework

The dependency list in the project file is commented into groups, and one of the group headers reads Models and Metrics Extras.

Underneath that header sit a transformer library with a wide version range, a tensor library with a range, and a vision library. All three are in the unconditional dependency list, not in an optional group.

So a plain install of the package downloads a deep learning framework whether or not the person installing intends to run a model locally. The comment and the placement disagree, and the placement is what an installer experiences.

That is a defensible choice if the framework's default path involves local models, which for an evaluation tool with a huggingface client in its dependency comments it partly is. It is an expensive one if you only want the metric implementations, and the two extras installer scripts sitting in the repository root suggest there is a split available somewhere, which makes the unconditional list harder to justify rather than easier.

Three constraint styles appear in one list, which is worth noting if you ever need to relax it. Most packages use a compatible-release operator, the numeric stack uses open lower and upper bounds, and exactly one entry carries a hard exclusion. The compatible-release style also means some libraries are confined to a single major line, which for a serialization library pinned to a version 22 series is tight by modern standards.

## Two dependency decisions are workarounds for someone else's bug

Three entries in the dependency list carry a comment explaining a decision rather than a preference, and all three are about upstream problems.

The first pins a transitive dependency. The library it constrains is not one this project depends on directly, it arrives through the datasets library, and the comment says the pin is a workaround for a tracked issue in this repository. Pinning a transitive dependency is a legitimate technique and the comment is honest about what it is doing.

The second installs a package whose entire purpose is a hotfix for a named 2023 vulnerability in that same constrained library. So the answer to that advisory was to add a third-party patch package rather than wait for an upstream release. That is a defensible response to a security fix with a slow release cadence, and it is also a permanent reminder in the manifest that the underlying library has an outstanding issue.

The third is the only hard exclusion in the list. A natural-language toolkit is held to a compatible-release range with one exact version removed, and the comment links the upstream issue that makes that version unusable.

Three dependency lines decided by other projects' bugs, with the reasoning recorded next to each. That is more care than most manifests show, and it doubles as a map of the project's constraints: anything upstream that moves will show up here first.

Two more entries are worth knowing about in an installation context. A microframework for serving web requests is pinned to a single patch-minor line, and a compression library and a sqlite-backed dictionary sit in the same group, which together suggest the local result cache is a compressed structured store rather than loose files.

## The lockfile and the build backend come from different tools

The repository carries a lockfile from one modern tool while the build system section names a different one entirely.

The lockfile is the one produced by a fast modern resolver. The build requirement is an unpinned single entry naming the traditional setuptools backend. So dependency resolution happens in one tool and packaging happens in another.

The combination works and is increasingly common, but it has a consequence worth stating: the lockfile is not an input to the build. Whatever resolves at install time is what gets installed, and the lock constrains a step that the build does not perform. The build requirement being unpinned compounds it, since a lock covering that step would have pinned it.

The rest of the tooling is more conventional and more interesting for what it reveals. Two shell scripts at the root install optional extras, and one of them is named for a shell package manager rather than for anything in this project. That is the sort of filename worth opening before running it, because it implies the framework's dependencies can be installed into an environment this repository does not otherwise have anything to do with.

The documentation build is configured through a mkdocs file and a readthedocs configuration. The web interface has its own directory and its own pre-commit script, separate from the Python one, so the frontend is linted on its own terms.

Finally, the citation the page asks you to use is cut off partway through the author list and never closes. The same citation exists as a file in the repository root, so the file is the copy to use and the inline block is a stale duplicate.

## Conclusion

HELM suits a group that wants its evaluation methodology traceable to a published paper and its leaderboard numbers reproducible from the same code, and that is willing to run a framework whose feature set has stopped growing. Three things to weigh before adopting it. New model support will not come from here, so anything requiring current providers needs a maintained fork or a different tool. The distribution is crfm-helm while the binaries are helm-prefixed, which is worth checking on a machine that also carries the unrelated tool of that name. And the base install pulls a deep learning framework whether or not you plan to run a model locally, so budget for that before installing into a constrained environment.

## FAQ

### What is HELM in the context of LLMs?

HELM stands for `Holistic Evaluation of Language Models`. It is an open source Python framework from Stanford's Center for Research on Foundation Models, for reproducible and transparent evaluation of foundation models including large language models and multimodal models.

### What is HELM Safety?

One of three current flagship leaderboards maintained with the framework, alongside HELM Capabilities and the vision-language leaderboard. Results come from evaluating recent models on notable benchmarks using this framework.

### Is HELM still being developed?

The page states that HELM entered maintenance mode on June 1, 2026 and links to a maintenance mode policy. The newest listed release is v0.5.16 from 2026-04-30, and the last push to the main branch was 2026-09-01.

### How do I install HELM?

With `pip install crfm-helm`. Note the split naming: the distribution is crfm-helm, while the three console commands it provides are helm-run, helm-summarize and helm-server.

### Which benchmarks does HELM support?

The README names MMLU-Pro, GPQA, IFEval and WildBench among datasets and benchmarks available in a standardized format. Its papers section covers language, vision-language, text-to-image, audio-language, table reasoning, medical and enterprise evaluation.

## Sources

- [License: Apache-2.0](https://github.com/stanford-crfm/helm/blob/main/LICENSE)
- [Project website](https://crfm.stanford.edu/helm)
- [README](https://github.com/stanford-crfm/helm/blob/main/README.md)
- [Releases](https://github.com/stanford-crfm/helm/releases)
- [stanford-crfm/helm on GitHub](https://github.com/stanford-crfm/helm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stanford-crfm-helm
