# Bumblebee: running Hugging Face models in Elixir with Axon and EXLA

> Bumblebee loads Transformers-format checkpoints from Hugging Face Hub and wraps them in Axon models you can run from Elixir, Mix.install or Livebook. The catch is coverage: a model only works if its architecture exists in the library.

**elixir-nx/bumblebee** — Pre-trained Neural Network models in Axon (+ 🤗 Models integration)

- Repository: https://github.com/elixir-nx/bumblebee
- Stars: 1,672 · Forks: 149
- Language: Elixir
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/elixir-nx-bumblebee

## What Bumblebee is for, and who it is aimed at

Bumblebee provides pre-trained neural network models on top of Axon, with integration with Hugging Face Models. The practical promise is that an Elixir developer can download a checkpoint and run a machine learning task in a few lines of code, without leaving the BEAM. The README frames the intended audience directly: the best way to get started is with Livebook, where Smart Cells perform different neural network tasks with few clicks, after which you tweak the code and deploy it. There are also single-file examples of running neural networks inside Phoenix and LiveView apps under examples/phoenix. So the target reader is someone building an Elixir or Phoenix product who wants inference in-process, and someone exploring models in a notebook who wants the exploration to turn into deployable Elixir rather than a throwaway Python script. If you are already comfortable serving PyTorch behind an HTTP boundary, Bumblebee is not solving a problem you have.

## How model loading actually works

The mechanism is worth understanding before you pick a checkpoint, because it explains most of the failure modes. A Transformers-format repository does not store an actual model, only trained parameters and a configuration file. The README describes config.json as a blueprint for how the model should be constructed: it names the model type and model-specific options such as the number of layers and their size. The implementation itself lives in library code, in both Transformers and Bumblebee. When you load a model, Bumblebee fetches the configuration, builds a matching model, then fetches the trained parameters and pairs them with that model. The consequence is stated plainly in the README: in order to use any given model, it needs to have an implementation in Bumblebee. Bumblebee is positioned as an Elixir counterpart of Transformers, so it imports those repositories only as far as its own architecture coverage goes. Parameter files come in several formats, and Bumblebee supports pytorch_model.bin and model.safetensors, while flax_model.msgpack and tf_model.h5 are not supported. Some repositories bundle several models in subdirectories, in which case Bumblebee.load_model({:hf, "model-repo", subdir: "..."}) selects the right one.

## Installing Bumblebee and running a first fill-mask task

Add Bumblebee and EXLA to your dependencies. EXLA is optional in the sense that the library installs without it, but the README calls it an important one because it compiles models just-in-time and runs them on CPU or GPU.

```elixir
def deps do
  [
    {:bumblebee, "~> 0.6.0"},
    {:exla, ">= 0.0.0"}
  ]
end
```

Then set Nx's default backend to EXLA in config/config.exs. Without this the models still work but you lose the compiled execution path.

```elixir
import Config

config :nx, default_backend: EXLA.Backend
```

For GPU execution the README points at the XLA usage documentation and says you must set the XLA_TARGET environment variable accordingly. In a notebook or a standalone script, Mix.install/2 both installs and configures everything in one call.

```elixir
Mix.install(
  [
    {:bumblebee, "~> 0.6.0"},
    {:exla, ">= 0.0.0"}
  ],
  config: [nx: [default_backend: EXLA.Backend]]
)
```

With that in place, the README's end-to-end example loads BERT from the Hub, builds a serving, and runs it. Loading returns a model_info struct and a tokenizer, which you then hand to a task-specific function.

```elixir
{:ok, model_info} = Bumblebee.load_model({:hf, "google-bert/bert-base-uncased"})
{:ok, tokenizer} = Bumblebee.load_tokenizer({:hf, "google-bert/bert-base-uncased"})

serving = Bumblebee.Text.fill_mask(model_info, tokenizer)
Nx.Serving.run(serving, "The capital of [MASK] is Paris.")
```

The result is a map with a predictions list, each entry carrying a score and a token. The README's own output ranks "france" first with a score of about 0.928, followed by "brittany", "algeria", "department" and "reunion". The pipeline shape matters more than the numbers: load, build a serving, run it. That serving is the unit you would embed in a Phoenix endpoint or a LiveView process.

## Checking whether your model is supported at all

This is the first thing to do, and the README offers three routes. Call Bumblebee.load_model({:hf, "model-repo"}) and see what happens. Use the repository inspector tool linked from the README, which runs a number of checks against the repository. Or read the code: open config.json in the model repository, copy the class name under "architectures", and search the Bumblebee codebase for that keyword. A match indicates the model is supported. Tokenizers have their own split. Transformers distinguishes slow tokenizers, implemented in Python, from fast ones, and the README's tokenizer section is where that distinction is drawn; a repository that only ships the Python implementation is a problem for an Elixir runtime. The README's model-support section does not promise universal coverage, and the phrasing is honest about it: Bumblebee imports Transformers repositories as long as they are implemented in Bumblebee. Treat architecture coverage, not download speed or API design, as the constraint that decides whether this library is usable for your project.

## Where Bumblebee is the wrong tool

Two hard boundaries show up in the README itself. The first is format: Flax and TensorFlow parameter files are listed as not supported, so a repository that only publishes flax_model.msgpack or tf_model.h5 cannot be loaded regardless of whether the architecture exists. The second is architecture coverage, and it is the more common blocker, because the Hub has far more model types than any single library implements. There is also a subtler mismatch for teams whose stack is Python-first. Bumblebee is an Elixir counterpart of Transformers, not a wrapper around it, so anything you gain in deployment simplicity you pay for in ecosystem reach: a new architecture appears in Transformers first, and in Bumblebee only once someone implements it. If your roadmap depends on recently published model types, that lag is a real cost. And if your team has no Elixir, adopting Bumblebee means adopting the language, the Nx stack and EXLA together, which is a much larger commitment than picking an inference runtime.

## How it compares with calling Transformers from Python

The obvious alternative is running Hugging Face Transformers in Python and calling it from your Elixir application over HTTP or a port. The difference in approach is structural rather than cosmetic. With Transformers you get the reference implementation, immediate access to new architectures, and the slow-tokenizer fallbacks that Bumblebee does not have. What you add is a second runtime, a process boundary, and serialization between them. Bumblebee collapses that: the model is built in your application's memory, the serving is a value you can run with Nx.Serving.run/2, and the deployment artifact stays a single BEAM release. For a Phoenix app already using Nx, that is a smaller system. For a team whose model is not implemented in Bumblebee, it is not an option at all, and the Python route remains the only path until coverage catches up.

## Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-08-19. Recent releases are v0.7.1 on 2026-07-22, v0.7.0 on 2026-05-15, and v0.6.2 on 2025-06-17. Note the gap between v0.6.2 and v0.7.0: roughly eleven months separate those two tags, which is worth weighing if you need a fix on a short timeline. The README's installation snippets still pin {:bumblebee, "~> 0.6.0"}, so the documented version lags the released one; check Hex for the current version rather than copying the snippet verbatim. Bumblebee is Apache-2.0, which is permissive and generally straightforward for commercial use, but the models you download carry their own licences on the Hub, and those are separate from the library's. Nothing in this repository grants you rights to a checkpoint. Verify the model card's licence before shipping anything to production.

## Conclusion

Bumblebee fits Elixir teams that already run Nx and want text, vision or audio inference inside a Phoenix app or a Livebook notebook without standing up a Python service. It is the wrong pick when your model's architecture is not implemented in the library, or when you need Flax or TensorFlow weights, since only PyTorch and Safetensors parameter files are supported. Before adopting it for a specific model, verify support first: open the repository's config.json, copy the class name under "architectures", search the Bumblebee codebase for that keyword, and confirm the tokenizer files are the fast kind rather than a Python-only slow tokenizer.

## FAQ

### How do I use Bumblebee in Elixir?

Add bumblebee and exla to your dependencies, set config :nx, default_backend: EXLA.Backend, then load a model and tokenizer with Bumblebee.load_model/1 and Bumblebee.load_tokenizer/1. You build a task pipeline such as Bumblebee.Text.fill_mask/2 and run it with Nx.Serving.run/2.

### How do I install Bumblebee?

The README shows two routes: add {:bumblebee, "~> 0.6.0"} and {:exla, ">= 0.0.0"} to your mix.exs deps, or use Mix.install/2 in a notebook or script with the same two packages and the Nx default backend set to EXLA.Backend in the config.

### Which model files does Bumblebee support in a Hugging Face repository?

Bumblebee reads config.json for the model blueprint and supports pytorch_model.bin and model.safetensors for parameters. The README lists flax_model.msgpack and tf_model.h5 as not supported.

### How can I tell whether a Hugging Face model works with Bumblebee?

Open the repository's config.json, copy the class name under "architectures", and search the Bumblebee codebase for that keyword; a match means the model is supported. The README also links a repository inspector tool that runs checks against the repository.

### Does Bumblebee need a GPU?

No, but EXLA is described in the README as an optional yet important dependency because it compiles models just-in-time and runs them on CPU or GPU. To use GPUs you must set the XLA_TARGET environment variable as described in the XLA usage documentation.

## Sources

- [elixir-nx/bumblebee on GitHub](https://github.com/elixir-nx/bumblebee)
- [Issues](https://github.com/elixir-nx/bumblebee/issues)
- [License: Apache-2.0](https://github.com/elixir-nx/bumblebee/blob/main/LICENSE)
- [README](https://github.com/elixir-nx/bumblebee/blob/main/README.md)
- [Releases](https://github.com/elixir-nx/bumblebee/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elixir-nx-bumblebee
