Model or dataset
FluxML/Metalhead.jl avatar
FluxML/Metalhead.jl

Metalhead.jl: Flux vision architectures, and where the pre-trained weights are

Computer vision models for Flux

348 stars68 forksJuliaNOASSERTION

At a glance

What is it?
Metalhead.jl collects standard computer vision architectures as pure Flux layers. Only a few of them ship pre-trained weights, and the README is explicit about which ones.
Who is it for?
Adopt Metalhead.jl if you are already working in Julia and Flux and want a ResNet, VGG or ViT defined in the same layer types you train with, rather than a wrapper around a foreign graph format. Do not adopt it if you need a pre-trained checkpoint for AlexNet, DenseNet, EfficientNet, MobileNet, ConvNeXt, the Inception family or UNet: the README's pre-trained column marks those as N, and no download path for them is documented.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 90 days ago.
What is it written in?
Mainly Julia, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Metalhead.jl is for, and who it is not for

Metalhead.jl is a model zoo for Flux.jl, the Julia machine learning library. Its stated purpose is to provide standard vision models that are built from pure Flux layers, and the README frames the implementations as reference material: they "represent the best-practices for creating modules like residual blocks, inception blocks, etc. in Flux." That second sentence matters more than the first. A reader who wants a frozen checkpoint to run inference on an image is using the package differently from a reader who wants to read how a residual block is expressed in Flux and copy the pattern.

The package also exposes a Layers module for building blocks used by the more complex architectures. So there are two audiences: people who want a named architecture they can instantiate and train, and people writing their own vision model who want the pieces. If you fall outside both, for instance if you want a Python-compatible inference server, this is the wrong dependency.

Which architectures actually ship pre-trained weights

The README's model table has a Pre-trained column, and it is the single most useful thing on the page. Most entries are marked N. The entries marked Y are ResNet, ResNeXt, SqueezeNet, WideResNet, VGG and ViT. Everything else, including AlexNet, ConvMixer, ConvNeXt, DenseNet, EfficientNet, EfficientNetv2, gMLP, GoogLeNet, all three Inception variants, MLPMixer, the MobileNet family, MNASNet, ResMLP and Xception, is marked N.

That asymmetry is the practical constraint on the whole package. A constructor marked N still gives you a working Flux model with the right layer topology, but you supply the training. If your plan was to load ImageNet weights for MobileNetv3 and fine-tune, the table says no before you write any code. UNet is listed separately under Other Models, also marked N, which makes sense: it is a segmentation architecture and the table does not claim a pre-trained checkpoint for it.

The spread of architectures is wide relative to the size of the project, and the README does not explain why pre-trained weights exist for some and not others. It also does not document where the weights come from, what dataset they were trained on, or what accuracy they reach. Treat the Y as "a checkpoint is available" and nothing more until you check the linked API page for the constructor you want.

Installing Metalhead.jl and running a first forward pass

Installation goes through the Julia package manager, as the README shows. From the Julia REPL, enter the package mode with the closing bracket and add the package:

julia
julia> ]add Metalhead

After that, the package is available to any environment you activate. The README points to a getting started guide under the project's documentation site at fluxml.ai/Metalhead.jl/dev/tutorials/quickstart/ for the first real use, and that is where the concrete walkthrough lives rather than in the README itself.

Because the package is built on Flux, a first use means bringing Flux into scope alongside Metalhead and constructing a model before feeding it a tensor. The README does not print a full inference example, so the exact call signature for each constructor is documented on its API page, linked from the model table. The pattern to expect is that you call the constructor by name, for example `ResNet`, and get back a Flux model you can call like any other layer chain. What you should see is a model object printed as a nested Flux layer structure. If you are after pre-trained weights, check the constructor's API page for the loading argument rather than assuming the default constructor downloads anything; the README's table only tells you a checkpoint exists, not how it is fetched.

The pre-trained gap is the real limitation

The most common reason a project like this disappoints is a mismatch between the architecture someone wants and the weights that are actually available. Here the mismatch is documented, but easy to miss if you skim the table and stop at the model names. Roughly two dozen architectures are listed and only a handful carry pre-trained weights.

There is a second, quieter limitation. The README does not document how weights are stored or fetched. The repository has an Artifacts.toml at the top level, which is the Julia mechanism for pinning and downloading binary artifacts, and a scripts/ directory, but the README does not describe what either contains or how a user triggers a download. If you are working in an environment without network access, or you need to know the size and provenance of a checkpoint before pulling it, the README is silent. That is not a defect in the architectures; it is a gap in the documentation that you should close by reading the API page and the Artifacts.toml before you plan around pre-trained weights.

A third case where this is the wrong tool: if your team is not already on Flux. Metalhead's value is that the models are pure Flux layers, so they compose with Flux training loops, optimisers and GPU handling. Outside that ecosystem the same architectures are available in other frameworks with far more pre-trained checkpoints, and you would be paying the cost of a Julia dependency for no integration benefit.

How it compares to torchvision and timm

The obvious comparison is torchvision, and the difference is not the architecture list, which overlaps heavily. It is the graph. A torchvision model is a torch.nn.Module, saved and loaded as a state dict, and the pre-trained checkpoint is the primary product; the model definition is almost incidental. Metalhead inverts that. The model definition is the product, written in pure Flux layers so that it is readable and composable inside Julia, and pre-trained weights are a bonus attached to a minority of entries.

timm sits between the two in spirit: it is a large collection of vision models with a strong emphasis on pre-trained weights and fine-tuning recipes. If your selection criterion is "which ImageNet checkpoint do I fine-tune," timm has far more to choose from and the README here does not claim otherwise. If your criterion is "I am writing Julia, I want to train a ResNet from scratch, and I want the model to be ordinary Flux code I can read and modify," Metalhead is aimed exactly at you. The two are not competing on the same axis, and picking based on the model-name list alone will mislead you.

Maintenance, releases and the licence question

The repository is not archived, and the last push was on 2026-07-01. The most recent tagged release is v0.9.6 from 2026-01-05, preceded by v0.9.5 on 2025-01-08 and v0.9.4 on 2024-09-10. The gap between v0.9.5 and v0.9.6 is roughly a year, while the gap between v0.9.4 and v0.9.5 is about sixteen months, so the release cadence has been slow and irregular rather than steady. Pushes to the default branch happen more often than releases, which is typical for a package whose API surface changes rarely once a set of architectures is in place.

For an upgrade decision, that cadence cuts both ways. You are unlikely to be forced into frequent migrations, and you are also unlikely to see a new architecture or a new pre-trained checkpoint arrive quickly. Pin a version in your Project.toml rather than tracking master if you depend on specific constructor signatures.

The licence is the loose end. The repository carries a LICENSE.md file, but the metadata reports the licence as NOASSERTION, meaning no standard licence identifier was recognised. The README does not state licence terms, and it does not state the licence or provenance of any downloaded weights. Since several architectures here are ports of published models, and pre-trained weights carry their own terms separate from the code, confirm both the code licence and the weight terms with whoever owns the deployment before you ship. That is a question for your legal team, not something this article can settle.

Editorial conclusion

Adopt Metalhead.jl if you are already working in Julia and Flux and want a ResNet, VGG or ViT defined in the same layer types you train with, rather than a wrapper around a foreign graph format. Do not adopt it if you need a pre-trained checkpoint for AlexNet, DenseNet, EfficientNet, MobileNet, ConvNeXt, the Inception family or UNet: the README's pre-trained column marks those as N, and no download path for them is documented. Before committing, check the pre-trained column for the exact constructor you plan to use, then confirm the download works on your machine, since the README does not document offline behaviour or caching.

Frequently asked questions

Does Metalhead.jl include pre-trained weights for every model?

No. The README's model table marks ResNet, ResNeXt, SqueezeNet, WideResNet, VGG and ViT as pre-trained. The remaining entries, including AlexNet, DenseNet, EfficientNet, MobileNet and Xception, are marked N.

How do I install Metalhead.jl in Julia?

The README shows entering package mode in the Julia REPL with the closing bracket and running ]add Metalhead. The package is then available in the active environment.

What is Metalhead.jl built on?

It is built on Flux.jl. The README states that the architectures use pure Flux layers and that the package also provides building blocks for more complex models in the Layers module.

Where is the Metalhead.jl getting started guide?

The README links to a getting started guide on the project documentation site at fluxml.ai/Metalhead.jl/dev/tutorials/quickstart/. The README itself does not include a full worked example.

Official sources

  1. FluxML/Metalhead.jl on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fluxml-metalhead-jl.svg)](https://hysenlabs.com/projects/fluxml-metalhead-jl)
Community notes

Community notes