Transformers.jl: Hugging Face BERT, GPT-2 and Llama 2 Weights Inside Julia
Julia Implementation of Transformer models
At a glance
- What is it?
- Transformers.jl is a Julia package that reimplements transformer architectures on top of Flux.jl and loads pretrained checkpoints from Hugging Face. It is a research and porting tool, not a production serving stack, and its documentation is thinner than its example folder.
- Who is it for?
- Adopt Transformers.jl if you already work in Julia and want to run or fine-tune a Hugging Face checkpoint without leaving the language, and start by reproducing one model from the example folder before trusting the HuggingFaceValidation results for the architectures you care about.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 61 days ago.
- What is it written in?
- Mainly Julia, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Transformers.jl Is For, and Who Actually Needs It
The package exists for one narrow situation: you are writing Julia, you want a transformer, and you do not want to shell out to Python. The README describes it as a "Julia implementation of transformer-based models, with Flux.jl", which places it inside the Flux ecosystem rather than beside it. Models are Flux layers, gradients flow through Flux's AD, and training loops look like any other Flux training loop.
The second half of the pitch is weight compatibility. The HuggingFace submodule downloads pretrained checkpoints from the Hub, so a Julia user can pull bert-base-uncased or a GPT-2 checkpoint and get a model whose parameters match the original. That is the real reason to pick this over writing attention from scratch: you inherit years of pretraining instead of a randomly initialized stack.
Who is it not for? Anyone whose team is already standardized on PyTorch or TensorFlow. The value here is language integration, and if you do not have a Julia codebase to integrate with, you are paying a porting cost for nothing. The example folder hints at the intended audience: AttentionIsAllYouNeed, BERT, GPT2_TextGeneration, T5, Llama2_example.ipynb, Dolly_example.ipynb, and a HuggingFaceValidation directory. Those are reproduction and porting exercises, which is what this package is best at.
How the TextEncoder and Model Pipeline Fit Together
There are two objects in the documented flow, and they are separate on purpose. A TextEncoder handles tokenization and preprocessing: adding special tokens, truncating or padding to a fixed length, and one-hot encoding into the integer arrays the model consumes. The model itself is a Flux layer built from the checkpoint's configuration.
The README example makes the split explicit. The call hgf"bert-base-uncased" returns both a textencoder and a bert_model. Encoding a batch of sentences produces a sample whose token field, when passed back through decode, yields the token strings including the [CLS] and [SEP] markers and the WordPiece continuation pieces written as ##led and ##zzy. That round trip is the clearest documented statement of what preprocessing does, and it is worth reading closely because it shows padding and special-token insertion are not optional extras you add later.
The model call returns a named tuple with a hidden_state field. Everything downstream, whether you are doing classification, feature extraction or generation, reads from that. The architecture is not hidden behind a task-specific wrapper, so you get the raw sequence of hidden states and decide what head to attach. That is flexible and also means the package does less for you than a pipeline API would.
Installing Transformers.jl and Running BERT on Two Sentences
Installation is a single package-manager command from the Julia REPL, using the pkg mode bracket. There is no build step, no Python environment and no separate weights download command documented in the README.
]add TransformersAfter that, the README's own example is the shortest real use. It loads a pretrained BERT, encodes two sentences as one batch of contiguous sentences, checks the tokenization round trip, and runs the model.
using Transformers
using Transformers.TextEncoders
using Transformers.HuggingFace
textencoder, bert_model = hgf"bert-base-uncased"
text1 = "Peter Piper picked a peck of pickled peppers"
text2 = "Fuzzy Wuzzy was a bear"
text = [[ text1, text2 ]] # 1 batch of contiguous sentences
sample = encode(textencoder, text) # tokenize + pre-process (add special tokens + truncate / padding + one-hot encode)The assertion in the README is the part to run first, because it verifies tokenization independently of the model. It expects the decoded tokens to come back as [CLS], peter, piper, picked, a, peck, of, pick, ##led, peppers, [SEP], then fuzzy, wu, ##zzy, was, a, bear, [SEP]. Note the ##led and ##zzy pieces: if your decode output does not contain those, the tokenizer is not the one the checkpoint was trained with, and the model output will be meaningless regardless of whether it runs.
@assert reshape(decode(textencoder, sample.token), :) == [
"[CLS]", "peter", "piper", "picked", "a", "peck", "of", "pick", "##led", "peppers", "[SEP]",
"fuzzy", "wu", "##zzy", "was", "a", "bear", "[SEP]"
]
bert_features = bert_model(sample).hidden_stateWhat you should see after the final line is a hidden_state array with one row per token in the batch. The first download of a checkpoint is the slow part; the README does not document a cache location or an offline mode, so if you need reproducible air-gapped runs, that is something to check in the source before committing.
The 0.1.x Break and What It Costs You
The README opens with a notice that the current version is almost completely different from the 0.1.x line, and tells old users to either update their code or pin the old version. That is a candid warning, and it should shape how you read any code you find online. Blog posts, older Julia discourse threads and pasted snippets from before the rewrite will not run against the current API.
The release history supports the same reading. v0.3.1 landed on 2024-06-28, v0.3.0 on 2024-06-09, and v0.2.8 back on 2023-10-01. The repository's last push was on 2026-07-31, which is more recent than the last tagged release, so master carries work that is not in any release. If you depend on this package, depend on a tagged version and read the CHANGELOG before moving, rather than tracking master.
The deeper limitation is documentation coverage. The README sends you to the dev documentation site and to the example folder, and explicitly suggests tagging the maintainer on Julia's Slack or Discourse, or opening an issue, if you have questions. That is a reasonable support model for a research package, but it means the API surface is not fully described in one place. Expect to read src/ for anything the examples do not cover.
When a Julia-Native Model Is the Wrong Choice
The clearest failure mode is serving. Nothing in the README describes batching strategies for concurrent requests, quantization, ONNX export, a server binary or a REST interface. If your goal is to put a model behind an HTTP endpoint for other services, the Python ecosystem around the Hub, which the checkpoints themselves come from, is the shorter path, and Transformers.jl gives you no advantage there.
The second case is architecture coverage. The example folder is the honest inventory of what has been exercised: AttentionIsAllYouNeed, BERT, GPT2_TextGeneration, T5, and two notebooks for Llama 2 and Dolly. A model family with no example directory is unproven, and the HuggingFaceValidation folder exists precisely because matching a reference implementation is a task in itself. Check for your architecture before you plan around this package.
The third case is a team with no Julia experience. The package assumes fluency with Julia's package manager, with Flux layer semantics, and with reading source when the docs run out. A Python-first team will spend more time on the language than on the model.
Transformers.jl Versus Calling Hugging Face from Python
The obvious alternative is to keep the model in Python and call it from Julia, either through a subprocess, a socket, or a binding. The difference is where the boundary sits. With Python you get the full Hub ecosystem: every architecture, the tokenizer implementations that ship with the checkpoints, quantization, and serving libraries. You also get a second runtime, a serialization boundary, and a dependency on a Python environment that has to be reproduced on every machine.
Transformers.jl moves the boundary to the weights. The checkpoint is the interface, and everything after it is Julia. That means no Python at runtime, gradients that flow into the rest of your Flux model, and the ability to fine-tune inside the same program that consumes the result. It also means you reimplement or re-verify anything the package has not ported yet.
If you are already in Julia and your architecture is covered by an example, the native path removes a whole class of deployment problems. If you are not in Julia, the Python path is not a compromise, it is simply the better fit. The choice is about where you want your language boundary, not about which implementation is faster, and the repository does not publish benchmark numbers that would let you decide on speed.
Licence, Maintenance and Upgrade Cost
The package is MIT licensed, which permits commercial and closed-source use and requires preserving the copyright notice and permission text. That is the least restrictive common option and it does not create copyleft obligations on your own code. This is a description of the licence text, not legal advice; if licence compatibility matters to your organisation, have counsel read the LICENSE file in the repository root.
Maintenance has to be judged from the commit history rather than from the release tags. The last push was on 2026-07-31, so the repository is being touched, but the most recent tagged release is v0.3.1 from 2024-06-28. Work is landing on master that has not been cut into a release. For a dependency you plan to pin, that gap is the thing to watch: you may be running code that is two years behind the branch.
Upgrade cost is dominated by the API churn described in the README. The 0.1.x to 0.2.x transition was a rewrite, and the README treats that as a live concern for anyone on old code. Budget for reading the CHANGELOG and the example folder at each version bump, and treat the example folder as the compatibility test suite you actually run.
Editorial conclusion
Adopt Transformers.jl if you already work in Julia and want to run or fine-tune a Hugging Face checkpoint without leaving the language, and start by reproducing one model from the example folder before trusting the HuggingFaceValidation results for the architectures you care about. Do not adopt it if you need a Python-compatible serving path, a stable API across releases, or documentation that covers every exported function, because the README itself warns that the current version is almost completely different from 0.1.x and points readers at the source and the example directory rather than a reference manual. Verify first that the specific architecture you need has an example under example/ and that its weights load through the hgf string macro.
Frequently asked questions
How do I install Transformers.jl?
From the Julia REPL, enter package mode with the closing bracket and run ]add Transformers. The README documents no other setup step, and the package is registered under the name Transformers.
Can Transformers.jl load pretrained Hugging Face models such as bert-base-uncased?
Yes. The README example calls hgf"bert-base-uncased" through the Transformers.HuggingFace submodule, which returns a textencoder and a model, and the model call exposes a hidden_state field.
Is Transformers.jl compatible with the 0.1.x version?
No. The README states that the current version is almost completely different from 0.1.x and advises users of the old version to update their code or stick to the old version.
Which transformer architectures have examples in Transformers.jl?
The example folder contains directories for AttentionIsAllYouNeed, BERT, GPT2_TextGeneration, T5 and HuggingFaceValidation, plus Dolly_example.ipynb and Llama2_example.ipynb notebooks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chengchingwen-transformers-jl)