Model or dataset
Tongjilibo/bert4torch avatar
Tongjilibo/bert4torch

bert4torch reimplements the transformer zoo, and its own table says where transformers wins

An elegent pytorch implement of transformers

1,325 stars166 forksPythonMIT

At a glance

What is it?
bert4torch is a PyTorch reimplementation of the transformer model family with a Keras-style training loop, a directory of ready-made training tricks, and a serve command that turns one model directory into a terminal chat, a Gradio page or an OpenAI-compatible endpoint. The most interesting document in the repository is not the feature list but the comparison table, whose last row is a concession.
Who is it for?
bert4torch fits someone who wants the transformer model family in PyTorch, with fine-tuning, distributed training and callbacks in a Keras-style loop, and who wants to serve a local model without writing a script per model.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 141 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The comparison table's last row is the honest part

The feature comparison against the reference implementation is nine rows long, and the first seven are close to even.

Progress bars, distributed data-parallel training, callbacks for logging and early stopping, streaming and batch output for large-model inference, and large-model fine-tuning all get a check on both sides. The notes explain the overlaps: the progress bar prints loss and your own metrics, distributed training is what PyTorch ships, and the model scripts are meant to be generic so you do not maintain one script per model.

Then two rows favour bert4torch. Ready-made training tricks, where adversarial training and similar additions are described as plug-and-play, and readable, reusable code with room to customise, with the training style attributed to Keras.

The ninth row goes the other way. Repository maintenance capability, influence, usage and compatibility all get a check for the reference implementation and a cross for this one, with the note that the repository is currently personally maintained.

A project that marks its own weakest column is rarer than it should be, and it tells you where to look before adopting it: not at the feature list, which is competitive, but at who answers an issue in two years.

The version number is regex-scraped out of the CLI module

The packaging script does not declare a version. It opens the command line module, searches for an assignment whose name is the version constant, and returns whatever string follows the quote.

So the single source of truth for the version is a line inside the CLI implementation, and the build reads it with a regular expression rather than importing it. That is unusual, and it has a practical consequence: renaming that constant or reformatting that line silently breaks packaging rather than failing loudly, because a version has to be found for the build to proceed at all.

The same script declares two console entry points, a full name and a two-letter alias, both pointing at the same function. Registering an alias alongside the full name is a small kindness for muscle memory and costs nothing.

Dependencies are read from a text file line by line, with only lines beginning with a hash treated as comments. That file is five lines, and it includes a Python 2 and 3 compatibility shim, which is a leftover given that the same file asks for a modern PyTorch. It also names the sibling training library at an exact version, so the two projects are locked together rather than ranged.

The declared torch floor is far below what it develops against

The dependency on PyTorch is declared as a bare lower bound, strictly greater than one point six. The page, separately, says development originally used one specific older release and has since switched to version two point zero, and invites feedback if other versions misbehave.

So the floor and the tested target are three major generations apart, with a sentence acknowledging the gap and asking for reports. That is a reasonable position for a library that must work across whatever its users have pinned, and an uncomfortable one if you are trying to work out which torch a given release was validated against.

The extras are where the pinning actually lives. There are five of them: two libraries for transformer models and distributed loading, a distributed training library, a reinforcement learning library, and a parameter-efficient fine-tuning library. Four of those five float, and the distributed training one is capped with both a lower and an upper bound.

That single capped dependency is the most informative line in the file, because it marks the one integration the author expected to break on a new major release.

The package index lags the branch, and the project says so

There are two install commands and they are not equivalent. One installs the stable release from the package index. The other installs directly from the repository, which is described as the latest version.

The page then adds a caution in the same list: the published package is slower than the development version on the git branch. So the documented default is not the code the project is working on, and anyone who reads that sentence will clone instead.

Two further cautions follow, and both are the kind that cost an afternoon. When you clone, watch the reference paths, and check whether the pretrained weights need converting.

The weights question is not incidental. The model loader takes either a hub name or a local directory, and in both cases it looks for the project's own configuration file alongside the checkpoint. Give it a folder containing only the original weights and it will not find what it expects.

That is why the sibling training library is pinned to an exact version rather than a range: the two are released together, and the release table lists them side by side for every version.

One serve command produces three different interfaces

The feature that distinguishes this project is a single command that deploys a model, with the interface chosen by a flag.

The command takes either a hub identifier, which downloads everything, or a local directory, in which case the configuration file must already be sitting next to the model. Then a mode flag selects the surface: a terminal chat, a Gradio page in a browser, or an OpenAI-compatible endpoint.

That third mode is the one that changes what you can do. Once the server speaks the OpenAI format, any client that speaks it can talk to a model on your own disk, without writing a serving script.

The comparison table makes the same point from the other direction, claiming that the model scripts are generic so no separate script has to be maintained per model. Which is the real claim underneath all three modes: the interface varies, the code behind it does not.

The serving example directory exists alongside the model examples, and the deployment commands in the page are given for the local directory case, with the Gradio and OpenAI variants differing only in the flag.

Version 0.6.2 dropped the transformers dependency and kept the extra

The release table pairs this project with its sibling and gives one line of notes per version, which is the most efficient changelog format available.

The most consequential entry is the newest one. It adds vision-language and optical character recognition models, and it removes the dependency on the reference implementation, adding its own tokenizer and processor classes in its place. Version 0.6.1, earlier in the year, added a paddle-based OCR model and removed hardcoded model configuration. Version 0.6.0, in late 2025, added a mixture-of-experts model and support for two common quantization formats. The version before that added a model family, fixed a download bug, and split the OpenAI client into its own module.

The split-out client is what makes the serve command's OpenAI mode clean, and it is the kind of refactor that shows up in a changelog as one line.

The extras list still offers an install of the reference implementation, and the documentation keeps a tutorial for loading a model through it. So removing the dependency and supporting it are both true, which is a reasonable outcome: the core stopped needing it, and the bridge stayed for people who still want it.

Weights need a second config file, and two directory names are misspelled

The model loader has three input forms. Give it a hub name and it fetches both the weights and the project's configuration file. Give it a local directory and it looks for checkpoint files and that same configuration file. Or pass the configuration path and the checkpoint path separately, which is described as equivalent to the first two.

So every model you point this at needs two things in the right place, and the second thing is not part of the original checkpoint. That convention is the project's central interface with the weights ecosystem, and it is why the page warns about conversion.

Smaller details are equally worth noting. Two of the example directories are misspelled in the same way, a classification example whose name drops a letter, and the misspelling appears in both the page and the repository, so a link to it works while a search for the correct spelling does not.

The documentation is split by language: a readthedocs site in English, tutorial pages written in Chinese, and three longer articles on a Chinese blogging platform. The repository also carries a group-chat image for its community, and the page links download statistics for the package, which is a small signal about how much attention the library has had.

Editorial conclusion

bert4torch fits someone who wants the transformer model family in PyTorch, with fine-tuning, distributed training and callbacks in a Keras-style loop, and who wants to serve a local model without writing a script per model. Verify five things first: which version you are installing, because the package index lags the git branch by the project's own admission; that your weights come with the config file the loader expects, since a plain folder is not enough; which torch version you are on, since the declared floor is far below the one the project develops against; whether you need the extras, particularly the parameter-efficient fine-tuning library for LoRA; and whether you are comfortable with a personally maintained repository, which the project's own comparison table says is the weaker side of the trade. MIT licensed, version 0.6.2, with the last push dated 2026-05-16.

Frequently asked questions

What is bert4torch?

It is an MIT licensed PyTorch reimplementation of the transformer model family, MIT licensed, covering BERT, RoBERTa, ALBERT, XLNet, BART, RoFormer, ELECTRA, GPT, GPT2, T5, GAU-alpha and ERNIE for fine-tuning, plus chat models such as ChatGLM, Llama, Baichuan and Bloom for inference and fine-tuning, with a Keras-style training loop.

How do I install bert4torch?

`pip install bert4torch` gets the stable release, and `pip install git+https://github.com/Tongjilibo/bert4torch` gets the latest development version. The page warns that the published package is slower than the git version, and that when you clone you should watch the reference paths and check whether the pretrained weights need converting.

How do I serve a model with bert4torch?

`bert4torch serve <model>` takes either a hub name or a local directory; a local one needs the project's configuration file already beside the model. A mode flag then chooses the surface: `--mode cli` for a terminal chat, `--mode gradio` for a browser page, or `--mode openai` for an OpenAI-compatible endpoint.

Does bert4torch still depend on transformers?

Version 0.6.2 removed the dependency from the core and added its own tokenizer and processor classes. An extra is still offered for installing the reference implementation, and the documentation keeps a tutorial for loading a model through it, so both statements are true at once.

What extras does bert4torch offer?

Five: an install of the reference implementation, a distributed loading library, a distributed training library capped at both ends of a version range, a reinforcement learning library, and a parameter-efficient fine-tuning library. The page notes that LoRA support depends on the last of those.

How does bert4torch compare to transformers?

Its own table marks both as equal on progress bars, distributed training, callbacks, streaming and batch inference, and fine-tuning. It claims the advantages on ready-made training tricks and on readable, reusable code, and concedes the point on repository maintenance, influence, usage and compatibility, noting the repository is currently personally maintained.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. Tongjilibo/bert4torch on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tongjilibo-bert4torch.svg)](https://hysenlabs.com/projects/tongjilibo-bert4torch)