# BiSeNet: a faithful reproduction that reports its own error bars, and then gets them ignored

> This is a PyTorch implementation of two published segmentation architectures, with six pretrained models and four deployment paths. What makes it worth reading is the fourth note under the results tables, where the author states that repeated training runs of the same configuration vary by about two points of the headline metric, which means most of the differences in his own table are inside the noise.

**CoinCheung/BiSeNet** — Add bisenetv2.  My implementation of BiSeNet

- Repository: https://github.com/CoinCheung/BiSeNet
- Stars: 1,640 · Forks: 336
- Language: Python
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/coincheung-bisenet

## The author's own caveat invalidates the comparison above it

The readme opens with three results tables, one per dataset, each giving the same model under four evaluation protocols, and each with a download link. Inference on one image is a single command:

```
python tools/demo.py --config configs/bisenetv2_city.py --weight-path /path/to/your/weights.pth --img-path ./example.png
``` Then come five notes, and the fifth one is the most important sentence in the document.

The model has a large variance. The stated consequence is that training it many times gives results varying within a relatively wide margin, and the example given is specific: training the second architecture on the driving dataset repeatedly produces a single-scale score anywhere between about seventy-three point one and seventy-five point one.

That is a two-point spread on a metric where the whole table is compressed into a range of about three and a half points. Apply it to the first table. On the single-scale protocol the two architectures differ by about half a point, and the best single-scale result in the second row is inside the first row's range. The multi-scale-with-crop numbers differ by about one and three quarter points, which is at the edge of the stated spread. The point is not that the second architecture is worse. The point is that this table cannot tell you which architecture is better, and the person who produced it is saying so in the readme rather than letting you draw the conclusion from the column ordering.

That is genuinely rare. Most reproductions present a table and let the ordering do the arguing. This one presents a table, then explains why the ordering is not evidence.

The other three notes are of the same character, and each one marks a place where these numbers are not comparable to the published ones. The frames-per-second figures were measured differently from the paper, with a pointer to the deployment directory. The dataset used for the second architecture is not the one the authors used, chosen because it matches the object detection split, so the results may differ from the paper. And the third dataset was never reported by the original authors at all, so there are no official training settings, and what is provided is described in the author's own words as a result that merely makes it work, with the suggestion that better settings would improve it.

That last one is the correct thing to publish. A number you cannot reproduce is worse than no number, and saying so converts a claim into a starting point.

## Six pretrained models on one tag named zero point zero point zero

The release history contains one entry. It is version zero point zero point zero, the release note reads as the original implementation, and it is dated August 2020.

Every pretrained model in the readme hangs off that one tag. Six download links, two architectures times three datasets, all resolving to files attached to a release that is nearly six years old. That is not necessarily wrong. Model weights do not rot the way code does, and pinning weights to a specific tag is arguably the correct discipline, since a tag is immutable while a branch is not.

But the code in this repository is not frozen. The last commit to the default branch is dated April 2026, and the repository has been moved forward in ways that matter: a second architecture was added after the first, which the project description states outright, and the directory listing still contains a directory named for the old state. So the code is six years newer than the tag, and the weights are six years older than the code.

The consequence is a class of problem that is easy to hit and hard to diagnose. If a change to the model definition, the state dictionary layout, the configuration format or the preprocessing pipeline landed after the tag, then a weight file from that tag may not load into current code, or may load and produce different output. Nobody would be at fault. The tag says these weights came from this code, and the code has moved since.

The reverse is worse. If you train with the current code and compare against the published numbers, you are comparing a 2026 training run against a 2020 evaluation, on a machine configuration from 2021, and the readme has already told you the noise floor is two points. Those three facts compound.

For anyone planning to use the weights, the honest advice is to test them on a small set of images you can check by eye before trusting them for anything, and to check the output shape and label mapping rather than assuming they match the current configuration files. For anyone planning to train, the situation is cleaner: the training code is the thing to evaluate, and the weights are a reference rather than a dependency.

## Four deployment paths from one training implementation

The deployment section lists four targets, and the list is longer and more varied than you would expect from a research reproduction.

The first is a GPU inference toolkit from the same vendor as the training hardware, with its own directory in the repository. The second is a mobile-focused inference library, and its presence is the most surprising item on the list, because it implies the author cared about running this on a phone, which for a real-time segmentation model is a plausible and interesting target. The third is a toolkit from a different hardware vendor, implying Intel hardware, with its own directory. The fourth is a general model serving system with its own directory.

Four directories, four toolchains, one training implementation. That is a lot of work that has nothing to do with the model, and it is the part of this repository most likely to have bit-rotted. GPU inference toolkits and mobile inference libraries both move quickly, and a conversion script written for one version of either is a script that stops working. A serving system has a plugin interface that changes between majors.

The four evaluation protocols in the results tables are the same shape of decision, applied to measurement rather than to deployment. A single-scale evaluation is the plain one. A single-scale crop evaluation crops the image before scoring. A multi-scale evaluation runs several input sizes and adds horizontal flipping. A multi-scale crop evaluation does both. And critically, the exact scales and crop size are not in the paper or in the readme, they are in the configuration files, which is the right place for them.

That matters more than it looks. A multi-scale evaluation with the wrong scales is a different number, and papers that report multi-scale results are notoriously hard to compare because the protocol is under-specified. Publishing the protocol as configuration files, in the repository, next to the model, means someone can actually check whether your number is comparable to theirs. It is the cheapest possible reproducibility work and almost nobody does it.

Together, the four protocols and the four deployment paths describe a project whose author was trying to make the work usable rather than merely correct. That is worth recognising, because it is also what created the maintenance surface.

## The speed comes from the integer column, and the comparison between architectures flips there

The frames-per-second figures in the first table are given for three precisions, and they are the most interesting numbers in the readme for a reason that has nothing to do with segmentation accuracy.

For the first architecture, the figures are roughly one hundred and twelve frames per second at full precision, two hundred and thirty-nine at half precision, and four hundred and thirty-five at eight-bit integer. For the second architecture, roughly one hundred and three, one hundred and sixty-one, and one hundred and ninety-eight.

Two things follow. The first is that the precision columns differ far more than the architecture rows do. Going from full precision to integer more than triples the throughput of the first model and nearly doubles the second. That is the single largest performance lever in the table, and it is a deployment decision rather than an architecture decision, which means it is available to you whatever model you pick.

The second is that the ordering of the two architectures depends on which column you read. At full precision the first architecture is faster by about nine percent. At half precision it is faster by nearly fifty percent. At integer precision it is more than twice as fast. So a claim about which of these two models is faster is not a claim at all until you say at what precision, and any summary that picks one column is making a choice silently.

The caveat that sits above all of this is the note that these were measured differently from the paper, with a pointer to the deployment directory. That is the right place for the detail, and it is also a warning: these are numbers from the author's own harness, on the author's own GPU, measuring the author's own conversion. They are comparable to each other and to nothing else. The readme says as much.

The eight-bit column also implies that the conversion path is part of the deliverable rather than an afterthought. Getting a model to a working integer kernel is where most of the engineering effort in deployment actually goes, and having four deployment directories in the repository is consistent with an author who did that work rather than just describing it.

## The tested environment is from 2021 and will not install cleanly today

There is a short platform section, and it is a time capsule.

The operating system is a long-term-support release from 2018. The GPU is a four-series card, which is fine, that is a good inference card. The driver is from a 2020 branch, the CUDA runtime is at version 10 or 11, the deep learning library is 8, the Python is 3.8 point 8, and the machine learning framework is 1.11 point 0. Every one of those numbers is a coherent set from late 2021, and the reason they appear together is that the results in the tables were produced on that stack.

The Python version is the first practical obstacle. A 3.8 interpreter is end of life, and the current releases of the machine learning framework do not support it at all, so a person following the readme exactly on a current machine will not be able to install the framework version the code expects. A person installing a current framework version will find that a six-year-old research codebase may or may not still work with it, and the readme does not say.

The driver and CUDA versions matter for a different reason. The eight-bit throughput numbers in particular depend on the conversion toolchain, and conversion output has changed across driver generations. A driver from 2020 and a driver from 2026 will not produce identical integer kernels, so the fastest numbers in the readme are the numbers most sensitive to the parts of the stack that are hardest to install.

None of this is a criticism of the author, who published the platform details rather than leaving them out, which is more than most reproductions do. The published numbers are the published numbers, and they were produced on that stack. The correct reading is that this repository documents a reproducible experiment rather than a supported package, and reproducing it means building something like that 2021 environment, probably in a container, which the repository does not appear to provide.

That is the single most useful thing you could add to this project if you wanted to help, and it is a few hours of work: a container definition pinning the Python, the framework and the CUDA versions, so that the variance note and the results tables could be checked by somebody other than the author.

## Fewer operations, many more iterations, and a checker script for your own data

Two short notes at the end of the training section, and one detail in the dataset section, are the most practically useful things in the readme for somebody who wants to train rather than download.

The first note is a direct trade. The second architecture has fewer floating point operations and requires substantially more training iterations to reach its result, and the first architecture's training time is shorter. So the efficiency claim in the architecture comparison is paid for in compute, and a team with a fixed compute budget will make a different choice from a team comparing final accuracy on an unlimited budget. That is exactly the kind of thing a results table cannot tell you, and it is in the readme in one sentence.

The second note is about batch size and memory. The author used an overall batch size of sixteen for every model, split across more than two GPUs, and explains why: the largest of the three datasets has one hundred and seventy-one categories, which needs more memory than the others. The distributed launcher script in the repository is the reference for how the split was done. This is the kind of operational detail that is invisible in a paper and determines whether your own run fits on your own hardware.

The dataset section has the feature that most impresses me, which is a checker. For a custom dataset you write a text file with one comma-separated pair per line, an input image path and a ground truth path. Then you run a script that prints information about your dataset. Nobody stops you skipping that step, but a script that validates a data layout before you spend a day of training on a typo is worth more than most of the framework code, and its presence suggests the author has lost a day to a missing file.

The fine-tuning path is also complete, using the framework's distributed launcher with a flag pointing at an existing checkpoint and the training script in mixed-precision mode. Being able to start from a released model rather than from scratch is the difference between a twenty-minute experiment and a two-day one, and it is one flag.

Finally, the inference entry points are ordinary and complete: a script that takes a configuration, a weight path and an image, and writes a result image, and a second that does the same for a video file or, by passing a camera identifier instead of a path, for a live camera. That last substitution is the kind of detail that tells you the code has been used for something other than a demo.

## Conclusion

This repository is the right starting point if you want to train or fine-tune either of these architectures yourself rather than download somebody else's weights, and it is unusually honest about the limits of its own results, which is rarer than it should be. It is a poor fit if you were planning to pick between the two architectures based on the numbers in the first table, because the author has told you the difference is inside the run-to-run spread, and the single release tag means the six weight files are frozen at a moment when the code has since moved. Reproduce the environment before you trust anything, pin the driver and runtime versions named in the readme rather than whatever your machine has, and read the variance note before you draw a conclusion from a one-point difference.

## FAQ

### What does the CoinCheung/BiSeNet repository contain?

A PyTorch implementation of two real-time semantic segmentation architectures, with training code, dataset preparation for three segmentation benchmarks, evaluation under four protocols, six pretrained models and deployment directories for four inference targets including a mobile library and two different vendor toolkits.

### How reliable are the mIoU numbers in the BiSeNet results tables?

The author states that the model has a large run-to-run variance and gives a specific example: the single-scale score for one configuration varies between about 73.1 and 75.1 across repeated training runs. That two-point spread is larger than most of the differences between the two architectures in the tables, so the readme's own caveat says the tables cannot settle the comparison.

### Where can I download the pretrained BiSeNet models?

All six models, two architectures across three datasets, are attached to a single release tagged 0.0.0 in August 2020. The code in the repository has continued to change since then, so the weights and the current code are not from the same point in the project's history and it is worth checking the output before relying on them.

### What are the four evaluation protocols in the BiSeNet readme?

Single scale, single scale with the image cropped, multi scale with horizontal flipping, and multi scale crop with flipping. The exact input scales and crop size used for the multi-scale protocols are in the configuration files in the repository rather than in the readme, which is what makes the published numbers comparable to another implementation's.

### Which inference runtimes does the BiSeNet repository support?

Four, each with its own directory: a GPU inference toolkit, a mobile-focused inference library, a toolkit from a different hardware vendor, and a general model serving system. The frames-per-second table gives figures at full, half and eight-bit integer precision, and the eight-bit column shows the largest gains, roughly four times the full precision throughput for one architecture.

## Sources

- [CoinCheung/BiSeNet on GitHub](https://github.com/CoinCheung/BiSeNet)
- [Issues](https://github.com/CoinCheung/BiSeNet/issues)
- [License: MIT](https://github.com/CoinCheung/BiSeNet/blob/master/LICENSE)
- [README](https://github.com/CoinCheung/BiSeNet/blob/master/README.md)
- [Releases](https://github.com/CoinCheung/BiSeNet/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/coincheung-bisenet
