Open-source project
digantamisra98/Mish avatar
digantamisra98/Mish

Mish is a 2020 activation function whose faster versions all live in other people's repositories

Official Repository for "Mish: A Self Regularized Non-Monotonic Neural Activation Function" [BMVC 2020]

1,298 stars128 forksJupyter NotebookMIT

At a glance

What is it?
The official repository for a self-regularized, non-monotonic activation function published at BMVC 2020. Its primary language is Jupyter notebooks rather than a package, its declared homepage is the paper PDF, it has published no releases, and the most useful page in it is a changelog of which frameworks adopted the function and a list of variants maintained by other authors.
Who is it for?
Read this repository as the record of a research result rather than as a library to install. The function reached PyTorch, MXNet, the TensorFlow ecosystem, medical imaging and the JVM, and it is almost never the thing you fetch from here: there are no releases, the declared homepage is a PDF on an external archive, and the usable implementations are the ones your framework already ships.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 77 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The homepage is the paper PDF, and the main language is a notebook

This repository is a landing page for a paper, and the metadata says so before you read a word of it.

The declared homepage is not a documentation site or a package page. It is the proceedings PDF hosted on an external conference archive. The description names the paper and its venue, and the readme links the official paper, an arXiv entry with a version number, a citation export, and a hosted notebook viewer.

The repository's primary language is Jupyter Notebook, which tells you where the substance lives. The root holds a notebook on layer accuracy, an implementation directory named after the function, an observations folder, a benchmarks folder, an experiments folder, a configuration folder, a loss landscapes folder, a folder of reviews and citations, a BibTeX citation file, a licence, and three media files including a poster image, a poster PDF and an animated image with an unrelated filename at the root.

There are no published releases. That is normal for this kind of repository and it has a consequence worth stating: nothing here is versioned, so if you need the function pinned, you take it from whichever framework you already depend on, not from here.

The adoption changelog is the most informative page in the readme

One expanded section lists where the activation function was added, and it is a better guide to the current state than anything else in the repository.

The framework list is long and covers several ecosystems. It appears in PyTorch, in MXNet, in TensorFlow JS and TensorFlow Swift, in ONNX tooling and in a Kotlin deep learning library, in TorchSharp, in OneFlow, in a JVM automation project, and in two smaller neural network libraries. On the application side it appears in MONAI for medical imaging, in a plaid implementation, in an image processing library and its network tool, in a language model implementation, and in Google's AutoML.

Also listed are a YOLO implementation and a high-performing convolution network in PyTorch, where the readme records a multi-scale configuration as the top result on a well known object detection test split and a single-scale configuration as third, at a stated date.

The distribution is the point. This is a research result that reached mainstream frameworks within months, which is unusual, and the practical consequence is that by now the argument over whether to use it has been settled for you by whatever your framework ships.

The changelog dates have no year on them, and one entry is out of order

Two defects in that changelog are worth pointing out, because they are the kind that make a document unreliable rather than merely imperfect.

First, every entry is stamped with a day and month and no year. The entries run from mid-year dates in the middle of the list to dates in the following months, and the years are simply absent. A reader can guess them from context, since the paper acceptance and the arXiv revision bracket the publication, but nothing in the file states it. For a project whose last commit is years after these entries, guessing is the only option.

Second, the list is not chronological. A June entry appears after a September entry, and the August entries sit before it. So even the relative order is unreliable, which means you cannot reconstruct the adoption timeline from the document that exists specifically to record it.

There is also one entry that is a promise rather than a fact. The PyTorch line states both that the function was added and that it will be added in a particular future release. A changelog entry of that shape is not evidence that a framework shipped it, and nothing else in the repository confirms the outcome. For any framework you care about, check its own release notes instead of this list.

Every faster implementation of Mish lives in someone else's repository

There is a second expanded section listing variants, and reading it tells you where the engineering work went.

A considerably faster version based on a compiled GPU kernel is credited to another author and hosted in his repository. A memory efficient experimental version is not in this repository at all; it lives inside a separate efficient network project, in a file about activation functions, pinned to a specific commit. Faster variants for both this function and its bounded cousin are credited to a third author in a building blocks repository. An alternative, experimental and improved version of the bounded variant was developed by a fourth and is available in Julia.

And an experimental initialisation method based on variance sits in a gist rather than in the repository.

So the situation is: the original author published the paper and the reference implementation, and every subsequent improvement to speed or memory was contributed downstream and is maintained by the person who wrote it. There is no single place where the best available implementation lives.

For a function whose appeal is cheap expressiveness, that is a real limitation. Before adopting it, work out which of these implementations your framework already vendors, because pulling one from a gist pinned to a commit is a different maintenance commitment from using the one in your deep learning library.

The evidence lives in benchmarks, landscapes and a citations folder

The folders at the root are the argument, once you know what each one is for.

A benchmarks directory holds updated PyTorch benchmarks and pretrained models, published as a separate entry in the changelog. A notebook on layer accuracy is the kind of analysis that asks whether trained layers behave differently with this activation, and it is the only file in the repository surfaced through a hosted notebook viewer.

A loss landscapes folder corresponds to a changelog entry about loss landscape exploration done in collaboration with another researcher. That work is interesting for a specific reason: the intuition for a smooth, well-behaved loss surface is one of the arguments made for this family of activation functions, and landscapes are how you either show it or fail to.

Then there is an observations folder, an experiments folder, a configuration folder, and a folder whose name is reviews and citations. That last one is unusual and useful: a research repository that curates what other people have written about the result is telling you the paper has been picked up, reviewed and argued with.

A BibTeX file sits beside them all, so citing the work does not require visiting the paper page.

The closest thing to an ablation is a morphology comparison with two rivals

Among the media entries there is one item that is closer to a scientific comparison than the rest, and it is worth separating from the talks.

A collaboration produced a comparison of the morphology of this activation, a competing activation and the rectified linear unit. Morphology here means the shape of the function itself, which is the thing that distinguishes the smooth, bounded, non-monotonic family from the piecewise linear one, and comparing shapes is a reasonable way to argue about optimisation behaviour without training anything.

The other entries in that section are podcasts, conference talks and videos: a podcast episode about the function, a talk on the function and non-linear dynamics at a machine learning company, and a further talk. A poster was accepted for a deep learning summer school hosted by several Canadian institutes, and the poster files are still sitting in the repository root.

So the public footprint is almost entirely 2020 events and one collaboration video. None of it has been updated, and there is no follow-up work listed, which is consistent with a result that landed, was adopted widely and then stopped.

Editorial conclusion

Read this repository as the record of a research result rather than as a library to install. The function reached PyTorch, MXNet, the TensorFlow ecosystem, medical imaging and the JVM, and it is almost never the thing you fetch from here: there are no releases, the declared homepage is a PDF on an external archive, and the usable implementations are the ones your framework already ships. Two things to check before you build on it. The performance work that would make it attractive, faster kernels and memory-efficient variants, is all maintained by other people in other repositories, so its performance characteristics depend on a downstream project you did not choose. And the claimed state of the art results in the readme are dated 2020 with month-only timestamps, so treat them as a snapshot rather than a current ranking.

Frequently asked questions

What is the Mish activation function?

A self-regularized, non-monotonic neural activation function presented at the British Machine Vision Conference in 2020, with an arXiv paper and an official proceedings PDF. The repository described here is the official one for that result: MIT licensed, with the implementation, notebooks, benchmarks and a BibTeX citation file, and no published releases.

Is Mish available in PyTorch?

The readme's changelog records a PyTorch entry that both states the function was added and says it would be added in a later release, which is a promise rather than a shipped fact and is not confirmed anywhere else in the repository. Either way, you would take the function from PyTorch rather than from this repository, which publishes no releases.

What is H-Mish and where are the faster versions of Mish?

H-Mish is the bounded variant, and the faster implementations of both are maintained by other authors in other repositories: a GPU kernel version, a memory efficient version inside an efficient network project pinned to one commit, faster variants in a building blocks repository, an alternative improved H-Mish available in Julia, and a variance based initialisation method published as a gist.

What benchmarks and analysis are in the Mish repository?

A benchmarks directory with updated PyTorch benchmarks and pretrained models, a notebook on layer accuracy, an observations folder, an experiments folder, a configuration folder, and a loss landscapes folder from exploration done in collaboration with another researcher. The declared homepage of the repository is the proceedings PDF rather than any documentation site.

How do I cite the Mish paper?

A BibTeX citation file sits at the repository root, and the readme links the arXiv entry with a version number and the official conference paper. The repository also maintains a folder collecting reviews and citations, which the readme presents as part of the record of the result.

Official sources

  1. digantamisra98/Mish on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/digantamisra98-mish.svg)](https://hysenlabs.com/projects/digantamisra98-mish)