conformer: version 1.0 dates from 2022 and the torch floor from 2020
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
At a glance
- What is it?
- conformer is an unofficial PyTorch implementation of the INTERSPEECH 2020 speech recognition architecture, and it contains model code only. The training lives in another repository, the version has not moved since 2022, and the usage example stops mid-statement.
- Who is it for?
- conformer is worth reading if you want the Conformer block itself rather than a training stack, and the reference implementation is small enough to read in an afternoon. Two things to be clear about.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 103 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The model is here, the training is in another repository
The opening paragraph describes the architecture in one sentence: Transformers capture global interactions, convolutional networks exploit local features, and Conformer combines the two to model local and global dependencies of an audio sequence in a parameter-efficient way.
Then the sentence that decides what this repository is for: it contains only model code, and you can train with Conformer at openspeech.
So the split is deliberate and clean. This is the block. The recipe, the data pipeline, the training loop and the decoding all live in the openspeech repository, and nothing here will train a model on its own.
That is reinforced by the reference list, which points at the Conformer paper, Transformer-XL, the reference Transformer-XL implementation, and espnet. So there are three projects in the same lineage here: this one holds the model, openspeech is the recommended trainer, and espnet is credited as prior art.
The words unofficial in the project description are also part of the same statement. This is not the reference implementation from the authors who published the architecture, it is one person's PyTorch version of it.
For someone reaching this repository expecting to fine-tune a speech model, the useful thing to take away is the block itself, and the block is small.
Version 1.0 dates from 2022, the torch floor from 2020, and the Python floors differ
The packaging metadata is where the age shows.
setup(
name='conformer',
packages = find_packages(),
version='1.0',
description='Convolution-augmented Transformer for Speech Recognition',
author='Soohwan Kim',
author_email='[email protected]',
url='https://github.com/sooftware/conformer',
install_requires=[
'torch>=1.4.0',
'numpy',
],
keywords=['asr', 'speech_recognition', 'conformer', 'end-to-end'],
python_requires='>=3.6'
)Three floors, three eras. `version='1.0'` and the only tagged release are from February 2022. The PyTorch floor of 1.4.0 is from that project in 2020. The Python floor in the metadata is 3.6, while the installation notes recommend 3.7 or higher.
There is no `pyproject.toml` in the repository, so this is setup.py-only packaging. The only supported install route is an editable install from a checkout:
pip install -e .No wheel, no build backend declaration, no optional extras for a specific CUDA build.
The dependency list is correspondingly thin: torch and numpy, with the installation notes sending you to the PyTorch site to choose the build that matches your environment and offering a separate note about problems installing numpy.
The last push to the repository is dated 2026-06-29, so the code has been touched this year. The version number has not been, which means the tag tells you nothing about the state of the tree.
The licence badge points at a different repository
The badge row at the top of the page links four things: a licence, PyTorch, PEP 8, and a Zenodo DOI. Three are correct.
The licence badge resolves to:
https://github.com/sooftware/jasper/blob/main/LICENSEThat is a sibling repository under the same owner, not this one. The Conformer page's licence link therefore sends you to read another project's licence text.
The licence itself is not in doubt. The packaging metadata records Apache 2.0, the repository tree contains a `LICENSE` file, and `setup.py` carries the full Apache License 2.0 header in the source, copyright 2021 Soohwan Kim, with the standard grant and warranty disclaimer. So the terms are consistent across three places and only the badge is mislinked.
It is a small thing, and the kind that survives because nobody clicks a badge to check whether the link resolves to the project it sits on. But it is also the sort of copy-paste inheritance that tells you how this repository was assembled: from another one, by the same author, doing the same kind of work.
The author contact is published in the page as well as in the metadata, the same address in both.
The usage example stops mid-statement and wires the model to a CTC loss
The usage section is a Python block, and it ends in the middle of a list literal. The last line reads `target_lengths = torch.LongTensor([9,` and stops there.
What survives is still informative about the intended shape of a batch:
batch_size, sequence_length, dim = 3, 12345, 80
criterion = nn.CTCLoss().to(device)
inputs = torch.rand(batch_size, sequence_length, dim).to(device)
input_lengths = torch.LongTensor([12345, 12300, 12000])The feature dimension is 80, which is the usual mel filterbank width for speech, so the inputs are random values standing in for acoustic features. The batch has three items and the lengths differ per item, 12345, 12300 and 12000, which is exactly what a CTC loss needs: each sequence has its own valid length rather than one padded length for the batch.
Two observations. The loss shown is a CTC loss rather than anything specific to this architecture, so the objective and the decoding live with the training project, not here. And the sequence length of 12345 is an unusual number for speech frames, small enough to be a character or token count in the original example it was adapted from. Either way, it is a hard-coded value in the snippet rather than something the model constrains.
Device selection is handled in the snippet as well, choosing cuda when available and otherwise cpu, and moving the criterion to the same device.
The newest tagged artefact is a Zenodo archive, not a release
Two tags exist. v1.0 was published on 2022-02-21. The other is named v1.0-zenodo, with the release title Zonodo release, and it is dated 2026-01-05. The page carries a DOI badge pointing at a Zenodo record with identifier 18154427.
A Zenodo record is how a research artefact gets a citable identifier and a fixed archive, which is what a paper, a dataset or a model contribution needs. It is not a distribution channel. Nobody installs from it, it has no upgrade path, and it does not tell you which PyTorch the code was run against.
The repository tree backs up the same reading. Alongside the code there is a `CITATION.cff` file, which is the citation metadata format, and a `docs/` directory. So the repository has been prepared to be cited and archived rather than to be consumed as a dependency.
That is a perfectly legitimate shape for this kind of contribution, and it is consistent with everything else on the page. A model block that has not needed to change since 2022 does not need releases; it needs an identifier.
The practical consequence is that if you depend on this code, pin a commit rather than a version, and expect to supply your own training loop, your own decoding and your own PyTorch version.
Small patches direct, features negotiated, and the documentation comes from docstrings
The contributing note draws a line that most projects draw and few state plainly.
You are told to feel free to proceed with small issues such as bug fixes and documentation improvements, and to discuss major contributions and new features with the collaborators in the corresponding issues first. So the default is to open an issue before writing anything substantial, and the exception is mechanical work.
The code style note explains why style is not cosmetic. PEP 8 is followed, and the sentence that matters is that the style of docstrings is important to generate documentation. The documentation is built from the docstrings, so a badly written docstring is a missing documentation page rather than a cosmetic problem.
Support runs through GitHub issues, with the author's email address also published on the page.
The tree is correspondingly small: a `.gitignore`, `CITATION.cff`, `LICENSE`, the `README.md`, a `conformer/` package directory, `docs/`, and `setup.py`. There is no test directory, no continuous integration configuration and no development dependency group, which is consistent with a reference implementation kept for reading rather than a package kept for building.
Editorial conclusion
conformer is worth reading if you want the Conformer block itself rather than a training stack, and the reference implementation is small enough to read in an afternoon. Two things to be clear about. It says so itself: the repository is unofficial, model code only, and training belongs to a different project, so this is not something you point at a corpus and run. And it is an archived artefact rather than a moving dependency: version 1.0 was published in February 2022, the packaging floor allows a 2020 PyTorch, and the newest tagged artefact is a Zenodo archive for citation rather than a release. Use it as reference code, pin your own PyTorch, and do not expect a maintained package.
Frequently asked questions
Can conformer be used to train a speech recognition model directly?
No. The repository states that it contains only model code, and points to the openspeech repository for training with Conformer. Installing it from source with `pip install -e .` gives you the model block and the two dependencies, torch and numpy.
What Python and PyTorch versions does conformer declare?
The packaging metadata sets `python_requires='>=3.6'` and `install_requires=['torch>=1.4.0', 'numpy']`, while the installation notes recommend Python 3.7 or higher and send you to the PyTorch site to pick a build for your environment.
What licence is conformer released under?
Apache 2.0, recorded in the packaging metadata, in the LICENSE file in the repository root, and in the Apache header inside setup.py. The licence badge on the page itself links to a different repository by the same author.
Is conformer the official implementation of the Conformer paper?
No. The project describes itself as an unofficial PyTorch implementation of the architecture from the INTERSPEECH 2020 paper, and the reference list cites espnet alongside the paper and Transformer-XL.
How current is the conformer repository?
The version has been 1.0 since the release dated 2022-02-21, and the newest tag is a Zenodo archive named v1.0-zenodo dated 2026-01-05. The last push to the repository is 2026-06-29, so the tree has moved while the version has not.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sooftware-conformer)