# NVIDIA NeMo Speech: A PyTorch Framework for ASR, TTS, and Speech LLMs

> NVIDIA NeMo Speech is a Python and PyTorch framework for researchers and developers building automatic speech recognition, text-to-speech, and speech large language model systems. It provides training infrastructure, pre-trained model checkpoints on HuggingFace, and NVIDIA GPU containers, and it narrowed its scope to audio and speech in 2026 after splitting from the broader NeMo framework.

**NVIDIA-NeMo/Speech** — A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

- Repository: https://github.com/NVIDIA-NeMo/Speech
- Website: https://docs.nvidia.com/nemo/speech/nightly/index.html
- Stars: 18,468 · Forks: 3,611
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-nemo-speech

## What NeMo Speech Targets: ASR, TTS, and Speech LLMs

NVIDIA NeMo Speech is built for researchers and PyTorch developers working on three categories of speech models. Automatic speech recognition (ASR) converts audio to text. Text-to-speech (TTS) converts text to audio. Speech LLMs combine these capabilities with large language models.

The framework provides tools to create, customize, and deploy new models by starting from existing code and pre-trained checkpoints. The pre-trained checkpoints are published on HuggingFace under the nvidia/nemotron-speech collection.

NeMo Speech is not a hosted API or a finished product. It is a research and development framework for engineers who are training, fine-tuning, or evaluating models. Users are expected to have Python, PyTorch, and CUDA installed and to be familiar with deep learning training workflows. The README says the framework works with Python 3.12 or above, PyTorch 2.7 or above, and an NVIDIA GPU with CUDA for training. CUDA is also recommended for inference, though not strictly required.

## The 2026 Repository Split: Why NeMo Speech Focuses on Audio

The repository carries an explicit note about a scope change. In 2026, the broader NeMo repository split into domain-specific repositories. NVIDIA-NeMo/Speech is the result for audio, speech, and multimodal LLMs. The final pre-split release was v2.7.3, available in the 26.04 NeMo NGC container.

From v3.0.0 onward, support for modalities other than audio and speech (such as text LLM training, vision, and multimodal outside of speech) is no longer in this repository. The README directs users who need those other modalities to v2.7.3 explicitly.

The practical effect of this split is that NVIDIA-NeMo/Speech is a narrower and more focused tool than the pre-split NeMo was. Engineers who used NeMo before 2026 for non-speech tasks need to check whether their use case is still covered here or has moved to a separate repository. For speech and ASR specifically, the split represents a cleaner separation of concerns and more focused release management.

## Installing NeMo Speech from Source with uv

The README recommends installing from source using uv, a Python package manager. This approach reproduces the actively-tested stack from the committed uv.lock file:

```bash
git clone https://github.com/NVIDIA-NeMo/Speech.git
cd Speech
uv sync --extra all --extra cu13
```

The --extra cu13 flag installs support for CUDA 13.x, which is the recommended version. For CUDA 12.x systems, replace --extra cu13 with --extra cu12. On Linux, these two extras are mutually exclusive: pass exactly one.

This installs the actively-tested stack (Python 3.13, PyTorch 2.12, CUDA 13.2) into a .venv/ directory with NeMo Speech in editable mode. Additional install groups are available: --group test adds the test suite, and --group docs adds the documentation build tools. After activation, run tools with uv run or activate the environment with:

```bash
source .venv/bin/activate
```

For teams who need different Python, PyTorch, or CUDA versions than the locked stack, the README also describes a pip fallback that installs over an existing environment. The package name on PyPI is nemo-toolkit, as shown in pyproject.toml.

For a fully reproducible build that matches the Dockerfile and CI, the README recommends adding --locked --python 3.13 to the uv sync command. The NVIDIA NGC container (26.07.00 NeMo Speech) is the pre-built alternative for teams who prefer container-based deployment.

## Pre-Trained Models: Parakeet, Nemotron ASR, and MagpieTTS

NeMo Speech provides several pre-trained model checkpoints on HuggingFace under the nvidia/ namespace. These checkpoints are the primary starting point for engineers who want to use the framework without training from scratch.

Parakeet-unified-en-0.6b is an English ASR model that supports both offline and streaming inference within a single checkpoint. The README notes a minimum streaming latency of 160ms. Canary v2 covers speech recognition and translation for 25 European languages.

Nemotron-3.5-ASR-Streaming-0.6B is a streaming ASR model with 40 language support and controllable latency ranging from 80ms to 1s. It is built on a cache-aware Fastconformer architecture, according to the README.

MagpieTTS v2607 is a TTS model covering 12 languages: Arabic, Korean, Portuguese (new in v2607), plus English, Spanish, German, French, Vietnamese, Italian, Chinese, Hindi, and Japanese. Nemotron VoiceChat is a separate model that combines ASR, an LLM backbone, and TTS for full-duplex voice conversations. The README notes VoiceChat was in Early Access as of 2026-03 and points to an application form.

All of these checkpoints are separate downloads from the NeMo Speech framework itself. Installing NeMo Speech does not automatically download any of them.

## Streaming ASR and the Latency Trade-Off

A notable feature of several NeMo Speech ASR models is controllable latency in streaming mode. Nemotron-3.5-ASR-Streaming-0.6B allows latency to be set between 80ms and 1s. Parakeet-unified-en-0.6b has a minimum streaming latency of 160ms. These are presented in the README as fixed constraints of the current checkpoints, not as configurable thresholds in the framework itself.

The ability to control latency is described as picking an optimal point on the latency-accuracy trade-off curve. Lower latency produces faster transcription results but typically at the cost of some accuracy, because the model has less audio context to work from when generating each word. Higher latency allows more context and generally produces higher accuracy.

For teams building real-time voice applications, this trade-off is the key variable to evaluate before choosing a checkpoint. Nemotron-3.5-ASR-Streaming-0.6B supports 240 to 2400 concurrent streams on a single H100 GPU at 1x throughput, according to the README. These figures are specific to that one checkpoint and GPU type.

## A Checkpoint Loading Concern and Where NeMo Speech Falls Short

PyTorch 2.6 changed the default behavior of torch.load to use weights_only=True. The README notes that some NeMo model checkpoints may fail to load with this default and require the environment variable TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1 to be set before running code that calls torch.load.

The README adds an explicit caution: this environment variable should only be set for trusted files, because loading files from untrusted sources with weights_only=False carries the risk of arbitrary code execution. Teams who want to load third-party or community checkpoints through NeMo Speech should be aware of this risk before disabling the PyTorch safety default.

NeMo Speech is not the right tool for CPU-only inference. CUDA is required for training, and the README recommends it for inference as well. Teams running on hardware without NVIDIA GPUs face a significant performance gap.

The framework is targeted at researchers and PyTorch developers, not at application teams that want a finished API. There is no REST API or simple inference endpoint included. Building a production-ready speech service on top of NeMo Speech requires additional engineering work for serving, scaling, and monitoring.

The examples/ directory in the repository contains example code organized by modality (asr/, tts/, speaker_tasks/, and others), which provides starting points for common tasks, but the examples are not a substitute for working through the full developer documentation.

## Maintenance, License, and the NGC Container Path

The repository received a push on 2026-09-17, within six months of this article, and the most recent major release is v3.0.0, published on 2026-08-07. This confirms the project is actively maintained on the post-split timeline.

The license is Apache-2.0, as stated in the LICENSE file and the pyproject.toml. The Apache 2.0 license permits commercial use, modification, and distribution with attribution. The THIRD-PARTY-NOTICES file in the repository root lists third-party components and their licenses.

For teams who prefer a container over a source install, the 26.07.00 NeMo Speech NGC container is the pre-built option for v3.0.0. The NGC container is built and tested by NVIDIA. For the previous NeMo (pre-split) version, the 26.04 NeMo NGC container corresponds to v2.7.3.

The technical documentation is hosted at docs.nvidia.com/nemo/speech, with a nightly build for the main branch and a stable build for v3.0.0. The developer documentation is the reference for API details, training recipes, and model configuration options that are not covered in the README.

## Conclusion

NVIDIA NeMo Speech is the right tool for researchers and ML engineers who are building or fine-tuning ASR, TTS, or speech LLM systems on NVIDIA hardware and who need pre-trained checkpoints as starting points. It is not suited to teams looking for a simple speech-to-text API wrapper or a CPU-only inference library: NVIDIA GPU and CUDA are required for training, and the install process through uv or pip adds non-trivial dependency overhead. Before starting, verify that Python 3.12 and PyTorch 2.7 or later are available, and decide whether to use the source install with uv or the pre-built NGC container (26.07.00 NeMo Speech NGC container for v3.0.0) for a reproducible baseline.

## FAQ

### What pre-trained speech models does NVIDIA NeMo Speech include?

NeMo Speech publishes pre-trained checkpoints on HuggingFace under nvidia/nemotron-speech. As of the README, these include Parakeet-unified-en-0.6b (English ASR, offline and streaming), Nemotron-3.5-ASR-Streaming-0.6B (40-language streaming ASR, 80ms-1s latency), Canary v2 (ASR and translation for 25 European languages), and MagpieTTS v2607 (TTS for 12 languages). The checkpoints are downloaded separately from the framework.

### Does NVIDIA NeMo Speech require an NVIDIA GPU?

An NVIDIA GPU with CUDA is required for training. The README recommends it for inference as well, though inference can run without it. The uv install defaults to CUDA 13.x (--extra cu13) or CUDA 12.x (--extra cu12). The pre-built NGC container (26.07.00) bundles a tested CUDA environment.

### How do I install NVIDIA NeMo Speech?

The recommended install is from source using uv: git clone the repository, then run uv sync --extra all --extra cu13 (or --extra cu12 for CUDA 12.x). This installs the tested stack (Python 3.13, PyTorch 2.12, CUDA 13.2) into .venv/. A pip fallback using the nemo-toolkit package name is also available for teams with an existing Python environment.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA-NeMo/Speech/blob/main/LICENSE)
- [NVIDIA-NeMo/Speech on GitHub](https://github.com/NVIDIA-NeMo/Speech)
- [Project website](https://docs.nvidia.com/nemo/speech/nightly/index.html)
- [README](https://github.com/NVIDIA-NeMo/Speech/blob/main/README.md)
- [Releases](https://github.com/NVIDIA-NeMo/Speech/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-nemo-speech
