# TheWhisper: Optimized Whisper Inference for Apple Silicon and NVIDIA GPUs

> TheWhisper is a Python package that provides fine-tuned Whisper models and an optimized inference engine for real-time speech transcription on Apple Silicon and NVIDIA GPUs. It adds streaming support, flexible chunk sizes from 10 to 30 seconds, and CoreML acceleration for on-device macOS deployments that the original Whisper models do not offer.

**TheStageAI/TheWhisper** — Optimized Whisper models for streaming and on-device use

- Repository: https://github.com/TheStageAI/TheWhisper
- Website: https://thestage.ai
- Stars: 898 · Forks: 56
- Language: Python
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/thestageai-thewhisper

## What TheWhisper Adds Beyond the Original Whisper Models

OpenAI's Whisper models are fixed to a 30-second chunk window. TheWhisper publishes fine-tuned model weights on Hugging Face that support 10, 15, 20, and 30-second chunk sizes. This matters for streaming: a 10-second chunk delivers partial results to the user faster than a 30-second chunk, which is important for real-time captioning or voice interfaces where latency is perceptible.

Beyond chunk sizes, TheWhisper provides two distinct inference backends: a CoreML backend for Apple Silicon that targets approximately 2 watts of power consumption and 2 GB of RAM for the large-v3 model, and a NVIDIA GPU backend that includes both a HuggingFace Transformers path and an accelerated TheStage AI ElasticModels path.

The repository also includes a local REST API with a FastAPI server and an Electron demo app for macOS called TheNotes, which is Apple-certified. A tutorial on the TheStage AI blog documents building a note-taking app using Electron and TheWhisper.

## Hardware Requirements and Support Matrix

The Apple path requires macOS 15.0 (Ventura) or later, a minimum of 2 GB RAM (4 GB recommended for the large-v3 model), Python 3.10 to 3.12, and an M1 or later Apple Silicon chip. Supported chipsets range from M1 through M4 Pro and M4 Max. iOS 18.0 or later is listed as a supported platform as well.

The NVIDIA path requires Ubuntu 20.04 or later, CUDA 11.8 or higher, driver version 520.0 or higher, Python 3.10 to 3.12, and a minimum of 2.5 GB RAM. Tested GPUs include the RTX 4090, RTX 5090, L40s, H100, A100, and Jetson-Thor. The TheStage AI optimized engines report 220 tokens per second on an L40s for the whisper-large-v3 model.

All four model variants (whisper-large-v3 and whisper-large-v3-turbo, on both Nvidia and Apple) support streaming, hardware acceleration, word timestamps, multilingual transcription, and all four chunk sizes.

The NVIDIA path has a separate install step for the TheStage AI optimized engines, which requires a platform token configured with: thestage config set -t <YOUR_API_TOKEN>. This token is generated from a profile on the TheStage AI platform. Without it, the HuggingFace Transformers path still works as a fallback.

## Installing and Running a First Transcription

The repository uses optional dependency groups to keep Apple and NVIDIA paths separate. Clone the repository first:

```bash
git clone https://github.com/TheStageAI/TheWhisper.git
cd TheWhisper
```

For Apple Silicon:

```bash
pip install .[apple]
```

For NVIDIA with the HuggingFace Transformers backend:

```bash
pip install .[nvidia]
```

A basic Apple inference call looks like this:

```python
from thestage_speechkit.apple import ASRPipeline

model = ASRPipeline(
    model='TheStageAI/thewhisper-large-v3-turbo',
    model_size='S',
    chunk_length_s=10
)

result = model(
    "path_to_your_audio.wav",
    return_timestamps="word"
)

print(result["text"])
```

For streaming from a microphone, the StreamingPipeline, MicStream, and StdoutStream classes handle the audio loop. The streaming example in examples/run_streaming.py demonstrates the full pattern.

The pyproject.toml defines the Apple extras as including mlx==0.25.2 and coremltools==8.3.0 alongside torch 2.7.0 and transformers 4.52.3. The NVIDIA extras include tensorrt 10.13.3.9 and torch 2.9.0. Both extras share the same base dependencies declared in the main dependencies block: editdistance, librosa, silero-vad, fastapi, uvicorn, and sounddevice, among others.

## The Streaming Architecture and Its Design Constraints

TheWhisper's streaming implementation uses a two-output model: each chunk returns an 'approved_text' value that has been committed to the transcript and an 'assumption' value that represents text that may change as more audio arrives. This is a common approach for streaming ASR that trades latency for correctness: the caller receives partial results quickly but must handle the possibility that assumptions are revised.

The step_size_s parameter on MicStream controls how often new audio chunks are pushed through the pipeline. Setting it to 0.5 seconds means the model sees fresh data twice per second. Shorter steps increase CPU load; longer steps increase perceived latency. The README does not document a lower bound on step_size_s, so tuning this value for a specific device requires empirical testing.

The examples/ directory includes separate scripts for Apple inference, NVIDIA inference, and streaming. The examples/server.py file shows the FastAPI-based REST API setup, which can serve remote clients over HTTP. Connecting a browser or Electron front end to the server requires that the client and server agree on the endpoint format documented in the electron_app/ directory.

For the Jetson-Thor embedded platform, the installation requires tensorrt==10.13.3.9 pre-installed on the device before adding the thewhisper packages. The Jetson path uses a separate JFrog index URL for Jetson-compatible packages. This makes Jetson-Thor the most complex supported target, and the README warns that edge deployments require platform-specific setup not covered by the standard pip install commands.

## Limitations and When to Use a Simpler Alternative

TheWhisper is not a drop-in replacement for OpenAI's Whisper API. It is a self-hosted solution that requires a compatible GPU or Apple Silicon machine. Running it in a cloud VM without GPU passthrough will not benefit from the accelerated paths.

The TheStage AI optimized engine for NVIDIA requires installing packages from a private JFrog registry (thestage.jfrog.io) and a TheStage AI API token obtained from the platform. For small organizations the ElasticModels path is documented as free, but the README references an enterprise license for production use without describing the terms in the repository itself.

For teams that only need standard batch transcription without streaming or variable chunk sizes, faster-whisper is a widely used alternative. faster-whisper uses CTranslate2 for inference and is available as a standalone Python package without a registry dependency. It does not offer CoreML acceleration for Apple Silicon, which is where TheWhisper has a distinct advantage.

## License, Maintenance, and the Electron Demo

The repository is licensed under MIT. The last push was on 2026-09-23, five days before this writing, and the repository has no formal GitHub releases. The package published to PyPI is named thestage-speechkit at version 0.1.0.

The Electron demo app, TheNotes for macOS, is a downloadable Apple-certified pkg that demonstrates the CoreML streaming path in a production application. Its source and build instructions are in the electron_app/ directory. A build tutorial is on the TheStage AI blog. The app serves as both a working reference and a testing surface for the CoreML backend.

The benchmark/ directory contains methodology and results comparing latency, memory use, power consumption, and accuracy (via OpenASR) across model variants and hardware. The README notes these figures cover all four model configurations on both Apple and NVIDIA hardware.

## Conclusion

TheWhisper is the right choice for teams deploying Whisper on Apple Silicon hardware or on NVIDIA GPUs where inflight streaming and sub-30-second chunk sizes matter. The CoreML path is specifically designed for on-device scenarios with tight power budgets, while the NVIDIA path targets server-side real-time transcription. Teams planning production deployments should review the enterprise license summary in the README: the MIT license covers the repository itself, but the TheStage AI optimized engines may have separate terms. The Electron demo app (TheNotes for macOS) shows one concrete integration path using Electron and ReactJS.

## FAQ

### What chunk sizes does TheWhisper support beyond the original Whisper's 30 seconds?

TheWhisper's fine-tuned models support 10, 15, 20, and 30-second chunk sizes. The original Whisper models are fixed to 30 seconds. Shorter chunks are particularly useful for streaming, where they reduce time to first partial result.

### Does TheWhisper require a TheStage AI account?

The HuggingFace Transformers backend for NVIDIA and the Apple CoreML path do not require an account. The TheStage AI optimized engines for NVIDIA do require an API token from the TheStage AI platform and packages from a private registry.

### What is the power consumption of TheWhisper on Apple Silicon?

The README states approximately 2 watts of power consumption and approximately 2 GB of RAM usage for the CoreML path on Apple Silicon, making it suitable for extended on-device deployments.

## Sources

- [Issues](https://github.com/TheStageAI/TheWhisper/issues)
- [License: MIT](https://github.com/TheStageAI/TheWhisper/blob/main/LICENSE)
- [Project website](https://thestage.ai)
- [README](https://github.com/TheStageAI/TheWhisper/blob/main/README.md)
- [TheStageAI/TheWhisper on GitHub](https://github.com/TheStageAI/TheWhisper)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thestageai-thewhisper
