VibeVoice after the TTS pullback: hour-long ASR, a 1.58 GB CPU model, and no releases
GitHub describes it as Open-Source Frontier Voice AI. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- Microsoft's VibeVoice is no longer the 90-minute multi-speaker TTS model the README announced in August 2025: the TTS code was removed from the repository on 2025-09-05, leaving speaker-labeled ASR checkpoints, a realtime TTS model, and a separate CPU build, with no GitHub releases to pin against.
- Who is it for?
- VibeVoice-ASR-7B is a defensible pick for teams that need speaker-labeled, timestamped transcripts of hour-long recordings across many languages and can accept holding transformers at 4.51.3 through the streamingtts extra. Skip it for long-form multi-speaker synthesis, since Microsoft removed that code on 2025-09-05 even though the 1.5B weights are still linked, and treat the streaming and BitNet paths as separate products with their own language and hardware limits.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The TTS row still links weights, but the TTS code was pulled out of the repository
On 2025-09-05 Microsoft stated that VibeVoice is an open-source research framework intended to advance collaboration in the speech synthesis community, and that after the release they discovered instances where the tool was used in ways inconsistent with the stated intent. Because responsible use of AI is one of their guiding principles, they removed the VibeVoice-TTS code from this repository. The model table still lists VibeVoice-TTS-1.5B with a Hugging Face weight link, and its Quick Try column reads Disabled. The 2025-08-25 entry records what that model was: long-form multi-speaker text-to-speech, synthesizing up to 90 minutes with up to 4 distinct speakers, accepted as an Oral at ICLR 2026. The split matters in concrete terms. Weight files remain reachable through the Hugging Face link, while the runtime for them does not sit in the tree. Top-level entries are .gitignore, CONTRIBUTING.md, Figures/, LICENSE, README.md, SECURITY.md, demo/, docs/, finetuning-asr/, pyproject.toml, vibevoice/, and vllm_plugin/, and the only fine-tuning directory is finetuning-asr/. A team planning a long-form multi-speaker TTS pipeline has to source the runtime elsewhere, and this repository names no place to source it from.
7.5 Hz continuous tokenizers are what make hour-long input tractable
VibeVoice is built on a pair of continuous speech tokenizers, one acoustic and one semantic, running at an ultra-low frame rate of 7.5 Hz. The claim attached to that rate is that the tokenizers preserve audio fidelity while boosting computational efficiency for processing long sequences, and it is the mechanism behind the long-input claims rather than a decorative number. Generation sits on a next-token diffusion framework: a Large Language Model reads the textual context and the dialogue flow, and a diffusion head adds the high-fidelity acoustic detail on top. That division has a practical consequence for anyone wiring this into a product. The language model half is what lets the ASR side attribute speakers and place timestamps, and it is why the model wants conversational context rather than isolated utterances fed in one at a time. The diffusion half is a distinct generation stage, which is why the family is not one artifact but a set of separately trained checkpoints. Each one carries its own Hugging Face weight link, its own file under docs/, and its own quick-try destination in the model table, so choosing a member of the family is choosing which of those three routes you are willing to maintain.
ASR returns Who, When, and What instead of a flat transcript
VibeVoice-ASR is designed to take one long file rather than many short windows. It handles 60-minute long-form audio in a single pass and generates structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), which is the difference from a flat transcript with speaker labels attached afterward. The framing is explicitly against models that slice audio into short chunks and lose global context in the process. On top of that it supports User-Customized Context, the same idea that appears in the streaming model as customized hotwords, so domain vocabulary can be pushed into decoding. Language coverage is the second headline: the model is natively multilingual across more than 50 languages, with the per-language breakdown at docs/vibevoice-asr.md#language-distribution. Around the checkpoint sit several documented routes, a Playground for trying it, a finetuning recipe in finetuning-asr/README.md, vLLM inference in docs/vibevoice-vllm-asr.md, a Hugging Face Transformers release dated 2026-03-06, and an Azure AI Foundry Labs integration dated 2026-03-12. None of those routes carry a version of their own, which matters when a transcript feeds something audited later.
Streaming is a separate model with ten languages instead of fifty-plus
The streaming model is not the 7B checkpoint with a flag flipped. VibeVoice-ASR-Streaming, released 2026-09-03, is a unified streaming ASR model that continuously transcribes who said what as speech arrives and emits text once per chunk, and it is published as its own Hugging Face collection with its own documentation at docs/vibevoice-asr-streaming.md. Two documented constraints follow from it being a separate model. First, the language list shrinks. The streaming release names 10 languages alongside support for customized hotwords, against more than 50 for VibeVoice-ASR, so picking streaming for latency means giving up most of the coverage the 7B checkpoint advertises. Second, the code surface is separate. The demo folder carries vibevoice_asr_streaming_inference_from_file.py and vibevoice_asr_streaming_fastapi_demo.py, and serving it has its own path through docs/vibevoice-vllm-asr-streaming.md. A deployment that wants both accurate batch transcription and live captions is therefore stitching together two checkpoints, two documents, and two sets of example scripts rather than configuring one model twice.
The CPU build lives in another repository with a different size story
VibeVoice-ASR-BitNet, released 2026-07-23, is an edge CPU inference engine for VibeVoice-ASR, and its code is a separate project at github.com/microsoft/VibeASR.cpp rather than anything under the vibevoice/ package. The compression detail is specific: heterogeneous quantization with I8_S and I2_S takes the model from 4.62 GB down to 1.58 GB and keeps real-time inference, RTF below 1, on 3 or more CPU threads with no GPU required. That gives the family two runtimes with two maintenance surfaces. The Python package installed from this repository does not carry the CPU engine, and the declared dependency list gives a clear reason to expect that: torch, transformers, accelerate, llvmlite, numba, diffusers, tqdm, numpy, scipy, librosa, ml-collections, absl-py, gradio, av, aiortc, uvicorn, fastapi, pydub, and requests, with no vLLM and no quantization packages listed. The 1.58 GB figure describes the quantized BitNet build, not the 4.62 GB original, so an on-disk budget written from that number has to say which artifact it assumes.
transformers gets a range in the base install and an exact pin in the streamingtts extra
Two version constraints land on transformers at once, and only one of them is negotiable. The base dependency is transformers>=4.51.3,<5.0.0, a range that permits any 4.x release above the floor, while the optional dependency group streamingtts pins transformers==4.51.3 with an exact equality. Adding the streamingtts extra to an environment that already carries a newer transformers therefore forces a downgrade to one specific release, and any other component in that environment needing a later transformers will collide with it. The package also declares an entry point under vllm.general_plugins mapping the name vibevoice to vllm_plugin:register_vibevoice, so vLLM finds the plugin through installed package metadata instead of an explicit command line option. The declared version is 1.0.0, requires-python is >=3.10, and setuptools is configured to include only the vibevoice* and vllm_plugin* packages, which leaves demo/ and finetuning-asr/ as trees you run from a clone rather than import from a wheel.
The experimental voice list is nine languages and eleven English styles, and it is labeled provisional
The one text-to-speech path with a runnable quick try is VibeVoice-Realtime-0.5B, open-sourced 2025-12-03 as a real-time text-to-speech model that supports streaming text input and long-form speech generation, reached through the Colab notebook at demo/vibevoice_realtime_colab.ipynb. Its voice catalog is where the caveats sit. On 2025-12-16 the project added experimental speakers for exploration, covering multilingual voices in nine languages, DE, FR, IT, JP, KR, NL, PL, PT, ES, plus 11 distinct English style voices, with the note that more speaker types will be added over time. The word experimental carries weight there, because these voices are described as being for exploration, they arrive through a separate helper, demo/download_experimental_voices.sh, and the 2025-12-03 announcement gives no duration figure for its long-form generation the way the 90 minutes was quoted for the removed TTS model. A production voice menu built on this set is built on voices the project itself frames as provisional, with the catalog expected to change.
No GitHub releases means no version to pin and nothing to roll back to
The repository has no GitHub releases, and the 1.0.0 declared in pyproject.toml is a packaging version rather than a published tag with release notes behind it. The last push is 2026-09-03, the same date as the VibeVoice-ASR-Streaming announcement, so the tip of main is that streaming work rather than a frozen stable snapshot. The README does not document a rollback procedure either, and with no release history there is nothing to roll back to. The reproducibility gap shows up at every layer. Pinning the Python package does not pin the checkpoint the model table links to, the streaming model in docs/vibevoice-asr-streaming.md carries no version string of its own, and the BitNet weights sit in a different repository with a different release cadence entirely. For a team that has to reproduce a transcript six months from now, the only durable record is one you write yourself: the exact Hugging Face checkpoint identifier plus the commit you installed.
Editorial conclusion
VibeVoice-ASR-7B is a defensible pick for teams that need speaker-labeled, timestamped transcripts of hour-long recordings across many languages and can accept holding transformers at 4.51.3 through the streamingtts extra. Skip it for long-form multi-speaker synthesis, since Microsoft removed that code on 2025-09-05 even though the 1.5B weights are still linked, and treat the streaming and BitNet paths as separate products with their own language and hardware limits. Before committing, confirm three things: which checkpoint your language actually sits in, whether the transformers version already in your environment survives the streamingtts exact pin, and how you will pin a checkpoint given that the repository has no GitHub releases and no release history to fall back on.
Frequently asked questions
Is Vibe voice free?
The repository is MIT licensed and every model in the table has a Hugging Face weight link, with no paid tier and no GitHub releases paywall. Trying the models goes through a Playground for VibeVoice-ASR, a Colab notebook for VibeVoice-Realtime-0.5B, and a demo folder of Python scripts, so the cost of trying it is hardware and setup rather than a license.
What is the Microsoft Vibe voice model?
VibeVoice is a family of open-source voice AI models covering both text-to-speech and automatic speech recognition. It uses continuous acoustic and semantic speech tokenizers at a 7.5 Hz frame rate and a next-token diffusion framework in which a Large Language Model handles textual context and dialogue flow while a diffusion head generates the acoustic detail.
how to install vibevoice locally
The README does not document a pip install line, so the concrete starting points are pyproject.toml, which declares the package name vibevoice at version 1.0.0 with requires-python >=3.10, and the docs/ files for the model you want. Weights are fetched separately through the Hugging Face links in the model table, and the runnable examples are demo/vibevoice_asr_inference_from_file.py and demo/vibevoice_realtime_colab.ipynb.
how to use vibevoice asr
VibeVoice-ASR handles 60-minute long-form audio in a single pass and returns structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for User-Customized Context. The project documents a vLLM path in docs/vibevoice-vllm-asr.md, a finetuning recipe in finetuning-asr/README.md, and a streaming variant that emits text once per chunk in docs/vibevoice-asr-streaming.md.
how to use vibevoice in comfyui
No ComfyUI integration appears in the repository, whose top-level entries are .gitignore, CONTRIBUTING.md, Figures/, LICENSE, README.md, SECURITY.md, demo/, docs/, finetuning-asr/, pyproject.toml, vibevoice/, and vllm_plugin/. The demo directory holds Python inference scripts, a Gradio demo, a FastAPI streaming demo, a Colab notebook, voice files, and a web directory rather than node definitions.
how to use vibevoice 7b
VibeVoice-ASR-7B is the long-form recognition checkpoint, listed in the model table with a Hugging Face link and a Playground as its quick try. It is the checkpoint that carries the more than 50 language claim, while the streaming model released 2026-09-03 names 10, so language coverage is the deciding difference between them.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-vibevoice)