VibeVoice
GitHub describes it as Open-Source Frontier Voice AI. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
microsoft/VibeVoice: 🎙️ VibeVoice: Open-Source Frontier Voice AI
GitHub describes it as Open-Source Frontier Voice AI. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
Repository scope
GitHub describes it as Open-Source Frontier Voice AI. The repository metadata lists Python as its primary language. The metadata lists the MIT license. The README describes the project this way: 2026-07-23: ⚡ We released VibeVoice-ASR-BitNet , an edge CPU inference engine for VibeVoice-ASR. Through heterogeneous quantization (I8 S + I2 S), the model is compressed from 4.62 GB to 1.58 GB with real-time inference (RTF Code ] [ Models ] [ Report ]
🎙️ VibeVoice: Open-Source Frontier Voice AI
The README section "🎙️ VibeVoice: Open-Source Frontier Voice AI" states: 2026-03-12: 🚀 VibeVoice-ASR is now integrated into Azure AI Foundry Labs ! You can now explore and test our unified speech-to-text capabilities directly through Microsoft Foundry.
🎙️ VibeVoice: Open-Source Frontier Voice AI
The README section "🎙️ VibeVoice: Open-Source Frontier Voice AI" states: 2026-03-06: 🚀 VibeVoice ASR is now part of a Transformers release ! You can now use our speech recognition model directly through the Hugging Face Transformers library for seamless integration into your projects.
🎙️ VibeVoice: Open-Source Frontier Voice AI
The README section "🎙️ VibeVoice: Open-Source Frontier Voice AI" states: 2026-01-21: 📣 We open-sourced VibeVoice-ASR , a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for User-Customized Context. Try it in Playground. - ⭐️ VibeVoice-ASR is natively multilingual, supporting over 50 languages , check the supported languages for details. - 🔥 The VibeVoice-ASR finetuning code is now available! - ⚡️ vLLM inference is now supported for faster inference; see vllm-asr for more details. - 📑 VibeVoice-ASR Technique Report is available.
Editorial conclusion
The repository README is the source for this review. It does not replace a local installation or an independent test.
Community notes