Model or dataset
vndee/local-talking-llm avatar
vndee/local-talking-llm

vndee/local-talking-llm: a local voice assistant built from Whisper, Ollama and Chatterbox TTS

A talking LLM that runs on your own computer without needing the internet.

892 stars189 forksPythonMIT

At a glance

What is it?
A Python project that chains speech recognition, a pluggable LLM backend and Chatterbox text-to-speech into one offline loop. It is a short console script rather than a product, and the README is honest about the parts that will bite you.
Who is it for?
Adopt it if you already have a machine that can run Ollama and you want a readable reference for wiring Whisper to Chatterbox through LangChain, not a finished assistant. Skip it if you need a daemon, a wake word, or anything that survives a restart.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 179 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: a conversational loop that never leaves your machine

Most voice assistant demos are thin clients for a cloud API. You speak, a request leaves your network, and a hosted model answers. vndee/local-talking-llm takes the opposite position: the README describes a system that listens, thinks and speaks with no internet dependency, and the repository is small enough to read in one sitting. The top-level entries are app.py, tts.py, pyproject.toml, requirements.txt, a Makefile and a tests directory. That is the whole product.

The audience is narrow and specific. This is for Python developers who want to see how speech-to-text, an LLM call and text-to-speech fit together without a framework hiding the seams. It is also for people who have a reason to keep audio and transcripts off someone else's servers. The README frames it as building something Jarvis-like, and the project name is literal: a talking LLM that runs locally. If you want a packaged desktop assistant, this is not that. If you want a working reference implementation you can modify, the shape is right.

Whisper in, LangChain in the middle, Chatterbox out

The README lays out three components. Speech recognition uses OpenAI's Whisper to turn spoken language into text. The conversational chain uses the LangChain interface with a pluggable backend, either a local model through Ollama (the README names Gemma3 and Llama-4) or a cloud model through MiniMax (MiniMax-M2.7). Speech synthesis uses Chatterbox TTS from Resemble AI.

The data flow is a straight line: record speech, transcribe to text, generate a response, vocalize it. The README includes a Mermaid flowchart with exactly those stages, from user speech input through transcription, the conversational chain, the generated response, the synthesizer, and audio output. There is no queue, no session store, no state machine. Each turn is independent.

Chatterbox is the May 2025 change. The README states the project previously used Bark and that the old implementation is preserved in the archive-2025-05-29 branch. The new model is listed as 0.5B parameters with voice cloning, emotion control, and neural watermarking on the generated audio. Those four claims come from the README's feature list, not from any independent measurement, and the repository does not include benchmarks that would let you check the speed claim yourself. Treat the parameter count as a sizing hint: a 0.5B TTS model plus Whisper plus an LLM is three models resident at once, and that is the real constraint on which machine this runs on.

Installing with uv and getting a first spoken turn

The README is unusually direct about the install path: it recommends uv over pip and warns that requirements.txt was generated by uv pip freeze and contains pinned versions that may not install correctly across different systems. That warning matters, because a pinned file with torch==2.6.0 and numpy==1.26.0 will fail on some platforms. Use pyproject.toml, which declares requires-python >=3.11.

Install uv, clone the repository, and sync the environment. The README also asks for NLTK sentence-tokenization data, which is a separate download from the Python packages.

bash
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/vndee/local-talking-llm.git
cd local-talking-llm
uv sync
source .venv/bin/activate
python -c "import nltk; nltk.download('punkt_tab')"

Then install Ollama, start it, and pull a model. The README uses gemma3 as the example and notes you can substitute any model you prefer.

bash
ollama pull gemma3

With the environment active and Ollama running, the entry point is app.py. The README's basic usage is a single command, and the console output is styled with rich.

bash
python app.py

You should expect the first run to spend time downloading model weights for Whisper and Chatterbox before any audio is processed. Subsequent runs reuse the cache. Optional flags documented in the README include --voice with a path to a 10 to 30 second WAV sample for cloning, --exaggeration and --cfg-weight for emotion and pacing, --model to choose a different LLM, and --save-voice to keep generated samples. If you have no local GPU, the README offers MiniMax as a cloud backend: set MINIMAX_API_KEY in your environment, and speech recognition and synthesis still run locally while only the LLM call leaves the machine.

Where the design runs out of road

The honest limitation is that this is a script, not a service. Running python app.py gives you a console loop. There is no daemon, no systemd unit, no Docker file in the repository listing, and no documented way to run it headless behind another process. If you want an assistant that is always listening, you are writing that yourself.

Resource pressure is the second issue, and it is structural rather than a bug. Whisper, the LLM and Chatterbox all want memory at the same time. The README offers MiniMax specifically for people without a local GPU, which is an admission that the fully local path has hardware requirements the project does not quantify. There is no table of minimum specs, so you find out by trying.

The third is platform. The dependency list leans on sounddevice, torchaudio and pyaudio-adjacent audio libraries, and the README gives activation commands for both POSIX shells and Windows, which suggests the author expects Windows users to hit friction. Audio device selection is not documented at all: the README does not explain how to choose an input device when the default is wrong, and that is the most common failure in local audio code. If your microphone is not the system default, you are reading app.py.

Finally, requirements.txt is a trap for the unwary. The README says so plainly, but the file is still in the repository, and a reader who reaches for it first will get an environment that may not resolve.

How it compares to a hosted assistant stack

The obvious alternative is not another open source repository but the cloud path itself: send audio to a hosted speech-to-text API, send the transcript to a hosted chat model, and stream the reply through a hosted TTS voice. That approach removes the memory ceiling entirely. You can run it on a laptop with no GPU, latency is predictable, and voice quality is typically higher than a 0.5B local model.

The difference in approach is where the models live, and that changes what you can do with the system. A hosted stack cannot answer when your network is down. It cannot process audio you are not permitted to upload. It also cannot be inspected: you get an API contract, not a LangChain chain you can edit. vndee/local-talking-llm gives you the chain, the TTS wrapper in tts.py, and the option to swap the LLM backend by changing a flag or an environment variable. The cost is that you own the hardware, the model downloads and the audio device configuration.

Within the local category, the meaningful comparison is to the project's own past. The archive-2025-05-29 branch holds the Bark-based implementation. The README's stated reasons for moving to Chatterbox are voice cloning, emotion control, a smaller model and watermarked output. If you need the older behavior for any reason, that branch is where it lives.

Maintenance, licence and what an upgrade costs you

The repository is not archived, and the last push was on 2026-04-04. That is roughly five and a half months before today's date, so it is recent enough to call maintained without stretching the word, but there are no retrieved releases, and the project is at version 0.1.0 in pyproject.toml. There is no changelog and no tagged release to pin against. Upgrades mean pulling main and re-running uv sync.

That re-sync is not cheap. The dependency set includes torch 2.6.0, torchaudio 2.6.0, transformers 4.46.3, diffusers 0.29.0, librosa, numba and scikit-learn, plus the three model families. A resolver change or a torch bump can invalidate a working environment, and because requirements.txt is a freeze rather than a constraint file, it will not protect you. The Makefile offers a lint target that runs pre-commit across all files, which is the only automated check visible in the repository listing alongside the tests directory.

The licence is MIT. That is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and licence text are kept. Two caveats are worth stating without pretending to be legal advice. First, MIT covers this repository's code, not the model weights it downloads. Whisper, Chatterbox and whichever Ollama model you pull carry their own licences, and those are what you need to check before shipping anything. Second, the README notes that Chatterbox output carries a neural watermark, which is a property of the generated audio rather than a restriction on your code, but it is something to be aware of if you plan to distribute the audio.

Editorial conclusion

Adopt it if you already have a machine that can run Ollama and you want a readable reference for wiring Whisper to Chatterbox through LangChain, not a finished assistant. Skip it if you need a daemon, a wake word, or anything that survives a restart. Before you commit, run uv sync, download the NLTK punkt_tab data, and check that your GPU can hold both the Whisper and Chatterbox models in memory at once.

Frequently asked questions

Can I use vndee/local-talking-llm without an internet connection?

Yes, that is the point of the project. The README describes a voice assistant that operates offline, with Whisper for speech recognition, a local Ollama model for the conversation and Chatterbox TTS for the reply. You still need a connection for the initial model downloads and for installing dependencies.

Can I run the LLM in vndee/local-talking-llm locally?

Yes. Ollama is the default backend, and the README's setup steps are to install Ollama, pull a model such as gemma3, and run python app.py. The README also lists MiniMax as a cloud alternative for machines without a local GPU.

What is the best LLM to run locally with vndee/local-talking-llm?

The README does not rank models. It names Gemma3 and Llama-4 as examples for the Ollama backend and MiniMax-M2.7 for the cloud backend, and the --model flag lets you point at a different Ollama model. Which one fits depends on the memory you have free alongside Whisper and Chatterbox.

What is the best conversational LLM for vndee/local-talking-llm?

The README does not name a best conversational model. It presents the LLM backend as pluggable, with Gemma3 and Llama-4 given as Ollama examples and MiniMax-M2.7 as the cloud option, and leaves the choice to the reader.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. vndee/local-talking-llm on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vndee-local-talking-llm.svg)](https://hysenlabs.com/projects/vndee-local-talking-llm)