Max Headbox: the browser holds the agent, the backend only touches the Pi
Tiny truly local voice-activated LLM Agent that runs on a Raspberry Pi
At a glance
- What is it?
- A GPL-3.0 voice activated assistant for a Raspberry Pi 5 that runs its models through Ollama on the LAN, with Vosk for the wake word and faster-whisper for transcription. The architecture choice that defines it is that the web app owns the agent logic while a Python service handles the hardware, and the FAQ is unusually candid about the shortcuts.
- Who is it for?
- Max Headbox is worth reading if you want a self contained voice assistant with no cloud dependency and you are comfortable wiring three processes and exposing Ollama on your network. Three things to decide first.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 60 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The browser owns the agent loop, Express only reaches the hardware
The architecture decision is stated plainly in the project's own question and answer section, and it explains most of the rest. The author wanted the web app to be the most important part of the project, containing the logic of the actual agent, and to use the Express and Python backend layer only for interacting with the Raspberry Pi hardware, meaning the microphone and the transcription services. The stated benefit is replaceability: that layer could be rewritten in a different stack and reconnected to the frontend if needed. The proof that this was a real refactor rather than a claim is in the history, since an update dated 9/3/2026 records that Ruby is no longer needed because the backend was rewritten in Express.js, and the original version was written in Ruby Sinatra. The frontend talks to Ollama directly, which is why the Ollama host has to be network reachable rather than localhost.
One npm script starts three processes and two of them are servers
The run command hides more than it looks. The manifest defines prod-start as a build followed by three things run at once: the Vite preview server, the whisper service, and the backend. The whisper service is the Python half, started as uvicorn against backend.whisper_service:app bound to 0.0.0.0 on port 8000 with two workers, and binding to all interfaces is deliberate because the browser on another device has to reach it. The backend is a single Node process, node backend/server.js, which is the Express server holding the WebSocket endpoint. Development uses the same three without the build step. Two dependencies explain parts of the design that would otherwise look odd: streaming-json is there to parse the model's output as it arrives, and nodemailer is a dependency at all, which tells you the agent can send mail through a tool.
Ollama must be bound to 0.0.0.0 before the browser can reach it
Out of the box Ollama listens locally, and the frontend runs in a browser somewhere else, so the project walks through overriding the unit. The edit is made with systemctl edit, the service stanza gains one environment line, and the daemon is then reloaded and the service restarted:
sudo systemctl edit ollama.service[Service]
Environment="OLLAMA_HOST=0.0.0.0"sudo systemctl daemon-reload && sudo systemctl restart ollamaTwo models are then pulled, a 1b and an e2b variant:
ollama pull gemma3:1b
ollama pull gemma4:e2bThe security consequence is worth stating without editorializing: that override makes the model server reachable by anything on your network, and the project does not add authentication to compensate. The environment file then has to point three variables at the right addresses, and the first two deliberately share one because the WebSocket app runs inside Express rather than beside it.
Three VITE_ variables, one address, and a recordings directory in RAM
Configuration is a short environment file with three network variables and one path:
VITE_BACKEND_URL=http://192.168.XXX.XXX:4567
VITE_WEBSOCKET_URL=ws://192.168.XXX.XXX:4567
VITE_OLLAMA_URL=http://192.168.XXX.XXX:11434The backend and the WebSocket share an address because the WebSocket app also runs on Express, which is why 4567 appears twice with two schemes, while Ollama keeps its own default port. If Ollama lives on a different machine you have to give it that machine's address, so this file is also the switch that turns a single Pi setup into a distributed one. The fourth setting is not a network value: the recording directory defaults to /dev/shm/whisper_recordings, which on Linux is a RAM disk, so captured audio never touches the SD card by default. Running on a different operating system means overriding it, and the documented example points it at a Desktop folder instead.
A tool is four properties, and a .txt file is inert until renamed
The extension mechanism is small enough to describe exactly. A tool is a JavaScript module in src/tools/ that exports an object with four properties: the tool's name, the parameters passed to the function, a describe field, and the function's main execution body. There is no registry file to edit and no decorator. The catch is that some tools cannot work from the browser alone, because the frontend cannot query the Pi's hardware directly, so those need backend API handlers exposing the data over REST, and the author keeps them in a backend/notions/ directory as a reference. The bundled tools ship with a .txt extension and are documentation rather than code, since importing one requires renaming it to .js. That is a deliberate, if unusual, choice: it makes the examples safe by default and puts the decision to enable a tool in your hands.
dangerous: true turns a tool into a YES or NO gate
There is one safety mechanism, and it is opt-in per tool. If you consider a tool dangerous and want confirmation before the agent executes it, you set the property dangerous: true when creating it. When the model selects that tool, the app asks for confirmation before running anything, and you reply with YES or NO. Nothing else is guarded: there is no allowlist, no permission model, no per-tool timeout, and no distinction between a tool that reads a file and one that sends mail, which matters because mail is one of the shipped dependencies. The design consequence is that the safety of the system is a property of the tools you wrote, not of the agent, so an agent with the ability to run arbitrary JavaScript modules from its own tool directory is exactly as safe as the least careful module in it.
Vosk for the wake word, faster-whisper for what you said
Two speech projects do two jobs, and the project's own answer admits the overlap is redundant. Voice activation, meaning the wake word, is handled by Vosk through its Node API. Transcription is handled by faster-whisper, which the project points at for efficiency and accuracy and links a separate setup tutorial for. Asked why not reuse faster-whisper for the wake word too, the answer is that the goal was to get the wake word system working and the question was deferred. The stated architecture is therefore two Python services, the uvicorn whisper service on port 8000 with two workers, plus whatever Vosk runs inside the backend, and a dependency tree that keeps model inference on the Pi. The animated character in the interface is a lightly modified Microsoft Fluent Emoji set, which is the only part of the UI that is not React.
A committed .env, a 0.0.0 version, and a candid FAQ
Repository hygiene is mixed in ways worth knowing. A .env file sits at the top level of the tree rather than being ignored, so whatever addresses and paths you put in yours will be visible to git unless you change that, and the documented template contains a LAN address placeholder rather than a secret, which makes it a habit to fix rather than an active leak. Alongside it sit a .nvmrc, an eslint config, a package lock file, a Vite config and an index.html, which is a conventional Vite and React layout. The package is marked private at version 0.0.0 with no GitHub releases, so there is nothing to pin, and the project is GPL-3.0 licensed. The final section of the documentation is the most useful part for a reader deciding whether to trust it: the author states the intention to migrate off Ollama to llama.cpp eventually, concedes that the UI animations do cost inference time, and answers a question about how much of the code was written with an AI assistant.
Editorial conclusion
Max Headbox is worth reading if you want a self contained voice assistant with no cloud dependency and you are comfortable wiring three processes and exposing Ollama on your network. Three things to decide first. It needs a Pi 5 with about 6GB free for the models, an always-on microphone, and an active cooler, and the author is explicit that the screen and case bundle is optional but the cooler is not. Second, the Ollama override that binds it to 0.0.0.0 is the step that makes the browser able to reach the model, and it also means your LAN can reach it. Third, a .env file sits committed in the repository root, so check what is in yours before you copy the layout. The last push was on 2026-08-03 and there are no releases.
Frequently asked questions
Can I run AI locally on a Raspberry Pi?
That is what this project does. It runs 100% locally on-device with no cloud AI provider, on a Raspberry Pi 5 tested on 8GB and 16GB models, using Ollama with gemma3:1b and gemma4:e2b pulled locally. The documentation asks for roughly 6GB available to run the models.
Can a Raspberry Pi be used for AI?
For voice interaction with a local model, yes, and this repository is a working example: a Raspberry Pi 5, a microphone for voice commands, and about 6GB free for the models. The optional screen, case and cooler bundle provides the display, with the active cooler described as the part you should not skip.
What hardware does Max Headbox need?
A Raspberry Pi 5, tested on 16GB and 8GB models, a microphone for voice commands, and optionally a GeeekPi screen, case and cooler bundle. On the software side it needs Node 22, Python 3 and Ollama, plus about 6GB of available space for the models.
How do I add a tool to Max Headbox?
Create a JavaScript module in src/tools/ that exports an object with four properties: name, parameters, describe and the execution body. Tools needing Pi hardware data also need an Express route exposing it over REST, and the bundled examples ship with a .txt extension that you rename to .js to import.
How does Max Headbox confirm a dangerous tool?
You set the property dangerous: true on the tool when you create it. When the model selects that tool the app asks for confirmation before executing it, and you answer YES or NO. The safeguard is per tool and opt-in, so it does not cover tools you do not mark.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/syxanash-maxheadbox)