Be More Agent: a local Raspberry Pi AI agent with a BMO voice
Local AI Agent running on Raspberry Pi
At a glance
- What is it?
- Be More Agent is an offline-first Python agent for Raspberry Pi that chains OpenWakeWord, Whisper.cpp, Ollama and Piper TTS behind a reactive face GUI. It is a framework for a character, not a finished product, and its setup script and config file carry most of the real decisions.
- Who is it for?
- Adopt Be More Agent if you already own a Pi 5 or a 4GB Pi 4 and want a wake-word-to-voice loop that never leaves the device, and if you accept that the character is assembled from PNG folders and WAV folders rather than configured through a UI. Do not adopt it if you need a documented API, a supported upgrade path, or cloud-grade model quality on a 2GB board.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 173 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Be More Agent actually is, and who it is for
Be More Agent turns a Raspberry Pi into a conversational device that runs entirely on the board. The README describes it as "100% locally" and lists the four stages of the loop: a wake word listener, speech-to-text, a local LLM, and neural text-to-speech, with face animations driven by the current state. The intended audience is narrow and specific. You need a Raspberry Pi 5 (recommended) or a Pi 4 with at least 4GB of RAM, a USB microphone and speaker, an LCD screen over DSI or HDMI, and a Raspberry Pi Camera Module if you want the vision path. That is a hardware build, not a software install, and the README treats it that way from the first line.
The framing that matters is the one the README uses for itself: a blank canvas. Faces are PNG sequences under faces/idle, faces/listening, faces/thinking, faces/speaking, faces/error and faces/warmup. Sounds are WAV files under sounds/greeting_sounds, sounds/thinking_sounds, sounds/ack_sounds and sounds/error_sounds. There is no plugin API and no theme file. Personality lives in the system_prompt_extras string in config.json plus whatever assets you drop into those folders. If you want a character that behaves exactly like the demo video, you are expected to draw it and record it yourself. The repository ships a start_agent.sh and a be-more-agent.desktop launcher, but neither is explained in the README, so a desktop shortcut exists without documentation of what it runs.
That design choice has a cost. Because the character is defined by directory contents rather than by a manifest, the agent loops through every image it finds in the folder for the current state. Adding a frame is a file copy; getting the timing right is not, since nothing in the README specifies frame counts, image dimensions or playback rate for the face animation. The same is true of sound categories, where the agent picks one file at random each time, which means the variety you get is exactly the number of files you drop in.
The pipeline from wake word to spoken reply
The mechanism is a linear chain, and each link is a separate project that Be More Agent wires together. OpenWakeWord handles detection with a local wakeword.onnx model, and the README notes it is offline and needs no access keys. Audio then goes to Whisper.cpp for transcription, which sits in the repository as a checked-out directory rather than a pip package. The transcript goes to Ollama for generation, and the reply goes to Piper TTS, which lives in a piper/ directory with its voice models. The GUI updates the face folder according to state, so the character is listening, thinking, speaking or idle as the chain progresses.
Two details in that chain are worth more than they look. First, the README claims the agent "automatically detects your microphone's sample rate and resamples audio on the fly to prevent ALSA errors", which is the kind of hardware-aware handling that usually gets skipped in hobby projects. Second, web search is conditional: when the LLM does not know the answer, the agent queries DuckDuckGo for current information. That is the only outbound network call in the described flow, and it is the one piece that contradicts the offline-first label when it fires. The README also lists a vision path using the Moondream model and a connected camera, configured through the vision_model key.
State is also persisted. config.json carries a chat_memory flag, and the repository root contains chat_memory.json, which is where conversation history is kept. That matters for two reasons. A long-running agent accumulates a transcript on disk, and the README does not describe a rotation or truncation policy for it. It also means the personality you set through system_prompt_extras is not the only thing shaping replies; prior turns feed back into the prompt. If you want a device that forgets between sessions, chat_memory is the switch to look at, and no other key in the example config appears to control persistence.
Installing Be More Agent on a Raspberry Pi
The README assumes Raspberry Pi OS and starts with system packages. Run these two commands to update the board and get git, which setup.sh needs in order to exist at all.
sudo apt update && sudo apt upgrade -y
sudo apt install git -yOllama is the brain, so it installs separately with the vendor script. The README then pulls two models, gemma:2b for text and moondream for vision. Note the discrepancy before you copy anything: the installation block pulls gemma:2b, while the example config.json in the same README sets text_model to gemma3:1b. Pull the model you actually intend to name in config.json, or the agent will ask Ollama for a model you never downloaded.
curl -fsSL https://ollama.com/install.sh| sh
ollama pull gemma:2b
ollama pull moondreamClone the repository and run the installer. According to the README, setup.sh installs system libraries, creates the necessary folders, downloads Piper TTS, and sets up the Python virtual environment. It also downloads the compiled BMO voice model and its JSON configuration from the Releases page into a local voices/ directory, so you do not need to fetch bmo.onnx and bmo.onnx.json by hand unless you skip the script.
git clone https://github.com/brenpoly/be-more-agent.git
cd be-more-agent
chmod +x setup.sh
./setup.shStart the agent inside the virtual environment. The setup script downloads a default wake word of "Hey Jarvis"; to use your own, train a model at OpenWakeWord, place the .onnx file in the root folder and rename it to wakeword.onnx. On the first run, agent.py creates config.json if it does not exist, and the README's example file is the shape you are editing: text_model, vision_model, voice_model, chat_memory, camera_rotation and system_prompt_extras. If you installed the BMO voice manually, point voice_model at voices/bmo.onnx instead of the default piper path.
source venv/bin/activate
python agent.pyWhat you should see is the warmup face sequence from faces/warmup, a greeting sound from sounds/greeting_sounds, and then an idle face waiting for the wake word. If instead you get a traceback about a missing search library, the README's troubleshooting section points at duckduckgo-search not being present in the active virtual environment.
The sample rate trap behind the "demonic" BMO voice
The troubleshooting section is the most honest part of the documentation, and it describes a real failure mode. If your custom voice sounds deep, slow or "demonic", the README states this is almost always a sample rate mismatch between the model and the audio player, not a Piper installation problem. By default agent.py expects medium quality models and plays audio at 22050 Hz. A model trained at 48000 Hz or 16000 Hz played back at the default rate will be pitched and paced wrong. The fix is to match the sample rate, which means either choosing a voice model trained at the expected rate or changing what the agent plays at. The README does not document a config key for the playback rate, so this is a source edit rather than a setting, and that is a meaningful gap for anyone who wants to swap voices without touching Python.
Other limitations are structural. The README's troubleshooting notes that a "No search library found" error means duckduckgo-search is missing from the virtual environment, so web search silently depends on a pip package that is easy to lose when you rebuild the venv. Shutdown errors such as Expression 'alsa_snd_pcm_mmap_begin' failed on Ctrl+C are described as normal, caused by cutting the audio stream mid-sample. There is no documented API, no service definition, and no upgrade procedure. Hardware is the hard boundary: a 2GB Pi 4 is below the stated 4GB minimum, and the vision path adds Moondream on top of the text model, so a board that just barely runs the conversational loop may not have room for camera work as well. If your goal is a text assistant on a cheap single-board computer, this project asks for more hardware than that goal requires.
How it compares with a cloud assistant or a plain Ollama script
The obvious alternative is a cloud voice assistant, and the difference is architectural rather than cosmetic. A cloud assistant keeps the models on someone else's hardware, which is why it works on a cheap board and answers quickly. Be More Agent pushes every stage onto the Pi, which is why the README specifies a Pi 5 or a 4GB Pi 4. The trade is capability for locality: a 1B or 2B parameter model running locally will not match a hosted model, and the README's own answer to that gap is the DuckDuckGo lookup when the model does not know something. You are buying privacy and no API fees, and paying in model quality and setup time.
The second alternative is writing your own script against Ollama and Piper. That is genuinely less work than it sounds, and if all you want is push-to-talk transcription and a spoken reply, Be More Agent's value is in the parts you would otherwise skip: the wake word model, the sample rate detection, the state machine that drives face folders, and the random selection of WAV files per category. If you do not want a character with a face and sounds, those are the features you are carrying for nothing. If you do want them, rebuilding them is the bulk of the project, and that is the honest case for adopting this repository rather than starting from a blank file.
A third comparison is a speech assistant framework that expects a cloud speech API. Those usually give you better recognition on noisy audio and a smaller install, because the heavy models are remote. Be More Agent's OpenWakeWord and Whisper.cpp stages are the price of keeping recognition local, and the README does not publish accuracy figures for either, so there is no way to judge from the documentation how well the wake word holds up in a room with a television on.
Maintenance, licence and what an upgrade costs
The repository is not archived, and the last push was on 2026-04-11, which is the same date as the v1.0-enclosure release containing 3D files for the BMO enclosure. The v1.0-voice release with the custom BMO voice model came earlier, on 2026-03-28. There is no changelog in the repository listing and no documented migration path between releases, so an upgrade means re-reading the README and re-checking config.json against whatever keys the current agent.py expects. The dependency surface is small and listed in requirements.txt: sounddevice, numpy, scipy, openwakeword, onnxruntime, ollama, duckduckgo-search and Pillow. Whisper.cpp and Piper are vendored directories, so they update on their own schedule, not through pip, and nothing in the README pins them to a version.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive licence, and it does not cover the model weights you pull: Ollama models, the Piper voice, the BMO voice model from Releases and the wake word model each carry their own terms, and the repository does not restate them. Check those separately before shipping a device that uses them. This is a description of the licence text, not legal advice.
The practical upgrade cost is the config file. Because agent.py generates config.json on first run and the README's example shows a fixed set of keys, a new release that adds or renames a key will leave an existing config.json without it, and there is no documented validation step that tells you which keys are missing. Keeping a copy of your config.json before pulling is the cheapest insurance available here.
What the README does not answer
Several things a buyer would want are simply absent. There is no stated latency figure for the full wake-word-to-speech loop, only the word "low-latency" applied to Piper. There is no memory or storage footprint for the model set. There is no rollback procedure if an upgrade breaks a working build, and no version pinning beyond what requirements.txt implies. The config.json example and the installation commands disagree on the text model, which is the kind of inconsistency that costs an evening. The face and sound folders are documented by structure rather than by specification, so image dimensions, frame counts and WAV formats are left to trial and error. The custom BMO voice is distributed through the Releases page rather than through a package index, so a manual install means two files placed into a voices/ directory you create yourself.
None of this is disqualifying for a hobby build. It does mean the README is a starting point rather than a manual. The project is honest about its failure modes where it documents them at all, which is more than many repositories of this size manage, and the troubleshooting entries for the demonic voice and the ALSA shutdown error are specific enough to act on. The gap is in the middle: between a working demo and a device you can hand to someone else, the documentation stops.
Editorial conclusion
Adopt Be More Agent if you already own a Pi 5 or a 4GB Pi 4 and want a wake-word-to-voice loop that never leaves the device, and if you accept that the character is assembled from PNG folders and WAV folders rather than configured through a UI. Do not adopt it if you need a documented API, a supported upgrade path, or cloud-grade model quality on a 2GB board. Verify two things before you commit hardware: that your chosen Piper voice model's sample rate matches what agent.py plays at, and that the model named in config.json is actually pulled in Ollama, because the README's installation block pulls gemma:2b while its example config points at gemma3:1b.
Frequently asked questions
Can I run Be More Agent for free?
Yes, in the sense that the software is MIT licensed and the stack it uses is local and open source: Ollama, Whisper.cpp, OpenWakeWord and Piper TTS. The README states there are no API fees and no cloud data usage, though the optional DuckDuckGo web search is an outbound call.
How can I build my own BMO with Be More Agent?
Install the agent on a Raspberry Pi 5 or a 4GB Pi 4, run setup.sh to fetch the BMO voice model and Piper, then replace the PNG sequences under faces/ and the WAV files under sounds/ with your own assets. The v1.0-enclosure release provides 3D files for the BMO enclosure.
What hardware does Be More Agent need?
The README lists a Raspberry Pi 5 (recommended) or a Pi 4 with at least 4GB of RAM, a USB microphone and speaker, an LCD screen over DSI or HDMI, and a Raspberry Pi Camera Module for the vision path.
Why does the BMO voice sound deep or slow in Be More Agent?
The README attributes this to a sample rate mismatch between the voice model and the audio player. By default agent.py expects medium quality models and plays at 22050 Hz, so a model trained at 48000 Hz or 16000 Hz will sound wrong until the rates match.
Is Be More Agent maintained?
The repository is not archived, and the last push was on 2026-04-11, the same date as the v1.0-enclosure release. No changelog or upgrade procedure is documented in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/brenpoly-be-more-agent)