Botium speech processing bundles seven containers and one opinionated default
Botium Speech Processing
At a glance
- What is it?
- A Docker Compose stack that puts Kaldi for recognition, MaryTTS for synthesis and SoX for file conversion behind a single HTTP API with Swagger docs, at the cost of 8GB of memory and a configuration surface the project tells you not to edit.
- Who is it for?
- Botium speech processing fits a team that needs speech in and out of an automated test suite, or a chatbot with a voice channel, and that has decided the quality ceiling of cloud vendors is not worth the per-minute price.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 32 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One compose file starts seven services, not one application
The default `docker-compose.yml` is the whole architecture in one document. It starts `nginx` on port 80 with the repository's `nginx.conf` mounted over the stock config, a `frontend` on the `botium/botium-speech-frontend` image that receives `BOTIUM_API_TOKENS`, `BOTIUM_SPEECH_PROVIDER_TTS: marytts`, `BOTIUM_SPEECH_PROVIDER_STT: kaldi` and two Google credential variables that ship empty, a `watcher`, two separate Kaldi services named `stt-en` and `stt-de`, a `tts` service running MaryTTS, and a `dictate` service. Six of them run as `user: 1000:1000`, and five share a single named `logs` volume. Recognisers are one container per language, which is the detail that drives both the memory bill and the language list.
Two languages by default, and the container count follows
The shipped configuration uses MaryTTS for synthesis, Kaldi for recognition and SoX for audio file conversion, with German and English included as languages. The cost note attached to the hardware requirements explains the consequence directly: memory usage can be reduced if only one language is needed for Kaldi, because the default configuration ships two. That maps onto the Compose file, where `stt-en` and `stt-de` are separate services rather than two model directories inside one container. Adding a third language is therefore not a configuration file edit in the usual sense, it is another service in the stack, another recogniser image pulled from the registry, and another entry in your resource budget. The stack is easy to extend; it is just not free to extend.
Cloud backends get slimmer compose files, not new features
The repository carries one compose file per cloud provider: `docker-compose-azure.yml`, `docker-compose-google.yml`, `docker-compose-ibm.yml`, `docker-compose-aws.yml`, `docker-compose-deepgram.yml` and `docker-compose-picotts.yml`, plus `docker-compose-frontend.yml` and `docker-compose-test.yml`. The slimming is mechanical. When you point the stack at a cloud recogniser, only the frontend service is needed, because the transcription happens on someone else's servers, and the two Kaldi containers disappear along with most of the 8GB requirement. For Azure you add a subscription key and region key to `docker-compose-azure.yml` and start with `docker-compose -f docker-compose-azure.yml up -d`. Deepgram works the same way with an API key. What you trade away is the thing that motivated the stack in the first place, which is that nothing leaves your own machine.
Configuration lives in .env, and the recommendation is not to edit it
Settings are read from environment variables, and the file to look at is `frontend/resources/.env`, which the README points you at for its comments. The advice that follows is the part worth heeding: do not change `.env` itself, create `.env.local` instead to overwrite defaults, because that prevents trouble on a future `git pull`. The distinction is the difference between a config that lives in your repository and one that does not, and it is what lets you keep the checked-in defaults as documentation of what the stack ships with. Two variants already exist at the root for other purposes, `.env.develop` for developer images and `.env` for released ones, which is how `docker-compose --env-file .env.develop up` selects a different image stream than the plain `docker-compose up -d`. Building your own images is also a compose decision, `docker-compose -f docker-compose-dev.yml up -d`, which the README flags as slow, and the Makefile at the root is the shortest path to it: a `docker_build` target that tags the frontend as `botium.speech:${SPEECH_VERSION}` and a `docker_publish` target that re-tags it onto `${AWS_REGISTRY_HOSTNAME}` and pushes. Both `develop` and `release` invoke those same two targets in order and differ by nothing in the file, so the two names are a convention rather than two behaviours.
Per-request credentials let one caller borrow a different cloud account
Certain sections of a JSON or multipart request body are read as overrides. A `credentials` object replaces the server default credentials for a cloud service, and a `config` object replaces the server default settings for that cloud API call. The documented request for Google shows both shapes, one that sends a service account `private_key` and `client_email`, and another that sends a `config` block switching the encoding to MP3 so you can post `sample.mp3` instead of `sample.wav`. An IBM request does the same with an `apikey` and `serviceUrl`. This is the mechanism that lets one shared server serve several tenants without hardcoding everyone's keys, and it also means a caller with API access can direct a request at an account that is not the one you configured.
Results are cached by hash, so a second identical call does no work
Both directions are cached, keyed by the MD5 hash of the audio for recognition and of the input text for synthesis. The stated reason is performance, and the stated way to defeat it is to empty the cache directories, `frontent/resources/.cache/stt` and `frontent/resources/.cache/tts`. Note the spelling in those two paths. The README writes `frontent` rather than `frontend`, and the same typo appears again in the request-specific configuration section, which is worth knowing before you go looking for a directory that the Compose file mounts under a different name at `./frontend/resources`. Caching by content hash means a repeated word or a repeated utterance is free, and that editing a file without changing its content will not produce a new result.
HTTP for files, websocket for streams, folders for batches
Three interfaces cover the three shapes of work. Plain requests use HTTP POST to `/api/stt/{language}` with an audio content type, and HTTP GET to `/api/tts/{language}?text=...` with the audio written out. The documented call for recognition is a POST to `http://127.0.0.1/api/stt/en` carrying a `Content-Type: audio/wav` header and the file attached with `-T sample.wav`, and the synthesis call is a GET on `http://127.0.0.1/api/tts/en?text=hello%20world` writing to `tts.wav`. Conversion is a third endpoint, `/api/convert/{profile}`, for audio file processing. Streaming opens a websocket at `/api/sttstream/{language}` and returns three URLs: `wsUri` to stream audio into, defaulting to wav chunks, `statusUri` to check whether the stream is still open, and `endUri` to finish and close. Batches avoid the API entirely: drop audio into `watcher/stt_input_de` or `watcher/stt_input_en` and collect transcripts from `watcher/stt_output`, drop text into `watcher/tts_input_de` or `watcher/tts_input_en` and collect audio from `watcher/tss_output`, with the `tss` spelling as given.
A token list is the only gate on the API, and Chrome blocks the test page
Securing the stack is one environment variable. `BOTIUM_API_TOKENS` holds the list of valid API tokens, separated by whitespace or comma, and the `BOTIUM_API_TOKEN` header is checked on every call. There is no per-user account model, no role separation and no rate limit described anywhere in the documentation, so a valid token is a full-access key to every endpoint, including the ones that accept cloud credentials from the request body. For manual testing the stack ships two browsers: `/` for Swagger UI, `/dictate/` for a rudimentary dictate.js speech recognition page, and `/tts` for the MaryTTS synthesis interface. The dictate page works with Kaldi only, and in Chrome it needs HTTPS because the service is published over plain HTTP, which you are told to arrange yourself, for example through an ngrok tunnel.
Editorial conclusion
Botium speech processing fits a team that needs speech in and out of an automated test suite, or a chatbot with a voice channel, and that has decided the quality ceiling of cloud vendors is not worth the per-minute price. Its defaults are deliberate and its toolchain is fixed, which is exactly what the README means by being highly opinionated about the included tools, and that cuts both ways: a Rasa or Botium integration gets a uniform API quickly, while anyone wanting a different recogniser has to leave the stack. Before deploying, count the containers, because the full Compose file brings up nginx, a frontend, a watcher, one Kaldi service per language, MaryTTS and a dictate test page, and the languages are chosen at build time. Check that your disk and memory budget covers 40GB and 8GB, decide whether Kaldi's default English and German are the languages you need, and set `BOTIUM_API_TOKENS` before the first request reaches port 80.
Frequently asked questions
What does Botium speech processing do?
It puts speech-to-text and text-to-speech behind one HTTP API. The default stack uses Kaldi for recognition, MaryTTS for synthesis and SoX for audio file conversion, and it ships prebuilt Docker images so nothing has to be installed natively.
How much hardware does Botium speech processing need?
The stated requirement is 8GB of RAM accessible to Docker and 40GB of free disk space for a full installation, plus internet connectivity, docker and docker-compose. Memory can be reduced by keeping only one language, since the default configuration ships two.
Can I run Botium speech processing without the local Kaldi services?
Yes. Compose files exist for Azure, Google, IBM, AWS, Deepgram and picoTTS, and using one of them makes the installation slimmer because only the frontend service is required. You add the provider's key to the compose file, such as an Azure subscription and region key, and start it with `docker-compose -f docker-compose-azure.yml up -d`.
How do I change the configuration in Botium speech processing?
Settings come from environment variables, documented in comments in `frontend/resources/.env`. The project advises against editing that file and recommends creating a `.env.local` file to overwrite defaults, so a later `git pull` does not conflict with your changes.
How is the Botium speech processing API protected?
The environment variable `BOTIUM_API_TOKENS` lists the tokens the server accepts, separated by whitespace or comma, and the `BOTIUM_API_TOKEN` HTTP header is validated on every call. No user accounts or per-token permissions are described.
Does Botium speech processing support streaming audio?
Calling `/api/sttstream/{language}` opens a websocket stream and returns three URLs: `wsUri` for pushing audio, which accepts wav chunks by default, `statusUri` to check whether the stream is still open, and `endUri` to stop streaming and close the websocket.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/codeforequity-at-botium-speech-processing)