Open-source project
common-voice/common-voice avatar
common-voice/common-voice

Common Voice: the Mozilla web app behind a public domain speech corpus

Common Voice is part of Mozilla's initiative to help teach machines how real people speak.

3,490 stars866 forksTypeScriptMPL-2.0

At a glance

What is it?
Common Voice is a self-hostable TypeScript web application for collecting speech donations and turning them into CC0 datasets. It is built for language communities and dataset maintainers, not for teams that just want to download a corpus.
Who is it for?
Adopt Common Voice if you run a language community that needs its own donation pipeline and you can operate a MySQL, Redis and object-storage stack, or if you want to contribute sentences, translations or code upstream. Do not adopt it if you only want the data: the quarterly dataset releases are the product for you, not this repository.
Can I use it commercially?
Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Common Voice actually is, and who the repository is for

The README opens by calling this repository "the web app for Mozilla Common Voice, a platform for collecting speech donations in order to create public domain datasets for training voice recognition-related tools." That sentence is the whole scope. The repository is the donation site and its backend, not the corpus itself. The datasets ship from a separate repository, cv-dataset, on a quarterly cadence, while the platform code and the sentence collection move on a monthly or as-needed schedule.

That split decides who should care. If you are training an ASR model, you want the dataset release, and installing this code gets you nothing. If you run a language community that needs its own recording and validation pipeline, or you want to contribute sentences, interface translations or code upstream, this is the right repository. The README is explicit that non-coders have a path: interface translations go through Mozilla's Pontoon platform, and the README states that direct pull requests changing localization content are not accepted.

The architecture visible in package.json and docker-compose.yaml

The root package.json declares a Yarn workspace with four members: common, server, web and maintenance. The build script reflects that layering. It runs build:maint, then build:common, then builds the server and the web app concurrently, which means the server and the browser client share the code in common rather than duplicating types and helpers.

The runtime picture comes from docker-compose.yaml. There are four services. db runs mysql:8.0.31 and sets MYSQL_DATABASE to voiceweb, MYSQL_USER to voicecommons and MYSQL_PASSWORD to voicecommons. redis runs redis:alpine. storage runs fsouza/fake-gcs-server on port 8080 with a filesystem backend and an external URL of http://storage:8080. The web service builds from docker/Dockerfile, mounts the repository at /code, reads DOTENV_CONFIG_PATH=/code/.env-local-docker, and exposes port 9000. A fifth service, bundler, builds from ./bundler, exposes port 9001 and runs npm ci, npm run build and npm start.

The web container's command is the interesting part: it waits for storage:8080 with docker/wait-for-it.sh, runs docker/prepare_storage.sh, then installs dependencies with yarn --frozen-lockfile before starting. That ordering tells you storage initialization is a prerequisite, and the fake GCS server exists so local development does not need cloud credentials.

Installing it and recording your first clip locally

The README does not spell out setup steps itself. It points to docs/DEVELOPMENT.md for the local development environment, and the repository ships .env-local-docker.example as the template for the environment file the web container expects. The compose file already names that path, so the first move is to create the file from the example.

bash
cp .env-local-docker.example .env-local-docker
docker compose up

The compose file maps the web service to port 9000 and the database to 3306, so once the containers report healthy you open the site at http://localhost:9000. The web container will run yarn --frozen-lockfile on first boot, which takes a while because it installs the whole workspace.

The root package.json also gives a non-Docker path. The engines field requires Node 18 or newer, and start depends on the common workspace being built first.

bash
yarn build:common
yarn start

Either way, the first real use is the same as on the hosted site: pick a language, read the sentence on screen, and record it. Sentences come from /server/data, where the README says the majority of text originates from user submissions in the Sentence Collector or from Wikipedia via the cv-sentence-extractor tool.

Where the project is thin: deployment, rollback and production storage

The compose file is a development harness, and it reads like one. Storage is fsouza/fake-gcs-server with a filesystem backend, which is convenient locally and not a production object store. Nothing in the README describes a production deployment topology, a migration procedure for the MySQL schema, or how to roll back a platform release if a monthly update breaks a running instance. The release notes are linked, but the README does not document rollback.

Version pinning is another trade-off. The root package.json pins TypeScript 5.2.2, eslint 8.57.1 and a set of resolutions that force specific webpack and loader-utils versions. That is a deliberate stability choice for a project with many contributors, and it also means upgrades arrive in batches through the pinned set rather than continuously. The engines field requiring Node 18 or newer is a hard floor: a build host on an older runtime will fail before anything else runs.

Finally, this is the wrong tool for anyone whose goal is data. Running the app produces recordings in your own instance, not a copy of the published corpus. If you need the corpus, the README's own structure points you at the separate dataset repository and its quarterly releases.

How Common Voice differs from a generic speech collection tool

The obvious alternative in this space is a general-purpose annotation platform, for example a self-hosted labeling tool where you upload audio and contributors transcribe or rate it. The approach differs at the root. Common Voice starts from text: sentences live in /server/data, sourced from the Sentence Collector and from Wikipedia through cv-sentence-extractor, and volunteers read them aloud. The unit of contribution is a spoken sentence, and the sentence inventory is a maintained asset in its own right, with its own workflow document, docs/SENTENCES.md.

A generic annotation tool has no equivalent of the language workflow, docs/LANGUAGE.md, or the community playbook the README links for building a language community around Mozilla voice tools. It also has no built-in licensing posture. Here the sentence text is released under CC0, the code under MPL-2.0, and the README notes that files matching europarl-VERSION-LANG.txt were extracted from the Europarl Corpus. That licensing chain is the reason the output can be published as public domain data, and replicating it on top of a general annotation platform would be your problem, not the tool's.

Licence, maintenance and upgrade cost

The repository is MPL-2.0, and package.json carries the same identifier under the name voice-web. MPL-2.0 is file-level copyleft: modifications to files already under the licence stay under it, while separate files you add can carry other terms. This is not legal advice, and anyone embedding the platform in a commercial product should read the licence text in LICENSE rather than a summary.

The content licensing is separate and worth noting because it constrains reuse. Sentence text from the Sentence Collector and from Wikipedia is CC0, while the europarl files carry their own provenance from the Europarl Corpus. If you fork the sentence data, you inherit that mix.

On maintenance, the last push to the default branch was on 2026-09-23 and the most recent release is release-v1.160.0 from 2026-09-21, with release-v1.159.0 in June 2026 and release-v1.158.0 in April 2026. The repository is not archived. That cadence has a cost: a monthly platform release means a self-hosted instance drifts quickly unless someone owns the upgrade, and the pinned dependency set means each upgrade touches several packages at once rather than one at a time.

Editorial conclusion

Adopt Common Voice if you run a language community that needs its own donation pipeline and you can operate a MySQL, Redis and object-storage stack, or if you want to contribute sentences, translations or code upstream. Do not adopt it if you only want the data: the quarterly dataset releases are the product for you, not this repository. Before committing, verify that the Node 18 engine requirement matches your build hosts, read docs/DEVELOPMENT.md for the setup path that the README only links to, and confirm your own deployment's storage backend, since the compose file wires in fsouza/fake-gcs-server rather than a production object store.

Frequently asked questions

What is Common Voice?

It is a Mozilla platform for collecting speech donations in order to create public domain datasets for training voice recognition-related tools. This repository is the web app that runs the donation site, not the dataset itself.

What is Common Voice Mozilla?

The README describes it as part of Mozilla's initiative to help teach machines how real people speak, hosted at commonvoice.mozilla.org. The code lives in this repository under MPL-2.0 and is developed as a Yarn workspace with common, server, web and maintenance packages.

How do I use the Common Voice dataset?

The dataset is not produced by this repository. The README states that dataset releases follow a quarterly cadence and points to the separate cv-dataset repository for dataset metadata, so that is where dataset usage starts.

Official sources

  1. common-voice/common-voice on GitHub
  2. License: MPL-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/common-voice-common-voice.svg)](https://hysenlabs.com/projects/common-voice-common-voice)