Dicio, an on-device Android assistant that still calls home for most skills
Dicio assistant app for Android
At a glance
- What is it?
- Dicio is a GPL-3.0 voice assistant for Android that runs speech recognition, intent matching and output on the device in fourteen languages, built on Vosk and OpenWakeWord. The privacy claim holds for the voice pipeline and stops at the skill implementations, four of which query DuckDuckGo, OpenWeatherMap, Genius and a Lingva instance over the network.
- Who is it for?
- Adopt Dicio if you want a voice assistant whose microphone audio never leaves the phone and you are willing to supply the wake word model, accept an automatic model download from a third-party host, and check which of the network-backed skills you are happy to have making requests.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 68 days ago.
- What is it written in?
- Mainly Kotlin, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Dicio is, and where the on-device claim stops
Dicio is a free and open source voice assistant for Android, and the README's description is precise enough to be evaluated directly. It supports many different skills and input and output methods, gives both speech and graphical feedback to a question, and interprets user input and, when possible, generates user output entirely on-device, which the README presents as privacy by design.
That last clause carries the weight, and the qualifier inside it is doing the honest work. The on-device claim covers the pipeline: speech to text, the matching of what was said against what the assistant can do, and the spoken or graphical response. It does not cover the data the skills fetch.
The skill list makes that boundary visible. Search looks up information on DuckDuckGo, with more engines described as coming later. Weather collects information from OpenWeatherMap. Lyrics shows Genius lyrics for songs. Translation translates between languages using Lingva. Those four are network requests to third-party services, and in three of them the query text is the thing you just said out loud. The remaining skills, opening an app, the calculator, contacts, timers, the current time, navigation, jokes, media control, wake word control, notifications and the flashlight, are local operations against the device.
So the accurate description is narrower than a fully offline assistant and more useful than a cloud one. Your audio and your intent stay on the phone. What you asked about does not, unless the skill that handles it is a local one. Anyone whose threat model is a microphone, or a stranger standing too close to the phone, is well served. Anyone whose threat model includes a search provider, this is not the tool.
Fourteen languages, and a speech model that arrives when you need it
The language support is the part of Dicio that required the most engineering, and the repository layout shows it. Fourteen languages are listed: Czech, Dutch, English, French, German, Greek, Italian, Polish, Russian, Slovenian, Spanish, Swedish, Turkish and Ukrainian. Adding a language is a documented contribution, with steps at the project's documentation site.
Speech to text runs on Vosk. The design decision the README explains is about phones rather than accuracy: in order to run on every phone, small models are used, weighing around 50MB. And the download starts automatically whenever it is needed, from the model list published at alphacephei, which is what lets the app language be changed seamlessly.
Two consequences follow, and both matter for a privacy-positioned application. The first is storage and bandwidth: a language switch means a fifty megabyte download on demand, and there is no offline install path described in the README. The second is that the model host is a third party, not the F-Droid repository or the app itself, so a first run on a restricted network can fail in a way that has nothing to do with the phone.
The alternative design would have been to bundle one or two models and treat the rest as an optional download from your own mirror. The README does not say whether such a setting exists, so if you are deploying this on a locked-down network, that is the first thing to look for in the settings rather than assume.
The number handling is a good illustration of what local processing buys. The calculator example in the README is deliberately awkward: what is four thousand and two times three minus a million divided by three hundred. That is not a calculator skill, it is a natural-language number parser and formatter, and it is one of four separate repositories the project maintains, described in the contributing section as the number parser and formatter.
OpenWakeWord, importing a .tflite model, and training your own keyword
Wake word detection uses OpenWakeWord, and the default keyword is Hey Dicio. The interesting part is how far the customisation goes, because there are three levels and the deepest one is unusual for an app.
The first is downloading a model. Other `.tflite` models are available from the OpenWakeWord project releases, and a community collection of home assistant wake words is linked as a source. The second is importing it, through Settings, then Input and output methods, then Import custom wake word, which selects the downloaded file. So a user who wants a different trigger word does not need to build anything.
The third is training a model for a keyword of your own. The README points at the OpenWakeWord automatic model training notebook and at a configuration file that is in this repository, under `meta`. Having a training configuration committed alongside the app is the detail that makes this level realistic: the app is not just a consumer of wake word models, it carries the recipe for producing one.
There is a practical caveat that the README does not spell out, and it applies to all three levels. A custom keyword changes what the assistant responds to and nothing else, so a model that triggers too readily turns a conversation in a café into a series of skill invocations. The training notebook is where you would set sensitivity, and nothing in Dicio's own documentation suggests how to tune it for a specific device's microphone.
What is not supported is the other direction. There is no way to run two wake words at once, and no always-listening mode with a different trigger model per profile. For a phone used in one environment that is a reasonable limitation; for a user who wants the assistant in the car and silent at home, it is a wall.
The two skills that read your device, and why they arrived in version 4.1
Two of the fifteen skills deserve a closer look because they are the ones with a privacy cost that is not obvious from the name.
The notifications skill reads all notifications currently in the status bar, and the example question is what are my notifications. On Android that requires the notification listener access, which is a permission that grants an app the text of every notification any other app posts, including messages from apps with no relationship to the assistant. The flashlight skill is harmless by comparison, toggling the phone flashlight on a spoken request.
The release history explains when these arrived. Version 4.1 is titled Notifications, flashlight, Turkish and was published on 2026-05-12. Before that, 4.0 in October 2025 added the translation and joke skills, and 3.2 in July 2025 added Swedish and wake word control. So the two most device-invasive skills are also the newest, added together with a third language, which is the shape of a release that bundles a batch of contributions rather than a carefully staged privacy review.
That is a criticism of the sequencing, not of the feature. Being able to ask what your notifications say by voice is genuinely useful, and there are reasonable ways to build it safely. What the project does not appear to do, from the README, is offer a scoped version, so a user who wants a timer but not their notification contents has no middle setting.
The telephone skill sits in between. It views and calls contacts, example given as call Tom, which needs contact access and can place a call without further confirmation. Whether it confirms before dialling is not stated in the README, and for a skill that initiates a phone call, that is a question worth answering in the source before you enable it.
The sentence compiler and Unicode CLDR plugins that make 14 languages tractable
The top-level directory listing explains how a fourteen-language assistant stays maintainable, and two entries are doing the heavy lifting.
`unicode-cldr-plugin` is a Gradle plugin for the Unicode CLDR data set. That is the reference data that describes how numbers, dates, currencies and measurement units are written in each locale, and it is exactly what an assistant needs in order to say four thousand and two in a way that sounds native in Polish and to format a timer duration in Turkish. Hand-maintaining that per language is hopeless; generating it from CLDR is the standard approach, and having it as a build plugin means the data is refreshed on a normal dependency cycle rather than by hand.
`sentences-compiler-plugin` is the other one, and it explains the architecture. A voice assistant has to know what can be said, and a natural way to build that is to declare, per language, the sentence patterns a skill can match and the answers it can return. The contributing section says the compiler for the sentences language files lives in a separate repository, and the plugin in this tree wraps that compiler as part of the Gradle build. So the language definitions are compiled during the build, and the assistant's grammar is data rather than hand-written Kotlin.
The third piece is the evaluation repository, described as the code for evaluating matching algorithms. That is the piece that decides whether a spoken phrase counts as a request, and separating it means the matching behaviour can be measured rather than assumed.
Taken together, the multi-repository split is not accidental fragmentation. Sentence definitions, number parsing, matching evaluation and the app are four things that change for different reasons and at different rates, and the app is the only one of the four that is an Android application. A contributor adding a language touches the sentence files, not the Kotlin.
Dicio against a cloud assistant and against a self-hosted voice stack
Two comparisons, and the first one is about where the audio goes.
A commercial phone assistant keeps the recognition and the response on a server. That buys accuracy, because a large model runs on hardware you do not have, and it buys breadth, because a server can call any API on the internet and act on the result. What it costs is that the recording leaves the device, and the terms under which it is retained and who may hear it are not something you can inspect. Dicio's design is the direct opposite on that axis: the microphone audio is processed locally, the grammar is a set of compiled sentence patterns rather than a general model, and the assistant can only do what someone wrote a sentence pattern for.
That last point is the real trade. A general model assistant will attempt anything you ask. Dicio will refuse anything it has no skill for, which means it is predictable, auditable and cheap to run, and it means a request it does not recognise simply fails. For a set of fifteen well-defined tasks, that is the right trade. For open-ended conversation, there is no version of this design that competes.
The self-hosted alternative is a different shape again. Assembling your own pipeline, a wake word detector, a speech engine, an intent router and a set of home automation skills, gives you control over every stage and lets you cover home devices, which Dicio's skill list does not. The cost is that you maintain it, and that the parts are harder to combine than a single app.
Where Dicio sits between them is worth stating plainly: it is the choice for someone who wants the microphone to stay on the phone and accepts a fixed, inspectable set of abilities. It is not a general assistant, and no release in its history suggests it is trying to become one.
GPL-3.0, a PRIVACY.md, and an organisation rename the links missed
The licence is the GNU General Public License version 3, with the text in the `LICENSE` file. That permits redistribution and modification, including commercially, on the condition that derived work carries the same licence and source. For an application distributed through F-Droid and the Play Store, that is the standard arrangement and it is also why the source is genuinely open rather than open on request.
The repository contains a `PRIVACY.md`, which a project making on-device claims should have and many do not. Its contents are not reproduced in the README, so it is the file to read before trusting the privacy positioning, along with the network-backed skills identified earlier.
There is a packaging story worth noting. Dicio is distributed three ways, through F-Droid, through GitHub releases, and through the Play Store, all under the application id `org.stypox.dicio`. A `fastlane` directory with per-locale metadata and phone screenshots, a `fastlane_deploy.sh` and a `Gemfile` at the top of a Kotlin project mean the store listing is version-controlled and released by script, which is how a small volunteer project keeps three distribution channels consistent.
The one inconsistency is naming, and it is worth being concrete about. This repository is `DicioTeam/dicio-android`, while the GitHub releases link in the README points at `Stypox/dicio-android`, the contributing section points at three more repositories under the `Stypox` account, the documentation lives at a domain under that name, and the application id is `org.stypox.dicio`. The project has moved to a team organisation and the identifiers have not all followed. Nothing is broken, but a contributor who follows the links from the README lands in the old organisation, and anyone automating a build should use the current repository rather than the one the badges point to.
On maintenance, the release titles are the most useful signal. They name capabilities rather than fixes, three releases in fifteen months, with the newest adding the two device-reading skills and a language. The project is small, volunteer-run, and moving at the pace of contributions, which is a fair description of what to expect from it.
Editorial conclusion
Adopt Dicio if you want a voice assistant whose microphone audio never leaves the phone and you are willing to supply the wake word model, accept an automatic model download from a third-party host, and check which of the network-backed skills you are happy to have making requests. Do not adopt it expecting every skill to be offline, since search, weather, lyrics and translation all call external services, and do not expect smart home control, which the skill list does not cover. Verify first by reading PRIVACY.md in the repository, switching the app language to confirm the Vosk download works on your network, and reviewing the notification skill before granting it status bar access.
Frequently asked questions
What is Dicio?
Dicio is a free and open source voice assistant for Android that supports many skills and input and output methods and gives both speech and graphical feedback to a question. It interprets input and, when possible, generates output entirely on-device, which the README presents as privacy by design.
Which languages does Dicio support?
The README lists Czech, Dutch, English, French, German, Greek, Italian, Polish, Russian, Slovenian, Spanish, Swedish, Turkish and Ukrainian. Instructions for translating Dicio to a new language are published in the project documentation.
What does Dicio use for speech recognition, and how big is the model?
It uses Vosk, with small models weighing around 50MB so it can run on every phone. The download starts automatically whenever needed from the model list at alphacephei, which is what allows the app language to be changed seamlessly.
How do I change the Dicio wake word?
OpenWakeWord listens for Hey Dicio by default. You can download other .tflite models and import them through Settings, then Input and output methods, then Import custom wake word, or train a model for a keyword of your own using the OpenWakeWord training notebook with the configuration in the repository's meta directory.
Does every Dicio skill work without a network connection?
No. The README says interpretation and output happen on-device when possible, but the search skill looks up DuckDuckGo, weather collects from OpenWeatherMap, lyrics come from Genius, and translation uses Lingva. The remaining skills, including the calculator, timers, contacts, flashlight and app opening, act on the device.
Where can I install Dicio, and what is it licensed under?
It is on F-Droid, on GitHub releases and on the Play Store, all under the application id org.stypox.dicio. The code is licensed under the GNU General Public License version 3, and the repository also contains a PRIVACY.md file describing its data handling.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dicioteam-dicio-android)