# ClickUi is one Python file with global-scope Whisper and Kokoro, and the settings you want live in the source

> A PySide6 desktop assistant for voice and chat with local or paid models, with SONOS output and browser driving, structured as a single clickui.py plus sonos.py. Whisper and Kokoro load into the global scope or the program raises, model selection is a Python edit rather than a settings change, and the Windows installer has no counterpart on macOS or Linux.

**CodeUpdaterBot/ClickUi** — The best way to use AI is on your own computer. Use local or paid API models, and ctrl+k to show/hide the chat UI. Experience the future of AI, and help build it too!

- Repository: https://github.com/CodeUpdaterBot/ClickUi
- Website: https://clickui.app
- Stars: 427 · Forks: 41
- Language: Python
- License: AGPL-3.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/codeupdaterbot-clickui

## The assistant is one file, so the model list is a source edit

The repository root lists twenty entries, and the application logic is two of them: clickui.py and sonos.py. Everything else is an asset, a dependency file or a document. favicon.ico, google_icon.png, redfin_icon.png, zillow_icon.png and voice_icon.png sit alongside four wav files for feedback and interface states, including recording_started.wav, recording_ended.wav and loading.wav, and example_audio_response.wav as a sample.

That shape has a direct consequence for configuration. Model and engine information is not read from a config file at startup; the open item on adding models to settings says a model has to be defined in Python right now to make available. Choosing between Ollama, OpenAI and Gemini, which the pinned dependency set carries as ollama, openai and google-genai, means opening clickui.py and editing it.

The same pattern shows up in the settings roadmap. Temperature, top P, repeat penalty, per-model max_tokens for Claude and thinking token effort are all listed as settings that should exist in a SettingsWidget, and none of them do yet. Voice name and model size selection for Kokoro is on the same list. So the widget exists and covers a subset, while the parameters most people reach for first are still Python edits.

For a project whose stated goal is to be readable by anyone who wants to build AI assistance from the ground up, that is a defensible trade. For anyone expecting an installer-plus-settings experience, it is the main gap.

## Whisper and Kokoro load globally, and the documented workaround is editing source

Two speech models are called critical dependencies that must be installed and loaded into the global scope before the program can run. Whisper does automatic speech recognition for the input, Kokoro does text to speech for the output in Voice Mode, and API keys for the model itself are configured alongside them.

The failure mode is stated without hedging. If they are not installed and properly configured, the result is runtime errors. There is no lazy import and no voice-mode-off flag, because the import happens at module scope.

The escape hatch the documentation offers is source modification: run without voice functionality and its dependencies by commenting the Whisper and Kokoro loading out, with the explicit consequence that voice mode will not work. So the two supported states are a fully working voice path with all speech dependencies present, or a text-only run achieved by editing the file. There is no configuration value between them.

A .voiceconfig file exists at the repository root, which suggests a place for voice settings, but the open items place Kokoro voice name and model size selection in the settings widget instead, and the interruption feature during Kokoro output is also unchecked. Whatever .voiceconfig holds, the parts a user would reach for are still being designed.

## requirements.txt freezes a hundred packages, including three browser automation stacks

The dependency file is not a loose list. Every entry is pinned with `==`, including the transitive ones: blis at 1.2.0, numba at 0.61.0, llvmlite at 0.44.0, h11 at 0.14.0, catalogue at 2.0.10. That is the shape of a fully resolved environment rather than a requirements file, which means a clean machine gets exactly the versions this was resolved against and nothing floats.

Three observations about what is inside it.

First, two separate browser automation stacks are pinned at once, playwright 1.50.0 and selenium 4.29.0, alongside beautifulsoup4 4.13.3 for the scraping path. A third is planned: the roadmap has WebUI Browser-Use functionality, with a note that the upstream browser-use web UI is too slow for real-time usage and a mini version might be needed instead. Three overlapping browser toolchains in one assistant, two of them already frozen in.

Second, the pins come from different years. openai-whisper is pinned as 20240930, a date version from September 2024, while PySide6 sits at 6.8.2.1 alongside scipy 1.15.2 and pydantic 2.10.6. A speech model frozen in 2024 running against a 2025 Qt stack is a combination that has to be re-resolved the moment anything moves.

Third, the Kokoro chain pulls phonemizer-fork 3.3.2 rather than upstream phonemizer, plus espeakng-loader 0.2.4 and curated-tokenizers 0.0.9. A forked phonemizer and a native espeak-ng binary are the two entries most likely to fail on a machine that was not the one this was resolved on.

## Install.bat is Windows only, and macOS gets a text file

The documented run command is a single line:

```python
python clickui.py
```

The scripted install path is not symmetric across platforms. The easy installation section is labelled Windows Only, and it is two steps. Install Anaconda or Conda, system-wide for all users, added to the PATH. Then run Install.bat, which is described as available from ClickUi.app and in the GitHub repository, and which downloads this git repository, runs the installation commands, and starts the program.

So the one-command setup belongs to Windows only, and it is a batch file. A Conda environment is not optional in that path, which means the Windows route brings its own Python and its own package manager.

What macOS gets is Mac_Prerequisites.txt, a text file, plus a separate requirements_mac.txt alongside the main requirements.txt and a conda_packages.txt. No shell installer appears at the root. A Linux user, or a macOS user, is left to create an environment, reconcile two or three dependency files, and run the script by hand, and the point where that reconciliation fails is during the speech import described earlier.

The version state matches the packaging. There are no GitHub releases, so there is no tag to pin a macOS or Linux environment against, and the last push to main was 2026-09-07.

## Voice mode documents its own crash on exit, and you cannot interrupt playback

The roadmap carries a specific failure rather than a vague wish. Exiting voice mode during transcription or audio playback sometimes quits the entire program, which is listed as an item to fix or revise along with the related start and stop toggling and resets. That is the single most damaging entry in the list, because it means the mode you cannot fall back to is also the mode that can take down a session mid-answer.

Two related gaps sit next to it. There is no interrupt capability during Kokoro output, so there is no barge-in: to speak over a spoken response you have to wait for it to finish or kill it. And because voice name and model size selection is not in the settings widget yet, the choice of which Kokoro voice reads your answers is made outside the interface.

Taken together, voice mode is the feature the project's own demo leads with, and it is the part with the least finished state management. The completed item in the list is smaller than any of these: replacing the keyboard library with pynput for cross-platform keyboard and mouse support, attributed to a contributor, because the previous library had issues on Mac and Linux. pynput 1.8.1 is in the pinned set.

For a text-only user the roadmap is honest about what voice mode is not yet. For anyone building on the voice path, the crash on exit is the item to test before anything else.

## Two windows, one attachment, no tests, and no cost tracking

The interface gaps are listed as precisely as the missing settings, which makes the current shape easy to reconstruct.

Two windows launch in the taskbar, one for each area, and merging them into one main window is an open item. After launching with the hotkey the cursor lands in the prompt area, and there is no arrow-key navigation: left arrow to open settings, down arrow to pop open the conversation window below. Navigation is a mouse operation today. ctrl+k is the documented way to show and hide the chat UI.

Attachments are limited to one file, with multi-file attachment listed as a capability to add. Code replies get no formatting and no per-block copy icon, which for an assistant meant to write code is the gap most likely to be felt on day one. Cost visibility is absent twice over: no model pricing table to total input, output, web search and file upload per message, and no per-message token tracking.

There is an open item to build tests for all functionality, naming prompt input chat, reply input chat and conversation history validation as the cases. With no tests in the repository and one source file carrying the logic, a refactor of clickui.py has nothing to catch regressions.

The tree also shows an unversioned state file, .voiceconfig, checked in at the root rather than in a per-user location, alongside .gitignore.

## AGPL-3.0, no releases, and features tracked as README checkboxes

The license is AGPL-3.0, stored as LICENSE.txt at the repository root. For anyone planning to embed this assistant in a product they distribute, that copyleft choice is the first thing to settle, and it is the part of the project with the least ambiguity.

The project governance is the unusual part. Contributors are asked to submit new features and ideas in the README as checkboxes so they can be added to the page, and pull requests go to main and are reviewed. The feature list is therefore both the roadmap and the contribution surface, and it is why the limitations above are stated in the first person by the maintainers rather than inferred by a reader.

The top of the page also asks for collaborators directly: leave Voice mode running, have conversations throughout the day, and build it out. That is an invitation to shape the project by use rather than a maintenance commitment, and it is consistent with a repository that has no GitHub releases, no tag to install, and a single source file that any contributor can read in an afternoon.

One claim on that page deserves measuring rather than repeating. The header calls ClickUi the starting ground for the most widely used computer-based AI assistant, something most people will have installed. Nothing in the repository substantiates a popularity claim, so treat the positioning as intent and judge the tool on its two dependency modes and its pinned environment.

## Conclusion

Use ClickUi if you want a desktop assistant you can read end to end in one file and are willing to run Whisper and Kokoro unconditionally, since the escape hatch is commenting out source lines. Do not plan around its installer on macOS or Linux, because Install.bat is Windows-only and a Mac prerequisites text file is all the other platforms get. Before anything else, check whether the model you want is reachable from the settings widget, because adding one currently means editing Python.

## FAQ

### What are ClickUi's critical dependencies and what happens without them?

Whisper for speech recognition and Kokoro for text to speech, plus configured API keys and engine information for the model. Whisper and Kokoro are loaded into the global scope, so if they are not installed and configured before running, the result is runtime errors. The documented workaround is to comment out the Whisper and Kokoro loading, which leaves voice mode non-functional.

### How do I add an AI model to ClickUi?

Currently by defining it in Python. The open item to add models through the SettingsWidget notes that a model has to be defined in Python right now to make available. The pinned dependencies cover Ollama, OpenAI and google-genai, so the local and paid API options exist in code rather than in configuration.

### Does ClickUi have an installer for macOS or Linux?

The scripted installation is labelled Windows Only and uses Install.bat together with a system-wide Anaconda or Conda install. macOS has Mac_Prerequisites.txt and a separate requirements_mac.txt, and there is no shell installer at the repository root, so those platforms are set up by hand.

### What is the known voice mode bug in ClickUi?

The roadmap lists an item to fix voice mode start and stop toggling and related resets, noting that clicking to exit voice mode during transcription or audio playback sometimes quits the entire program. Interrupt capability during Kokoro output is also still an open item, so there is no way to speak over a spoken response.

### What license does ClickUi use and are there any releases?

The license is AGPL-3.0, stored as LICENSE.txt at the repository root, and the repository has no GitHub releases. The last push to main was 2026-09-07, so there is no tag to pin an environment against.

## Sources

- [CodeUpdaterBot/ClickUi on GitHub](https://github.com/CodeUpdaterBot/ClickUi)
- [Issues](https://github.com/CodeUpdaterBot/ClickUi/issues)
- [License: AGPL-3.0](https://github.com/CodeUpdaterBot/ClickUi/blob/main/LICENSE)
- [Project website](https://clickui.app)
- [README](https://github.com/CodeUpdaterBot/ClickUi/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/codeupdaterbot-clickui
