# pub-local-jarvis: MiniCPM-o on the desktop, watching the screen and listening to the audio

> A Windows desktop pet that runs MiniCPM-o 4.5 locally over DXGI screen frames and WASAPI system audio, deciding every second whether to stay silent or speak. The installer bundles Python and a CUDA runtime, the model arrives as a 6.32 GiB download, and the current version is text only.

**LYiHub/pub-local-jarvis** — Windows 本地多模态 AI 桌面桌宠，支持屏幕与音频感知。

- Repository: https://github.com/LYiHub/pub-local-jarvis
- Stars: 494 · Forks: 98
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/lyihub-pub-local-jarvis

## One release, and 6.32 GiB of model that does not ship in it

There is a single GitHub release, `v0.1.2`, published on 2026-07-24, and the last push to the default branch carries the same date. The Python package version in `pyproject.toml` also reads 0.1.2, so tag, tree and package agree.

What does not ship is the model. The installer is `AI-Jarvis-Setup-0.1.2-x64.exe`, and the weights arrive on first launch: roughly 6.32 GiB, downloaded with resume support and hash verification, landing in `%LOCALAPPDATA%\AIJarvis\models\MiniCPM-o-4_5-gguf`. The instruction to not force-kill the program during download or verification is repeated in the troubleshooting section, and a corrupted download is resolved by deleting the incomplete model directory and starting over.

The hardware floor is narrow on purpose: 64-bit Windows 10 or 11, an x64 processor with AVX2, and at least 12 GiB of free disk. The installer bundles the Python backend and a compiled C++ inference runtime, so Python, Git, CMake, Visual Studio and CUDA are all absent from the requirements list.

## The installer carries its own CUDA and falls back to CPU without asking first

The Windows build ships an NVIDIA CUDA inference runtime and uses it in preference to anything on the system, then decides how much of the model to offload based on currently available VRAM. That last part is the useful design choice: with 8 GiB of VRAM the model is not loaded in full, and the number of GPU layers is computed at startup rather than hardcoded.

When CUDA initialization fails, or the machine has no compatible NVIDIA card, the program switches to CPU on its own and says so through a pet bubble notice. Functionality survives, and the documentation is specific about the cost: first-frame perception and text generation may take tens of seconds up to several minutes.

Two hardware notes follow from that. RTX 50 series cards are advised to use CUDA 13.x with current drivers. And if a browser, a game or another AI program is holding VRAM, the advice is to close it before retrying rather than to lower a setting. Building the installer at all needs CUDA Toolkit 13.1 or newer, since the CUDA runtime libraries get copied inside the package, which is what lets end users skip the Toolkit entirely.

## Confirming the GPU path means reading the CUDA row, not the 3D curve

Because the fallback is automatic, the only reliable way to know which path you are on is to check. The application waits for a ready indicator before perception begins, and then you have two ways to confirm. In Task Manager, open the GPU performance graph and look at the CUDA compute graph and the dedicated GPU memory figure. Looking only at the default 3D curve will mislead you, since inference does not necessarily register there.

The second way is to run the driver query while the application is running:

```powershell
nvidia-smi
```

The signs of a CPU fallback are all three together: CPU utilization climbing while GPU memory does not clearly increase, plus the pet showing its CPU notice. During setup, the same command doubles as a gate. If it is missing or errors, the instruction is to fix the driver first and explicitly not to continue with Python package installation, since installing packages will not repair a driver problem.

A success signal also exists in the log: the line reporting that NVIDIA CUDA acceleration is enabled. Missing `cudart64_*.dll`, missing `cublas64_*.dll`, or an outright initialization failure point at the CUDA Toolkit 13.x Runtime and Development components. Missing `VCRUNTIME140.dll` or `MSVCP140.dll` points somewhere else entirely, at the Visual C++ 2015-2022 x64 redistributable, followed by a Windows restart.

## Two isolated contexts: structured perception, and LISTEN or SPEAK

The model is MiniCPM-o 4.5 in GGUF form, and three components inside it do the work: a language model, a vision model for the screen, and an audio model for sound. Capture is DXGI for frames and WASAPI for system audio, and both feed two separate contexts that never mix.

One context does structured perception and yields scene, game, course and memory results. The other is the full-duplex context, and it runs a continuous decision about one frame per second plus the most recent second of audio. The model chooses `LISTEN`, meaning keep observing and do not interrupt, or `SPEAK`, meaning emit one short sentence grounded in what is currently on screen or audible. Because the input stream keeps advancing, this is not question and answer; the model can hold off until there is something worth saying.

The architecture text makes one thing explicit: structured scene judgment and full-dup conversation use mutually isolated model contexts. That separation is what keeps a background scene classification from contaminating the conversational thread.

## Voice output is disabled and game mode never touches the controls

Two capability boundaries are stated plainly in the current version. There is no voice broadcast and no live voice conversation. Interaction is text, and the desktop pet opens a dialog beside itself on `Ctrl+M`, where you can put a question to the local model and read the reply while it can still see the live screen.

Game companionship is the second boundary. It reads the scene and responds with prompts and interaction through a transparent overlay chosen not to interfere with your input, and per-game companion configurations are supported. What it does not do is control the game or act in your place. There is no input injection anywhere in the described behavior.

That framing also explains the upstream choice. The inference layer comes from a pinned fork of a llama.cpp and ggml based project providing MiniCPM-o GGUF inference, integrated with the language, vision and audio models and with voice output switched off. The same switch is why the audio model exists but speech does not: it is there to let the model hear, not to let it speak.

## Local memory keeps the timeline and drops the frames

The privacy story is a data-retention boundary rather than a promise. Screen frames and system audio feed the current inference only, and an ordinary run does not keep raw captures long-term. What persists is derived: an activity timeline and a daily summary assembled from stable scene perception results, with the original frames and audio left out of long-term memory.

Course mode is narrower still. It saves only the key frames that were selected, the organized knowledge points, and a description for each frame, then produces a Markdown note with those summaries once the session ends. Mid-course it offers content reminders.

Two controls sit on top. Double-clicking the pet pauses or resumes screen and audio perception entirely, and while paused the model cannot keep obtaining frames or sound, which is the switch to use in front of other people. And the visual schedule image is off by default: it only reaches the network when you configure a compatible image generation API and actively generate one, and only then are that day's review and project role reference images sent to the API you configured.

## Source needs six toolchains, and release builds verify by uninstalling

Running from source reverses the convenience of the installer. It wants Python 3.12 or newer, Git, CMake 3.24 or newer, Visual Studio C++ Build Tools, Node.js on the current LTS line, and npm, plus CUDA Toolkit 13.1 or newer to produce a Windows installer.

```powershell
cd desktop
npm run deps:install
cd ..
.\start-real.cmd
```

Once the desktop app is open you click the start control for the assistant, and after the service is ready the pet can be dragged into position while `Ctrl+M` opens and closes the dialog. The release build path adds one step:

```powershell
cd desktop
npm run deps:install
npm run build
```

The artifact lands at `desktop/dist/AI-Jarvis-Setup-<version>-x64.exe`, and release builds are stated not to carry source, compilers or local run data. What is notable is the verification step. `npm run verify:installer` performs an isolated install into a temporary directory, strips Python, virtual environment and build tool variables from the environment, checks that the runtime is still self-contained, and uninstalls afterwards.

```powershell
npm run verify:installer
```

A heavier variant also exercises the parts an installer test cannot reach:

```powershell
npm run verify:installer -- -FullStartup
```

That one covers the first model download, the native inference process, and backend health.

## The Python packaging declares an Anthropic extra the local-first story ignores

The backend package is `aijarvis-backend`, described as a local-first control plane, and its dependencies are unremarkable: FastAPI, uvicorn, pydantic and pydantic-settings, plus the Hugging Face hub client and `hf-xet` for the model download, and tqdm for progress. A console script named `jarvis-backend` points at the app entry point.

One line does not fit that picture. An optional extra named `llm` pulls in `anthropic`, pinned to the 0.116 range and above. Nothing in the README, which is otherwise emphatic that inference and perception happen on the machine and that data does not need to go to a cloud, says what that client is for. It may be a development aid or a path for a capability the README does not cover, but that is inference on my part rather than something stated, and guessing would be worse than leaving the question open.

Two other details are worth knowing before you build. Pytest is pointed at `tests/unit` with asyncio in auto mode, so the Python tests are unit-scoped. And the linter excludes `third_party/runtime/vendor` explicitly, which tells you the vendored runtime is treated as untouchable code rather than something meant to be cleaned up.

## Conclusion

This is a Windows-only local inference stack with a detailed fallback path, aimed at people who want continuous screen and audio awareness without sending frames anywhere. It needs AVX2, 12 GiB of free disk, a 6.32 GiB model on first launch, and an installer that already carries Python and a CUDA runtime, which is why the exe is the sane starting point rather than the source tree. Two things to check before trusting it. The current version is text only, with no voice output, and game mode observes rather than controls anything. And the Python packaging declares an optional Anthropic client extra that the local-first description never accounts for, so find out what consumes that extra before you rely on the privacy claim.

## FAQ

### What does pub-local-jarvis need to run on Windows?

64-bit Windows 10 or 11, an x64 processor with AVX2 support, and at least 12 GiB of free disk. The first launch needs network access to download roughly 6.32 GiB of MiniCPM-o 4.5 model files, which are not bundled in the installer.

### Does the pub-local-jarvis installer require Python or CUDA to be installed first?

No. It bundles the Python backend and a compiled C++ inference runtime, and carries an NVIDIA CUDA runtime that decides GPU offload from available VRAM. Only building the installer yourself needs CUDA Toolkit 13.1 or newer.

### What happens when pub-local-jarvis cannot use the GPU?

It switches to CPU on its own and shows a notice in the pet bubble. Features still work, but first-frame perception and text generation may take tens of seconds up to several minutes. Task Manager's CUDA compute graph and `nvidia-smi` reveal which path is active.

### Does pub-local-jarvis keep my screen and audio recordings?

Frames and system audio feed the current inference only. Long-term local memory holds an activity timeline and daily summaries derived from stable scene results, with raw frames and audio excluded. Course mode saves only selected key frames, organized knowledge points and frame descriptions.

### Can pub-local-jarvis speak aloud or control games for me?

No. The current version is text interaction only, with no voice broadcast or live voice conversation, and `Ctrl+M` opens the dialog beside the pet. Game companionship observes the scene and offers prompts through a transparent overlay; it does not control the game or act on your behalf.

## Sources

- [Issues](https://github.com/LYiHub/pub-local-jarvis/issues)
- [License: MIT](https://github.com/LYiHub/pub-local-jarvis/blob/main/LICENSE)
- [LYiHub/pub-local-jarvis on GitHub](https://github.com/LYiHub/pub-local-jarvis)
- [README](https://github.com/LYiHub/pub-local-jarvis/blob/main/README.md)
- [Releases](https://github.com/LYiHub/pub-local-jarvis/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lyihub-pub-local-jarvis
