AI Jarvis (LYiHub/pub-local-jarvis): a Windows desktop pet that watches your screen and listens to system audio
Windows 本地多模态 AI 桌面桌宠,支持屏幕与音频感知。
At a glance
- What is it?
- AI Jarvis is a local-first multimodal desktop companion for 64-bit Windows 10/11, built on MiniCPM-o 4.5 GGUF and shipped as an NSIS installer. The interesting part is the full-duplex loop; the boring part is that it is Windows-only and text-only on output.
- Who is it for?
- Adopt AI Jarvis if you run 64-bit Windows 10/11, have 12 GiB of disk to spare, and want a screen-and-audio-aware companion whose inference stays on the machine. Do not adopt it if you need macOS or Linux, voice output, or a model that acts on the desktop; the README states the current version is text-oriented and the game companion only observes and hints.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem AI Jarvis targets: context you have to type out by hand
Most local assistants wait for a prompt. You describe what is on screen, paste an error, or summarize a lecture from memory. AI Jarvis inverts that. According to the README, it continuously ingests the desktop picture and the system's playing audio, then decides on its own whether to stay quiet or say something. The two outputs are LISTEN and SPEAK, and the model picks between them roughly once per second.
The audience is narrow and specific. The environment requirements are 64-bit Windows 10/11, an x64 CPU with AVX2, at least 12 GiB of free disk, and a network connection on first launch to pull about 6.32 GiB of model weights. macOS and Linux are not mentioned anywhere in the README, and the desktop shell is Electron with a C++20 native layer, so this is not a portable Python script you can run on a server. If you want a local multimodal assistant on a Mac, this is the wrong repository.
The three named use cases are desktop companionship, game companionship through a transparent overlay, and lecture note-taking. The last one is the most concrete: course mode saves selected keyframes, organized knowledge points and frame descriptions, then emits a Markdown summary. That is a real artifact you can keep, unlike the ambient commentary.
How the full-duplex loop and the two isolated contexts fit together
The README describes the model as MiniCPM-o 4.5 in GGUF format, combining a language model, a vision model (VPM) and an audio model (APM). Capture happens in the native layer: DXGI for the screen and WASAPI for system audio. The README's diagram shows these two streams splitting into separate consumers:
DXGI 屏幕画面 + WASAPI 系统音频
|
+-- 结构化感知上下文 -> 场景 / 游戏 / 课程 / 记忆
|
+-- 全双工上下文 -> LISTEN / SPEAK -> 模型回复That isolation is the design decision worth noting. Structured scene classification (is this a game, a lecture, neither) runs in one context, while the conversational LISTEN/SPEAK decision runs in another. Keeping them apart means a wrong scene label does not contaminate the dialogue history, and the dialogue does not have to re-derive scene state on every turn. It also means two prompt budgets instead of one, which matters when the model is quantized and running partly on CPU.
Inference itself comes from tc-mb/llama.cpp-omni, a fork of llama.cpp/ggml. The README states the project pins that upstream version and connects the LLM, VPM and APM with voice output disabled. That pin is a deliberate trade: reproducible builds and a known-good MiniCPM-o integration, in exchange for manual work whenever upstream changes. The pinned fork is also the reason audio goes in but never comes out.
Installing AI Jarvis on Windows and opening the chat window
The README recommends the installer over a source build. Download AI-Jarvis-Setup-0.1.2-x64.exe from the releases page, run it, then launch AI Jarvis and click the button labeled 启动 AI 贾维斯. The installer bundles the Python backend and a compiled C++ inference runtime, so Python, Git, CMake, Visual Studio and CUDA are not prerequisites for the end user. The model weights are not bundled; first launch downloads, resumes and verifies MiniCPM-o 4.5 automatically.
If you would rather run from source, the README lists Python 3.12+, Git, CMake 3.24+, Visual Studio C++ Build Tools, and a current Node.js LTS with npm. Building the Windows installer additionally needs CUDA Toolkit 13.1 or newer, though the CUDA runtime is copied into the package so end users never install the Toolkit. From the repository root:
cd desktop
npm run deps:install
cd ..
.\start-real.cmdAfter the desktop window opens, click 启动 AI 贾维斯. Once the service reports ready you can drag the pet to reposition it. Ctrl+M opens or closes the chat dialog, and the README says replies take the live screen into account. Double-clicking the pet toggles privacy mode, which pauses screen and audio perception so the model can no longer obtain either.
For NVIDIA users, the README suggests confirming the driver before anything else:
nvidia-smiIf that prints a card name, driver version and memory size, the driver is working. On a successful start the runtime log shows NVIDIA CUDA 加速已启用. If CUDA fails to initialize the program falls back to CPU and says so through the pet bubble; the README warns that first-frame perception and text generation can then slow from tens of seconds to minutes. RTX 50 series cards are advised to use CUDA 13.x with a current driver.
Where AI Jarvis breaks down: CPU fallback, text-only output, and no control
The most likely failure is not a crash. It is a quiet downgrade. CUDA initialization can fail on an otherwise fine machine, and the program responds by moving to CPU and showing a bubble. The README's troubleshooting table maps symptoms to causes: CUDA initialization failed or insufficient video memory means close the memory-hungry programs and restart; missing cudart or cublas DLLs mean update the NVIDIA driver and possibly install CUDA Toolkit 13.x; missing VCRUNTIME140.dll or MSVCP140.dll means install the Visual C++ 2015-2022 x64 redistributable. A model verification failure is handled by exiting, deleting the incomplete model directory and re-downloading.
There is a subtler trap in confirming acceleration. The README explicitly tells you not to trust the default 3D graph in Task Manager. You should look at the CUDA compute graph and dedicated GPU memory, or run nvidia-smi while the program is working. If CPU usage climbs while GPU memory stays flat and the pet shows a CPU notice, you are on the fallback path.
The capability boundary is just as important. The README states the current version is text-oriented and does not provide voice playback or real-time voice conversation, so the audio model listens but the assistant never speaks aloud. Game companionship only observes, hints and interacts through a transparent overlay; it does not control the game or act in your place. And 8 GiB of video memory should not be forced into a full model load, since AI Jarvis decides GPU layer count from currently available memory. If a browser or another AI program is holding memory, the README says to close it first.
How AI Jarvis differs from a scripted screenshot pipeline
A common alternative is a scheduled script that captures a screenshot on a timer and posts it to a vision model API, then writes the answer to a file or a chat window. That approach is portable and easy to debug, and it works on any OS. The difference is architectural, not cosmetic.
A screenshot pipeline has no notion of whether it should speak. It fires on the timer, spends a request, and produces output whether or not anything changed. AI Jarvis keeps a continuous input stream and makes a per-tick LISTEN/SPEAK choice, which is the whole point of the full-duplex framing in the README. It also splits structured scene perception from the dialogue context, so the memory timeline is built from stable perception results rather than from raw frames. The README states that original screen images and audio are not written into long-term memory; only the activity timeline and daily summaries are kept.
The trade runs the other way too. A screenshot script is stateless and cheap to reason about. AI Jarvis holds a native capture layer, a pinned llama.cpp fork, a FastAPI orchestration backend and an Electron shell, all of which must agree before anything works. When it misbehaves, the README points you at %LOCALAPPDATA%\AIJarvis\runtime\native-worker.log, which is a much longer path than reading a cron log.
Privacy posture, licence, and what a version bump costs you
The privacy model is stated plainly: screen and audio go into current inference only, and ordinary operation does not persist the raw capture. Course mode keeps selected keyframes, organized knowledge points and frame descriptions. The one optional network path is the visual schedule image, which is off by default and only contacts a user-configured image generation API when you actively generate one; at that point the day's review and the character reference image are sent. If that distinction matters to your deployment, the config directory and the settings that gate the image API are what to audit, not the README prose.
The source is MIT licensed, which is permissive and imposes no copyleft on your own code. Two caveats sit outside the MIT grant. The repository carries a THIRD_PARTY_NOTICES.md, and the bundled inference path derives from a pinned fork of llama.cpp-omni plus MiniCPM-o weights from OpenBMB. Model weights are downloaded at first launch rather than shipped, so whatever terms attach to MiniCPM-o 4.5 apply to that download separately from the MIT code. This is a description of the repository layout, not legal advice; read the notices file and the upstream model terms yourself.
Upgrade cost is dominated by the pinned fork and the frozen runtime. The pyproject.toml pins the backend to Python 3.12 or newer with fastapi, huggingface-hub, pydantic and uvicorn ranges, and the README notes that the release build freezes the Python backend and compiles a static MSVC/CPU native runtime. Moving to a newer MiniCPM-o or llama.cpp means re-validating that whole chain. There is a verify path for exactly this: npm run verify:installer installs to a temp directory, strips Python and build-tool environment variables, checks the self-contained runtime and uninstalls. Adding -FullStartup also exercises the first model download, the native inference process and backend health.
Editorial conclusion
Adopt AI Jarvis if you run 64-bit Windows 10/11, have 12 GiB of disk to spare, and want a screen-and-audio-aware companion whose inference stays on the machine. Do not adopt it if you need macOS or Linux, voice output, or a model that acts on the desktop; the README states the current version is text-oriented and the game companion only observes and hints. Before trusting it, verify two things yourself: that the native-worker log shows CUDA acceleration rather than a silent CPU fallback, and that the model directory under %LOCALAPPDATA%\AIJarvis\models\MiniCPM-o-4_5-gguf passed its hash check on first launch.
Frequently asked questions
Is ChatGPT basically Jarvis?
No, and AI Jarvis is a narrower thing than the comparison suggests. It is a local multimodal desktop pet that runs MiniCPM-o 4.5 on your own Windows machine and decides on its own when to speak, rather than a hosted chatbot you prompt. The README also states the current version is text-oriented with no voice playback.
What is Jarvis used for?
In this project the README names three uses: desktop companionship with short reminders based on the live screen and system audio, game companionship through a transparent overlay that observes and hints, and lecture recording that produces a Markdown note with keyframes and knowledge points.
What is Jarvis AI used for?
Here it is used as a local full-duplex assistant: the model receives about one frame per second plus the last second of system audio and chooses between LISTEN and SPEAK. Structured scene perception and the dialogue context are kept separate, and the resulting timeline feeds local memory and daily summaries.
Can I make AI like Jarvis?
You can run this one rather than build it from scratch. The README's recommended path is the AI-Jarvis-Setup-0.1.2-x64.exe installer, which bundles the Python backend and the compiled C++ inference runtime; only the model weights are downloaded on first launch. Source builds need Python 3.12+, CMake 3.24+, Visual Studio C++ Build Tools and Node.js.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lyihub-pub-local-jarvis)