Artemis runs a local LLM, voice and image stack behind one daemon proxy
破限本地AI女友后宫,openclaw/claude code+画图语音向量数据库+live2D+桌宠+酒馆角色卡导入+前端,QQ+Telegram双通道,8G显存可跑🩵uncensored Fully offline AI girlfriends harem Openclaw/Claude code+Local LLM+GPT-SoVITS+ComfyUI image+Live2D+desktop pet+SilllyTavern Character card import+frontend | Dual channels for QQ & Telegram | Dynamic 8G VRAM scheduling+mem0 qdrant, can run offline
At a glance
- What is it?
- A Windows-first assistant that pairs llama.cpp with GPT-SoVITS voice and ComfyUI images, reachable from QQ, Telegram, a Live2D desktop pet and a browser chat on port 19270, with one character memory per persona. Everything is local, and the setup is a chain of PowerShell scripts rather than a package.
- Who is it for?
- Adopt Artemis if you want a character assistant that runs with the network off, you are on Windows with an NVIDIA card, and you are willing to sequence four setup scripts and download models first. Leave it if your GPU is not an NVIDIA one supported by the default scripts, since AMD users are routed to a separate folder, or if you want a hosted service with someone else's maintenance.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One daemon proxy sits between every channel and llama.cpp
The component list is long, and the reason it is not a mess is the daemon in the middle. OpenClaw drives the agent side, QQ Bot and Telegram Bot are two channels into it, and the generation stack behind it is llama.cpp for text, GPT-SoVITS for voice and ComfyUI for images, with a Sakura desktop pet and Live2D for the on-screen character.
The routing is documented in one line. Real LLM chat streams replies through the daemon's /api/chat endpoint to llama.cpp's /v1/chat/completions, with no fake fallbacks, which is a claim worth taking seriously: it means a failure surfaces as an error rather than as a canned reply pretending to be the model.
The model selector in Settings offers local llama, DeepSeek and Grok, all routed through the daemon proxy. That is the one place where a remote provider can be involved, and it is a choice in a dropdown rather than the default path.
Four channels reach the same agent, which is the architectural point. QQ, Telegram, the Live2D pet and the browser chat are not separate clients with separate logic; they are four ways to talk to one process, and per-character memory has to survive all four.
The web frontend is the easiest one to try because it needs no bot account: a browser chat at http://127.0.0.1:19270 that connects directly to the local daemon proxy and then to the llama.cpp server.
quick_setup.ps1 is step zero and it writes the config.yaml every later script reads
The prerequisite section starts with a wizard rather than a package install, which tells you what kind of project this is.
powershell -ExecutionPolicy Bypass -File quick_setup.ps1Four things happen inside it. You choose the default agent language, Chinese, Japanese or English, and the wizard copies the matching AGENTS_ file to DEFAULT_AGENT.md. It auto-detects installed tools: ComfyUI, GPT-SoVITS, llama.cpp and embedding models. It prompts for any path it cannot find. And it generates config.yaml with all the paths, already in the shape download-models.ps1 expects.
That last point is the reason the wizard is not optional. Every later step reads config.yaml rather than asking again, so a wrong path here becomes a wrong path everywhere.
There is a second option for the dependency stage, which clones the two external inference frameworks for you.
powershell -ExecutionPolicy Bypass -File setup-deps.ps1and the same thing on Linux or macOS.
bash setup-dependencies.shEither script detects the OS and GPU environment, clones ComfyUI from its upstream repository and GPT-SoVITS from its upstream repository, then writes the new paths back into config.yaml. Neither pip nor a package manager is used for those two, because they are external inference frameworks that carry their own environments.
After that, the order is fixed: download-models.ps1, then setup-llama.ps1, then start.ps1. There is no single all-in-one step in the documented path.
The ComfyUI workshop wants 12GB of VRAM while the llama path is documented at 8GB
The hardware table names a specific machine: an NVIDIA GeForce RTX 5070 Laptop with 8GB of VRAM, and an Intel Core i9-14900HX with 24 cores. The web chat section repeats the 8GB figure, saying 8GB of VRAM can run fully without stopping.
The constraint is the reason this project has a scheduling story at all. Text generation runs through llama.cpp, image generation runs through ComfyUI, and the two compete for the same card. The ComfyUI workshop console is documented as needing 12GB or more, so it is the piece that pushes past the 8GB configuration.
That is why the release notes and repository layout contain tuning artefacts rather than just code. LLAMA_TUNING.md, MODEL_PATHS.md and OPTIMIZATION_LOG.md sit at the top level next to artemis_headroom_proxy.py, which is named for headroom management rather than for anything user-facing.
Two more hardware facts are worth having in front of you. The scripts are written for NVIDIA, and the README says AMD users should look in the AMD_GPU folder instead. And the requirements file states that the torch version has to match both the CUDA driver and the GPU architecture, with the specific note that an RTX 5070 Blackwell card needs torch 2.7 or newer with the cu130 build.
If your card is neither of those, the honest answer from this documentation is that the default path does not target you.
Character cards are dropped in, and each character keeps its own memory
Characters are data, not code, and there are three ways to get one.
Dynamic characters auto-load from skills/harem/, and the studio displays persona, tags and a greeting per character, so adding a character means adding a directory rather than editing the application. Card import handles the other route: drag-drop or select SillyTavern PNG or JSON character cards and the metadata and persona are parsed automatically, which is how people bring characters they already have.
Hot-swapping is the feature that makes more than one character practical. You switch from a sidebar dropdown, and memories and chat context are preserved per character. Isolated memory per character is stated as the design, so switching does not leak one persona's history into another's conversation.
What backs that memory is worth being precise about. requirements.txt lists the mem0 bridge for vector memory with qdrant-client and sentence-transformers, and both lines are commented out with a note that memory search degrades if they are not installed. So the vector layer is optional by construction, while the chat history, settings and character state are persisted in browser localStorage.
That split defines the ceiling. Conversations and character state live in one browser profile on one machine, and memory search is the feature that disappears first if you skip the optional install.
PowerShell and shell scripts come in pairs, and two .cmd files close the loop on Windows
The repository root is mostly scripts, and the pattern is consistent enough to read at a glance. Every setup stage has a PowerShell and a shell variant: setup-all, setup-openclaw, setup-dependencies, download-models and setup-llama all appear as a .ps1 and a .sh pair. download-llama.ps1 and claude-code.ps1 and claude-code.sh sit alongside the same idea, and claude-code-ccr.ps1 and claude-code-ccr.sh are the second pair of those.
The Windows-specific parts are not optional extras. artemis_studio.ps1 and artemis_studio.py are both present because the studio launcher is a PowerShell script driving a Python process. shiki-start.cmd and shiki-stop.cmd are the manual start and stop pair, and restart_llama_rea.bat is a batch file for restarting the llama process on Windows.
The Python side is small, which is the other structural fact. shiki_daemon.py is the process everything talks to, artemis_bridge.py is the bridge the Live2D pet is controlled through, artemis_headroom_proxy.py handles VRAM headroom, and artemis_studio.py backs the studio. The requirements list is numpy, pyyaml, Pillow, scipy, pystray for the system tray, and flask for the embedding server. Project modules under skills/ need no pip install because PYTHONPATH covers them.
The Node side is nearly empty: package.json declares two dependencies, @phosphor-icons/web and playwright.
The documentation set is unusually candid for a project of this kind, with USAGE.md, TOOLS.md, IDENTITY.md, SOUL.md, USER.md and HEARTBEAT.md at the root alongside docs/, memory/, live2d/ and media/.
Three language READMEs, a LICENSE file at the root, and releases every week or two
The user-facing surface is translated in full: README.md in English, README_CN.md and README_jp.md, with the language choice made in the setup wizard and applied by copying the matching AGENTS_ file rather than by a runtime switch.
Licensing needs one sentence of its own. A LICENSE file sits at the repository root, and no licence identifier is attached to the project in the repository record, so read that file before you redistribute anything built on it.
The maintenance picture is one of frequent releases. Versions 1.8.6, 1.8.7 and 1.8.8 shipped on 2026-09-10, 2026-09-19 and 2026-09-30 respectively, and the last push to the repository was on 2026-09-26. Release notes are absent from this repository view, which means the version numbers tell you how often it moves but not what changed, so read the release page before upgrading a working setup.
One detail about the documentation itself: the hardware table is cut off partway through the CPU line in what is available here, and the video walkthrough is referenced by its identifier rather than explained, so the hardware requirements beyond GPU and CPU are not written down in the part I could read. Anyone buying hardware for this should read the repository directly before deciding.
Editorial conclusion
Adopt Artemis if you want a character assistant that runs with the network off, you are on Windows with an NVIDIA card, and you are willing to sequence four setup scripts and download models first. Leave it if your GPU is not an NVIDIA one supported by the default scripts, since AMD users are routed to a separate folder, or if you want a hosted service with someone else's maintenance. Verify three things first: that config.yaml was generated by quick_setup.ps1 with real paths to ComfyUI and GPT-SoVITS, because every later script reads it, that you have the VRAM for what you plan to run, since the ComfyUI workshop wants 12GB or more while the llama path is documented at 8GB, and that you read the LICENSE file at the root, since no licence identifier is attached to the project here. Releases 1.8.6, 1.8.7 and 1.8.8 shipped in September 2026, and the last push was on 2026-09-26.
Frequently asked questions
What does Artemis actually run on my machine?
llama.cpp for text, GPT-SoVITS for voice and ComfyUI for images, driven by a local daemon proxy, plus Live2D and a desktop pet for the character. Channels include QQ, Telegram and a browser chat at http://127.0.0.1:19270, and the project states that no audio, images or chat data leaves the machine.
How do I set up Artemis from scratch?
Run quick_setup.ps1 first, which picks the agent language, detects ComfyUI, GPT-SoVITS, llama.cpp and embedding models, and writes config.yaml. Then clone the external frameworks with setup-deps.ps1 or setup-dependencies.sh, and finish with download-models.ps1, setup-llama.ps1 and start.ps1.
What hardware does Artemis need and does it work on AMD?
The documented machine is an NVIDIA GeForce RTX 5070 Laptop with 8GB of VRAM, and the ComfyUI workshop console needs 12GB or more. The default scripts are written for NVIDIA, and the README points AMD users to the AMD_GPU folder. Torch has to match the CUDA driver and GPU architecture, with 2.7 or newer plus cu130 for a Blackwell card.
How do I add my own character to Artemis?
Either drop a SillyTavern PNG or JSON character card into the interface, which is parsed for metadata and persona, or add a directory under skills/harem/ that the studio auto-loads with its persona, tags and greeting. Each character keeps its own memories and chat context when you hot-swap.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/momori777-artemis)