# cyber-doctor (Cyber Huatuo): a Gradio medical agent that wires an LLM to Neo4j and a knowledge base

> cyber-doctor is a Python project that assembles a multimodal health assistant from an OpenAI-compatible LLM, a Neo4j knowledge graph, a file-backed knowledge base and edge-tts voice output. It installs from a git clone and runs on port 7860, but the README leaves deployment, safety and data-handling questions open.

**ZeroTang05/cyber-doctor** — 赛博医生项目——”赛博华佗“，基于多模态大模型的多功能智能体，一键搭建本地多模态大模型。接入医疗健康相关的知识图谱和知识库后可以进行疾病初诊，病历分析，专业知识问答等功能，成为你的私人医生。赛博华佗项目能帮助实现医疗资源的跨地域传播，让更多人借助大模型改善健康水平。"Cyber ​​Huatuo" - Easy to build a personal doctor agent based on LLM and Knowledge Graph/Knowledge Database.

- Repository: https://github.com/ZeroTang05/cyber-doctor
- Stars: 499 · Forks: 53
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zerotang05-cyber-doctor

## The gap cyber-doctor tries to fill

The README frames the problem in terms of uneven access to medical resources: people in less developed regions travel to major cities for care. The project's answer is a configurable assistant that performs basic disease triage, medical record analysis and professional question answering. The stated target user is anyone who cares about their own health, and the README also claims the same architecture can be pointed at any domain by swapping the fine-tuned model and the retrieval sources.

That second claim is the more interesting one. What cyber-doctor actually packages is a wiring diagram: a Gradio front end, an LLM client layer, a retrieval layer that can draw on Neo4j or uploaded files, and a voice layer. The medical knowledge graph is a default, not a hard dependency. Teams evaluating it should read it as an agent scaffold with a medical demo attached, not as a medical device.

## How the agent routes text, images, audio and retrieval

The repository layout shows the split clearly. app.py builds the Gradio interface and handles multimodal input. client/ holds the model layer: LLMclientbase.py defines the base class, LLMclientgeneric.py wraps chat generation, and clientfactory.py builds the right client for the provider. Two provider implementations ship in-tree, client/ourAPI/ for OpenAI-SDK-compatible endpoints and client/zhipuAPI/ for ZhipuAI. The README notes that the team lacked the capacity to test every provider, and asks for issues and pull requests to fix API compatibility problems.

Retrieval is split across three directories. rag/ handles file-backed knowledge, kg/ handles the Neo4j graph, and Internet/ runs the web search chain: Internet_chain.py coordinates keyword extraction, crawling and retrieval, Internet_prompt.py does the prompt-side keyword extraction, and retrieve_Internet.py calls the search API. The .env.example exposes YOUCOM_API_KEY for the You.com search and research API, which is the only web search credential the sample file names.

Audio is a separate path. audio/audio_extract.py extracts the target text and language for speech synthesis, and audio/audio_generate.py wraps edge-tts. Speech recognition uses whisper, listed in requirements.txt alongside openai-whisper. The README says the voice mode supports multiple dialects and defaults to audio output once you enter that mode. Output generation for documents lives in ppt_docx/, which builds plain-text PPT and Word files using python-pptx and python-docx.

## Installing cyber-doctor and running a first query

The README's startup path is four steps: clone, configure the API, fill in the YAML config, create the environment, run app.py. Python 3.10 is the recommended version and the README states python>=3.10 is required. Conda is suggested but not mandatory.

Start by cloning and creating the environment:

```bash
git clone https://github.com/Warma10032/cyber-doctor.git
conda create --name myenv python=3.10
conda activate myenv
pip install -r requirements.txt
```

Next, copy .env.example to .env and fill it in. The sample file defaults to ZhipuAI's endpoint and the glm-4-flash model, and it lists alternative base URLs for DeepSeek, Qwen, SiliconFlow, Doubao and a local Ollama server. The three keys you must set for basic chat are LLM_BASE_URL, LLM_API_KEY and MODEL_NAME. The image, video and multimodal fields are separate and, according to the .env.example comments, currently only ZhipuAI is supported for those.

```bash
cp .env.example .env
# then edit .env: LLM_BASE_URL, LLM_API_KEY, MODEL_NAME
```

The README then directs you to fill in config/config-web.yaml, which is selected by the PY_ENVIRONMENT=web setting in .env. Only after that do you launch:

```bash
python app.py
```

The README says the interface is then reachable at http://localhost:7860. A first real use is the text chat tab: type a symptom description and the agent answers through whichever provider you configured. If you want the knowledge graph path, Neo4j has to be installed and running first, and the graph connection details go into config/config-web.yaml.

## Loading the medical knowledge graph into Neo4j

Graph retrieval is optional but is the part the README documents in most detail, because the import step is awkward. The project uses the OpenKG dataset "面向家庭常见疾病的知识图谱" (a knowledge graph for common household diseases), and the README states that using this particular graph requires no changes to config/config-web.yaml. The community edition of Neo4j is sufficient.

The procedure is to rename the .dump file to your target database name, stop Neo4j, then load the dump:

```bash
neo4j-admin database load <database-name> --from-path=/path/to/dump-folder/ --overwrite-destination=true
```

Two warnings matter here. The README explicitly flags that --overwrite-destination replaces the data in your existing database, so do not run this against a graph you care about. And if the load reports an unsupported version, the message tells you to run neo4j-admin database migrate before starting the service again. The README's own note that no correct tutorial for this step could be found in the Chinese-language internet is a fair signal of how fiddly it is.

## Where the project is thin

Three gaps stand out. First, there is no evaluation story. The README lists capabilities such as medical record recognition and knowledge-graph-enhanced answers, but publishes no accuracy figures, no test set and no comparison against a plain LLM baseline. The tests/ directory exists in the repository listing, but the README does not describe what it covers. For a system whose stated purpose includes disease triage, that absence is the single largest reason to treat it as a prototype.

Second, provider support is uneven by the maintainers' own admission. The README says the team lacked the ability to apply for and test multiple APIs and that bugs are likely, which is why it invites issues and pull requests. If you are relying on a specific endpoint, expect to debug the client layer yourself.

Third, the multimodal generation features are narrower than the feature table suggests. Image generation, image understanding and video generation are all ZhipuAI-only according to the .env.example comments, so a deployment built on a different provider loses those paths entirely. The README does not document fallback behaviour when a configured model is unavailable.

## cyber-doctor versus a plain RAG stack

The obvious alternative is assembling LangChain, a vector store and Gradio yourself, which is essentially what cyber-doctor does internally. The difference is the graph. A conventional RAG pipeline chunks documents, embeds them into a vector index such as the faiss-cpu dependency listed here, and retrieves by similarity. cyber-doctor adds a second retrieval path through Neo4j and Cypher, so relationships between diseases, symptoms and treatments can be traversed rather than merely matched by embedding proximity. For medical content, where "what treats what" is a relation and not a topic, that is a real architectural difference and the main reason to pick this project over a generic template.

The cost is operational: you now need a running Neo4j instance, a correctly imported dump and a YAML file that ties the two together, in addition to the API key a plain RAG stack needs. If your knowledge is prose rather than a graph, the graph path buys you nothing and you are carrying the setup burden for free.

## Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-08-16, roughly a month before this writing, so the codebase is receiving changes. There are no retrieved releases, which means there is no tagged version to pin and no changelog to read before upgrading. Upgrades therefore mean pulling from main and re-reading the diff, and because requirements.txt pins exact versions for most packages (openai==1.51.0, gradio==4.44.1, langchain==0.3.3, torch==2.4.1, transformers==4.45.2), a pull that changes those pins can cascade into a dependency conflict. Pinning your own environment and upgrading deliberately is the cheaper path.

The licence is GPL-3.0, shown in the README badge and present as LICENSE in the repository root. Copyleft at that strength matters if you plan to distribute a modified version or offer it as a network service, since the obligations differ from permissive licences. That is a question for your own legal review, not something the README addresses.

## Conclusion

Adopt cyber-doctor if you are prototyping a domain assistant and already have an OpenAI-compatible key or a local Ollama endpoint, because the whole stack is a git clone, a .env file, config/config-web.yaml and python app.py on port 7860. Do not adopt it as a clinical product: the README describes disease triage and record analysis but documents no evaluation, no safety guardrails and no data-retention policy, and the project is GPL-3.0, so check how that interacts with your distribution before shipping. Verify first that your chosen provider actually works with the client factory, that Neo4j community edition is running if you want graph retrieval, and that you are comfortable with the knowledge graph import step overwriting an existing database.

## FAQ

### What is cyber-doctor?

cyber-doctor, also called Cyber Huatuo, is a Python project that builds a multimodal health assistant from an OpenAI-compatible LLM, an optional Neo4j knowledge graph, a file-backed knowledge base and edge-tts voice output, with a Gradio interface. The README describes disease triage, medical record analysis and professional question answering as its main uses.

### Which Python version does cyber-doctor need?

The README states python>=3.10 and recommends 3.10 specifically, suggesting conda for environment management. The install sequence is conda create --name myenv python=3.10 followed by pip install -r requirements.txt.

### Does cyber-doctor require Neo4j?

Only for the knowledge graph retrieval feature. The README lists downloading Neo4j as an option and notes that the free community edition is enough, and that the graph connection details go into config/config-web.yaml. The rest of the interface runs without it.

### Which model providers can cyber-doctor use?

The .env.example lists any endpoint compatible with the OpenAI SDK, including ZhipuAI, Doubao, SiliconFlow, DeepSeek and Qwen, plus a local Ollama server. Image generation, image understanding and video generation are documented as ZhipuAI-only.

### What port does cyber-doctor run on?

The README says that after running python app.py the interface is available at http://localhost:7860. That port comes from the Gradio interface in app.py rather than from a documented configuration key.

## Sources

- [Issues](https://github.com/ZeroTang05/cyber-doctor/issues)
- [License: GPL-3.0](https://github.com/ZeroTang05/cyber-doctor/blob/main/LICENSE)
- [README](https://github.com/ZeroTang05/cyber-doctor/blob/main/README.md)
- [ZeroTang05/cyber-doctor on GitHub](https://github.com/ZeroTang05/cyber-doctor)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zerotang05-cyber-doctor
