Model or dataset
SCIR-HI/Huatuo-Llama-Med-Chinese avatar
SCIR-HI/Huatuo-Llama-Med-Chinese

BenTsao (HuaTuo): instruction-tuning LLaMA, Bloom and Huozi with Chinese medical knowledge

Repo for BenCao [original name: HuaTuo (华驼)], Instruction-tuning Large Language Models with Chinese Medical Knowledge. 本草(原名:华驼)模型仓库,基于中文医学知识的大语言模型指令微调

4,995 stars498 forksPythonApache-2.0

At a glance

What is it?
BenTsao, formerly HuaTuo, is a research repository that publishes LoRA adapters and training scripts for Chinese medical question answering. The adapters are small; the base models are not, and the data quality caveats come from the authors themselves.
Who is it for?
Adopt BenTsao if you already have a 7B base model, a GPU with enough memory for half-precision inference, and a research reason to compare medical instruction tuning against a general Chinese model. Skip it if you need a supported product with documented data provenance: the README states the training set contains errors and is still being revised, and the knowledge base construction code was not published at the time of writing.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 89 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What BenTsao solves, and who it is actually for

A general Chinese language model asked a clinical question tends to answer in the register of a forum post. BenTsao, renamed from HuaTuo in May 2023, attacks that gap by instruction-tuning existing 7B base models on Chinese medical question-answer pairs. The README describes the pipeline plainly: build a medical knowledge graph and a literature corpus, generate question-answer data with the GPT-3.5 API, then fine-tune. The repository ships the second half of that pipeline plus the resulting LoRA weights.

The audience is narrow and specific. This is for researchers and engineers who want to study how much medical instruction tuning changes a base model's answers, and who are willing to supply their own base weights. It is not a hosted medical assistant. There is no inference endpoint, no packaged application, and no data-collection code for the knowledge graph: the README states that the knowledge base and dataset construction code was still being organized and would be released later. The published training set is described as more than eight thousand examples.

LoRA on half-precision base models: the mechanism

Every base model in the collection is adapted with LoRA rather than full fine-tuning. The README gives the reason directly: LoRA on a half-precision base model is a trade-off between compute and model quality. The practical consequence is that the repository distributes adapters, not merged checkpoints. A downloaded adapter folder contains adapter_config.json and adapter_model.bin, and inference loads those on top of a separate base model.

The base models are listed as Huozi 1.0 (a Bloom-7B derivative from HIT-SCIR), Bloom-7B, Alpaca-Chinese-7B, and LLaMA-7B. Adapters are distributed by Baidu Netdisk and, for several of them, Hugging Face. The data split matters when you pick an adapter: some adapters were trained on the medical knowledge base, some on the literature corpus, and at least one on both. The literature work covers liver cancer only, and the README says only the single-disease parameters for liver cancer are open; models for sixteen hepatobiliary and pancreatic diseases are described as planned.

Inference is driven by infer.py with four arguments: base_model, lora_weights, instruct_dir, and prompt_template. The prompt template is not interchangeable across models. Huozi and Bloom use templates/bloom_deploy.json, while LLaMA and Alpaca use templates/med_template.json for the knowledge-base adapters and templates/literature_template.json for literature adapters. Pairing a template with the wrong base model is the most likely way to get incoherent output, and nothing in the scripts detects the mismatch for you.

Installing BenTsao and running a first inference

The quick start assumes Python 3.9 or newer. Dependencies are pinned in requirements.txt, which fixes transformers at 4.30.1, peft at 0.3.0, bitsandbytes at 0.37.2, and accelerate at 0.20.1. Those pins are old enough that a fresh environment may resolve slowly, so install them in a dedicated virtual environment rather than into a shared one.

bash
pip install -r requirements.txt

That command installs the training and inference stack, including gradio and wandb, which the scripts import even if you only intend to run inference. Next, download a LoRA adapter and unzip it. The README states the extracted layout must contain exactly these two files:

text
lora-folder-name/
  - adapter_config.json
  - adapter_model.bin

With a base model and an adapter in place, the repository provides ready-made scripts. The knowledge-base path is the shortest:

bash
bash ./scripts/infer.sh

The README states that infer.sh calls infer.py with placeholder paths, and that you must replace base_model, lora_weights, and instruct_dir before running. Test cases live in ./data/infer.json and can be swapped for your own file as long as the format matches. The underlying call looks like this:

bash
python infer.py \
  --base_model 'BASE_MODEL_PATH' \
  --lora_weights 'LORA_WEIGHTS_PATH' \
  --use_lora True \
  --instruct_dir 'INFER_DATA_PATH' \
  --prompt_template 'TEMPLATE_PATH'

Choose TEMPLATE_PATH from templates/bloom_deploy.json for Huozi and Bloom, or templates/med_template.json for LLaMA and Alpaca adapters trained on the knowledge base. A second script, ./scripts/test.sh, is offered as a reference. To train on your own data instead, the README points at ./data/llama_data.json as the format to copy and ./scripts/finetune.sh as the entry point.

The data quality caveat comes from the authors

Most model cards bury this. BenTsao states it twice. The README says the training set, although built with knowledge injected, still contains errors and incomplete entries, and that the dataset will be iterated with better strategies. It repeats the point for the literature data: the one thousand liver cancer samples in ./data-literature/liver_cancer.json are described as limited in quality, with further iterations planned and release as a public dataset later.

That matters more here than in a general-purpose model, because the failure mode is a fluent, confident, wrong clinical answer. The evaluation table in the README is explicitly dated March 2023, which means it predates the Huozi adapters added in August 2023 and the literature adapters. There is no published evaluation covering the current adapter set.

A second limitation is the knowledge-tuning method. The README describes a three-stage process: fill retrieval parameters from the question (a head term and an attribute), query the knowledge base with those parameters, then generate an answer from the retrieved knowledge. This is retrieval-augmented generation with the retrieval step exposed to the model. It is a more interesting research direction than plain supervised fine-tuning, but it inherits the knowledge base's coverage: if the head term or attribute is missing, the pipeline has nothing to ground on. Sample data for this mode is in data/knowledge_tuning_data_sample.txt, and the README does not document rollback or fallback behavior when retrieval returns nothing.

Picking between BenTsao and a general Chinese model

The obvious alternative is not another medical model but the base models themselves. Alpaca-Chinese-7B and Huozi are both general Chinese question-answering models, and BenTsao's contribution is the adapter layered on top. The difference in approach is that a general model answers from pretraining alone, while BenTsao adds a supervised signal drawn from a structured medical knowledge graph and, for the literature adapters, from the conclusion sections of 2023 Chinese papers on liver cancer. Whether that signal helps is exactly what the repository exists to test.

The trade-off is operational. A general model is one download. BenTsao is a base model plus a separate adapter plus a template that must match both, and the adapter is distributed through Baidu Netdisk for most variants, which is inconvenient outside mainland China. If your goal is a working Chinese medical chatbot rather than a research comparison, the extra moving parts buy you nothing unless you can measure the difference yourself.

Compute, licence, and what upgrades cost

The README gives concrete training numbers: LLaMA instruction tuning ran for ten epochs on a single A100-SXM-80GB, taking roughly 2 hours 17 minutes, with about 40 GB of memory used at batch_size 128. The README states that 24 GB cards such as the 3090 or 4090 should support the work if you reduce batch_size. Those figures describe training; inference memory is not documented.

Code is Apache-2.0, which is permissive for the scripts and utilities. The licence does not travel with the weights. The base models carry their own terms: Bloom-7B, LLaMA-7B, Alpaca-Chinese-7B, and Huozi are separate projects with separate conditions, and the LoRA adapters are derivative of them. If you plan to use any of this commercially, check the base model licence and the adapter's own terms rather than assuming Apache-2.0 covers the stack. This is not legal advice.

Upgrade cost is mostly environmental. The pinned transformers and peft versions date from mid-2023, and moving to a current stack would require re-testing the adapter loading path in infer.py and generate.py. The repository also includes export_hf_checkpoint.py and export_state_dict_checkpoint.py for converting checkpoints, which is the path you would take if you want to serve a merged model instead of loading LoRA at runtime.

Editorial conclusion

Adopt BenTsao if you already have a 7B base model, a GPU with enough memory for half-precision inference, and a research reason to compare medical instruction tuning against a general Chinese model. Skip it if you need a supported product with documented data provenance: the README states the training set contains errors and is still being revised, and the knowledge base construction code was not published at the time of writing. Before you commit, run the infer script on your own questions with the matching template and check the answers against a medical reference, because the repository's own evaluation table is dated March 2023.

Frequently asked questions

What is BenTsao, and how does it relate to HuaTuo?

They are the same project. The README states the model was renamed from HuaTuo to BenTsao on 2023-05-12, and the repository publishes instruction-tuned Chinese medical language models built on LLaMA, Alpaca-Chinese, Bloom, and Huozi.

How do I install BenTsao and run inference?

Install the pinned dependencies with pip install -r requirements.txt, download a LoRA adapter, then run bash ./scripts/infer.sh after replacing the base_model, lora_weights, and instruct_dir placeholders in the script. The prompt template must match the base model: templates/bloom_deploy.json for Huozi and Bloom, templates/med_template.json for LLaMA and Alpaca knowledge-base adapters.

What is the Chinese name for the llama animal?

The repository does not cover this. The project's Chinese name, 本草, refers to a class of Chinese medical texts, and the README does not discuss the animal.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. SCIR-HI/Huatuo-Llama-Med-Chinese on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scir-hi-huatuo-llama-med-chinese.svg)](https://hysenlabs.com/projects/scir-hi-huatuo-llama-med-chinese)