MedicalGPT: a training pipeline for medical LLMs, not a medical chatbot
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。
At a glance
- What is it?
- shibing624/MedicalGPT is a Python training framework that walks a base model through continued pretraining, SFT, reward modeling, RLHF, DPO, ORPO, GRPO and OPD. It ships training scripts, not a hosted medical assistant.
- Who is it for?
- Adopt MedicalGPT if you have a base checkpoint, a domain corpus and GPUs, and you want the PT/SFT/DPO/ORPO/GRPO/OPD stages wired into runnable scripts rather than assembled from papers. Do not adopt it expecting a medical chatbot: there is no hosted product, no login, and the repository is a training codebase.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
MedicalGPT is a training pipeline, and the name oversells the medical part
The repository solves a narrow problem: given a base language model and a domain corpus, run the post-training stages that turn it into an instruction-following model for that domain. The README describes the intended flow as pretraining, supervised finetuning, RLHF with reward modeling and reinforcement learning, DPO, and standalone OPD. Medical is the worked example, not a constraint. Nothing in the feature list restricts the data to clinical text, and the release history shows the project tracking general model families (Llama-3, Qwen-2, Qwen-2.5, Qwen3.5) rather than medical datasets.
The audience is therefore engineers and researchers who already have a checkpoint and a corpus and need the training loop. It is not for clinicians who want an answer engine, and it is not for teams that want an API. The repository has no homepage, and the search phrases people use around login, app and free tiers point at a product that does not exist here. If you arrived looking for a free medical AI to query, this is the wrong repository.
The four-stage framing comes from the ChatGPT pipeline as described in Andrej Karpathy's State of GPT talk, which the README cites directly. DPO and ORPO are attributed to their papers. That lineage is worth knowing because it tells you what the code is modelling: a research pipeline that has been made runnable, not a product architecture.
How the training stages connect: PT, SFT, RM, RL, DPO, ORPO, GRPO, OPD
The stages are sequential and each one consumes the artifact the previous stage produced. Continued pretraining runs on raw domain documents to shift the model's distribution toward the domain, and the README marks it optional. Supervised finetuning then runs on an instruction dataset to align the model with instruction intent and inject domain knowledge. From there the pipeline forks.
One branch is RLHF, split into reward model training on a human preference ranking dataset, then reinforcement learning that uses that reward model to update the SFT policy. The README names the HHH principle (helpful, honest, harmless) as what the reward model is meant to capture. The other branch is DPO, which the README frames as optimizing the language model directly against preferences without a separate reward model, and describes as easier to implement and train than RLHF. ORPO, added in v1.9, is a monolithic preference method that the README attributes to the ORPO paper. GRPO arrived in v2.4 for both LoRA and full-parameter training, and the release note describes it as a pure RL method. OPD, on-policy distillation, arrived in v2.7 with its own entry point at training/opd_training.py and a launch script at scripts/run_opd.sh.
The concrete consequence of this layout is that you choose a branch, not a menu. If you run DPO you do not need the reward model stage; if you run RLHF you do. The scripts directory holds one launcher per method (the README points at scripts/run_orpo.sh for ORPO), so the configuration surface is shell scripts plus the training entry points rather than a single config file. That is a real design choice with a cost: there is no one place to read the full set of hyperparameters.
Installing MedicalGPT and running a first SFT job
The repository has no published package, so installation means cloning and installing requirements.txt. Note the version floors: transformers>=5.6.0 and trl>=0.29.0 are recent, and peft>=0.19.1 is pinned high. The file also carries a commented-out bitsandbytes line with the note that it is not available on pypi, so quantized training needs that dependency sourced separately.
git clone https://github.com/shibing624/MedicalGPT.git
cd MedicalGPT
pip install -r requirements.txtAfter this, the training entry points live under training/ and the launchers under scripts/. The README does not give a single canonical SFT command in the text provided; it points at the scripts directory, where each method has a shell script that sets the model path, data path and hyperparameters before invoking the Python entry point. Open the relevant script and edit the paths rather than inventing flags.
The launcher reads its arguments from the script body, so the model name, dataset path and output directory are all edited in the file. What you should see is a training run that writes checkpoints to the output directory named in the script and logs to tensorboard, which is in requirements.txt. The demo directory holds the inference side: demo/inference.py for a local model, demo/gradio_demo.py for a web UI, demo/fastapi_server_demo.py and demo/openai_api.py for serving, and demo/chatpdf.py for retrieval over files, which the v1.7 release note describes as combining a finetuned LLM with a knowledge base.
Data preparation is the part the README covers least in the text provided. The data/ directory holds samples, and v2.6 added toolcall examples there. Expect to convert your corpus into the format the chosen script expects before anything runs.
Where MedicalGPT breaks: data, templates and the missing evaluation story
The most common failure is upstream of the code. Every stage after PT needs a dataset in a specific shape: instruction pairs for SFT, ranked preferences for reward modeling and DPO, tool-call traces for the v2.6 function-call path. The README does not describe a schema validator, so a malformed preference file surfaces as a training error rather than a clear message. Budget more time for data conversion than for the training run.
Conversation templates are the second trap. The v2.5 release notes list qwen3, qwen3_5, qwen3_nothink and qwen3_5_nothink as templates added for the Qwen3.5 family. If your base model is not among the supported families, you are writing a template, and a wrong template silently degrades output instead of failing loudly.
Third, the repository is a training framework, so it does not answer the question a medical team actually asks: is the resulting model safe or accurate? There is no evaluation harness described in the repository, no clinical benchmark, and no reported metric. The DISCLAIMER file at the repository root exists for this reason. If your requirement is a validated clinical decision aid, this project does not produce one, and no amount of preference optimization changes that.
Maintenance is worth stating plainly. The last push was on 2026-06-03, which is more than six months before today, so this is not a repository being updated continuously right now. The release cadence in April 2026 was fast (2.5.0, 2.6.0, 2.7.0 within two weeks), and then activity stopped. Treat the code as stable-but-idle and check open issues before depending on a specific path.
MedicalGPT against Axolotl and LLaMA-Factory
The obvious alternatives are general fine-tuning frameworks such as Axolotl and LLaMA-Factory, and the difference is scope rather than quality. Those tools aim to cover as many models and datasets as possible behind a unified config, usually YAML, with a single training binary that dispatches on the config. MedicalGPT does the opposite: it implements a specific named sequence of post-training methods and gives each one its own script and entry point.
That means MedicalGPT is easier to read if you want to understand what RLHF, DPO, ORPO, GRPO or OPD actually do in code, because the stages are separated rather than abstracted behind a trainer registry. It is harder to use if you want to sweep models, because there is no unified config and no CLI described in the repository. A general framework will likely have broader model coverage and more contributors; MedicalGPT's advantage is that the pipeline order is explicit and the method implementations are traceable to the papers the README cites.
One more distinction matters for this domain. Neither Axolotl nor LLaMA-Factory is aimed at medical text specifically, but neither is MedicalGPT in any enforced sense. The medical framing here is documentation and example data, not a constraint in the code. If your corpus is legal or financial, the same scripts apply.
Licence, upgrade cost and the DISCLAIMER file
The repository is Apache-2.0, which permits commercial use and modification and includes a patent grant. That covers the code in this repository. It does not cover the base model weights you fine-tune, which carry their own licences from their publishers, and it does not cover the datasets you train on. Those are the two places where a medical deployment most often hits a licensing problem, and neither is addressed by the Apache-2.0 grant here.
The DISCLAIMER file in the repository root is the project's own statement about intended use, and it should be read before any deployment that touches patient data. This is a description of what the files say, not legal advice; get your own review.
Upgrade cost is shaped by the version floors. transformers>=5.6.0 and trl>=0.29.0 mean that a pinned environment from a year ago will not install cleanly, and moving forward may break the training scripts if upstream APIs shifted. The commented-out bitsandbytes line adds a manual step for anyone doing quantized training. Because there is no packaged release on PyPI, upgrading means pulling the repository and re-reading the scripts you depend on, particularly if you use the OPD path added in v2.7 or the tool-call path added in v2.6, both of which are recent and therefore less exercised than the SFT and DPO stages.
Editorial conclusion
Adopt MedicalGPT if you have a base checkpoint, a domain corpus and GPUs, and you want the PT/SFT/DPO/ORPO/GRPO/OPD stages wired into runnable scripts rather than assembled from papers. Do not adopt it expecting a medical chatbot: there is no hosted product, no login, and the repository is a training codebase. Before committing, verify three things: that the transformers and trl floors in requirements.txt resolve against your environment, that the conversation template you need exists among the qwen3, qwen3_5, qwen3_nothink and qwen3_5_nothink entries the v2.5 release notes list, and that the OPD and tool-call paths added in v2.6 and v2.7 match the data format you actually have. The DISCLAIMER file in the repository root is the boundary that matters most: nothing here produces a clinically validated model.
Frequently asked questions
Is there a ChatGPT for medical, and is MedicalGPT it?
MedicalGPT is not a hosted chat product. It is a training pipeline that implements continued pretraining, supervised finetuning, RLHF, DPO, ORPO, GRPO and OPD so you can train a model yourself; the demo directory provides inference and serving scripts, but there is no login or app in the repository.
Is MedicalGPT a free medical AI?
The code is Apache-2.0, so it is free to use and modify, but running it requires your own base model, your own dataset and GPUs. There is no free hosted medical assistant here; the cost moves to compute and data preparation.
Which AI is best for medical diagnosis?
The repository does not make or support a claim about diagnostic accuracy. It provides training stages and a DISCLAIMER file at the repository root, and no clinical benchmark or evaluation result appears in the repository.
What is GPT medical, and what is MedicalGPT?
MedicalGPT is described in its README as training a medical GPT model with a ChatGPT training pipeline, implementing pretraining, supervised finetuning, RLHF with reward modeling and reinforcement learning, DPO, ORPO, GRPO and standalone OPD.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/shibing624-medicalgpt)