MedLLMsPracticalGuide is one README, and four of its six update entries are commented out
[Nature Reviews Bioengineering🔥] Application of Large Language Models in Medicine. A curated list of practical guide resources of Medical LLMs (Medical LLMs Tree, Tables, and Papers)
At a glance
- What is it?
- A curated bibliography of large language models in medicine, kept as a single file with no code and no tooling, where the visible update log has two entries seventeen months apart, four version announcements are hidden inside an HTML comment, and the contribution instructions link to a different awesome-list repository.
- Who is it for?
- MedLLMsPracticalGuide is worth using as what it is: a well-organised index of the medical language model literature, arranged by pipeline stage, by task, by clinical application and by failure mode, and maintained by people who wrote the survey it accompanies. That taxonomy is the real contribution and it is better than most.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 85 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The repository has three files and none of them is code
The root of this repository contains a licence file, a readme, and an image directory. That is the whole project.
There is no source directory, no scripts, no data files, no tests and no continuous integration. The language field on the repository records the primary language as unknown, which is what happens when a repository has no code to classify.
That is not a criticism of the format. An index of the literature is legitimately a document, and the document here is the point. But it has consequences for anyone planning to rely on it or extend it, and they are worth stating before anyone does.
The first is that there is no schema. A curated list lives or dies on whether every entry is formatted the same way, and with no validation script and no pull request template, that consistency rests entirely on reviewers noticing. The second is that there is no history beyond the readme's own update section, so you cannot see when a particular entry was added or whether a given model row has been revisited since it was written. The third is that a list this size, spanning pipeline stages, tasks, applications and failure modes, is going to accumulate entries at a rate no single-file editor can keep coherent.
The companion survey is the durable artefact. The repository is the living index around it, and the two were published at different times under different titles.
Half the update log is inside an HTML comment
The readme describes itself as an actively updated list. The evidence for that is a section of dated entries, and half of the evidence is not rendered.
Two entries are visible. The first is the release of the repository and the survey in November 2023. The second is April 2025, recording that the paper was published in Nature Reviews Bioengineering and that the repository had reached 1,500 stars.
Between them sit four more entries, and all four are wrapped in an HTML comment so they do not appear on the rendered page: a 1,000-star milestone in October 2024, and version six, five, four and three updates in July, May, March and February 2024. Each of those version announcements points at the same arXiv identifier, which is how a reader can tell they are revisions of the same preprint rather than separate releases.
So the visible history is two entries seventeen months apart, and the substantive curation history is invisible. That is an odd way to handle a changelog: commenting out old entries is the opposite of archiving them, and it means the page that would tell a reader how current the list is tells them much less than it could.
The dates underneath still tell the story. The newest visible entry is April 2025. The last push to the repository is dated 2026-07-10, so the file has been touched more recently than the log admits, and no entry explains why. There are no releases either, so there is no tag to compare a fork against.
For a list whose value depends on being current, this is the part to check before you cite it.
The contribution instructions point at a different repository
The contributing section invites additions twice, and the second invitation has the wrong target.
The text asks people who want to add a model or a piece of work either to email two individuals or to open a pull request. Two email addresses are given, one at Oxford and one at Rochester, both matching the affiliations listed for the paper's first and fourth authors. Routing every contribution through two personal inboxes is workable for a paper of this size, and it is consistent with a small number of people maintaining a document.
The pull request link is the problem. The link text points at this repository's pull request page. The address underneath points at a different repository entirely, one whose name describes a curated list of multimodal models in medical imaging. Somebody copied this section from that list and changed the visible text without changing the target.
The consequence for a contributor is small and annoying: following the printed instruction takes you to a project you did not mean to open, where your pull request would be off-topic and closed. Nobody is harmed, but the first thing the page asks a new contributor to do does not work.
It is also a signal about the review process. There is no contributing document beyond this paragraph, no code of conduct file in the root, no template describing what a good entry looks like beyond one line of format, and no statement of who reviews a submission or how long it takes.
The submission format asks for a code link the tables cannot supply
The entire contribution specification is one line of markdown, given as a template:
* [**Name of Conference or Journal + Year**] Paper Name. [[paper]](link) [[code]](link)Read as a template it is a good one. It demands a venue and a year, a paper link and a code link, and putting the venue in bold at the front means every entry carries its own provenance and a reader can filter the list by where something was published.
The problem is that the template is narrower than the content it governs. The list covers papers, models and datasets, and the tables include leaderboards as well. A dataset has no code repository. A leaderboard has neither a paper nor a repository in the ordinary sense. A model card is neither. So for the majority of what the list actually contains, the required format has a slot that cannot be filled, and the readme does not say what to put there instead.
The venue field has a related ambiguity. It asks for a conference or a journal plus a year, which does not accommodate a preprint, and the project's own companion paper spent its first year as an arXiv preprint under a different title than the one it eventually published under. A list about a fast-moving field will accumulate a lot of preprints, so the format will be applied to a case it was not designed for.
Neither issue is fatal. Both are the kind of thing that is settled by the two people reading submissions, which is the real review mechanism here.
The taxonomy is the contribution, and it is better than the list
The section structure is the most valuable thing in the repository, and it is worth setting out because it is what a reader actually uses the document for.
Building is split three ways: pre-training from scratch, fine-tuning general models, and prompting general models. Medical data is split three ways: clinical knowledge bases, pre-training data, and fine-tuning data. Downstream tasks are split by output shape rather than by subject, into generative tasks with summarisation, simplification and question answering, and discriminative tasks with entity extraction, relation extraction, text classification, natural language inference, semantic textual similarity and information retrieval.
Clinical applications are the longest list: retrieval-augmented generation, medical decision-making, clinical coding, clinical report generation, medical education, medical robotics, medical language translation, and mental health support.
Then the failure modes, which is the section most curated lists omit. Hallucination. Lack of evaluation benchmarks and metrics. Domain data limitations. New knowledge adaptation. Behaviour alignment. And ethical, legal and safety concerns. A reader who wants to know what goes wrong with these systems gets a named category for each failure rather than having to infer it from a list of successes.
Future directions close it, beginning with new benchmarks, interdisciplinary collaborations, and multimodality.
That is a real editorial contribution, and it outlives any individual entry. It is also the part that ages worst, because a taxonomy encodes what its authors thought the field's axes were in 2023.
Two of the five research questions ask the same thing
The survey scopes itself with five numbered questions. They are short and they are the clearest statement of intent in the document.
The first asks how medical language models should be built. The second asks what the measures of downstream performance are. The third asks how they should be used in real-world clinical practice. The fourth asks what challenges arise from their use. The fifth asks how we should better construct and utilise them.
The fifth is the first with a second clause. How to construct a model and how to utilise a model are different problems, one about training and one about deployment, and the third question already covers utilisation in clinical practice. So of five questions, two overlap and a third overlaps with the fifth, which leaves three distinct axes: construction, evaluation, and clinical use.
That is a fine survey. It is worth noticing because the questions are also the list's implicit structure, and a taxonomy built on five questions where two are duplicates has an axis it will quietly double-count.
Alongside them sit two goals. The first is surpassing human-level expertise. The second is that emergent properties appear as medical model size scales up. Both are empirical hypotheses rather than aspirations, and the second in particular is falsifiable in a way the first is not. Neither has a success criterion attached on the visible page, and each is followed by an empty container where a figure would go.
The paper changed title between preprint and journal, and the link set reflects both
The project is two artefacts with two names, and the readme carries both.
The preprint is titled as a survey of large language models in medicine, covering progress, application and challenge, and it lives at an arXiv identifier from November 2023. The journal article is titled as an application of large language models in medicine, published in Nature Reviews Bioengineering, and the update log records that publication in April 2025. Same work, different title, and the repository's declared homepage is the arXiv abstract page rather than the journal article.
The readme handles this by linking both, and by giving each of the six version announcements the same arXiv identifier, which is the correct way to signal that versions two through six are revisions of one preprint. That is a careful piece of linking. It is undermined slightly by four of those six announcements being invisible.
The author list is long, with two core-contributor and corresponding-author markers explained in a footnote, and the affiliations run to fourteen institutions. Most are universities across the United Kingdom, North America and Hong Kong. Three are companies: a cloud and AI group, an internet company, and a cloud provider. A mix like that is normal for a survey of this kind and says something useful about who the field's tooling comes from.
One navigation error is worth mentioning because readers use the contents list as a map. The leaderboard entry is spelled with the letters transposed, and the typo has propagated into that entry's own anchor, so a reader who copies the fragment link lands on a URL that does not match the visible text.
Editorial conclusion
MedLLMsPracticalGuide is worth using as what it is: a well-organised index of the medical language model literature, arranged by pipeline stage, by task, by clinical application and by failure mode, and maintained by people who wrote the survey it accompanies. That taxonomy is the real contribution and it is better than most. What it is not is a maintained registry. The repository has three entries in its root, the update log that would tell you how fresh the list is has half of its history commented out, and the newest visible entry is from April 2025. So treat it as a snapshot with a good structure and check dates on anything you cite. Two things to do rather than assume. If you want to contribute a paper, use the pull request page for this repository rather than the link printed in the contributing section, which points at an unrelated imaging list. And if you rely on it for a literature review, verify that the models and datasets in the tables have been updated since the version six announcement, because nothing in the repository enforces that and nothing in the visible log reports it. For anything touching clinical decisions, the accompanying journal article and its own limitations are the right source, not a link list maintained on a best-effort basis.
Frequently asked questions
What is MedLLMsPracticalGuide?
A curated list of practical guide resources for medical large language models, covering papers, models, datasets and leaderboards, together with the survey it is based on. It exists as a single README file in a repository whose only other entries are a licence and an image directory.
How current is the medical LLM list in MedLLMsPracticalGuide?
The newest visible update entry is dated 2025-04-08, recording the journal publication of the survey. Four earlier entries covering versions three through six and a milestone are wrapped in an HTML comment and do not render. The last push to the repository is dated 2026-07-10, later than any visible entry, and there are no releases to compare against.
How do I add a paper to MedLLMsPracticalGuide?
The format is a single markdown line with the venue and year in bold, then the paper title, a paper link and a code link. The contributing section also gives two email addresses or points at a pull request page, though the pull request link printed there targets a different repository.
What are the two stated goals of medical large language models in the guide?
Surpassing human-level expertise, and the emergence of properties as model size scales up. Both are stated as goals on the page, and neither is given a success criterion there.
What challenges does MedLLMsPracticalGuide catalogue?
Hallucination, lack of evaluation benchmarks and metrics, domain data limitations, new knowledge adaptation, behaviour alignment, and ethical, legal and safety concerns. Downstream tasks are split into generative and discriminative groups, and clinical applications into eight named areas.
Is the MedLLMsPracticalGuide paper the same document as the arXiv preprint?
It is the same work under two titles. The preprint is titled as a survey of large language models in medicine, and the journal article in Nature Reviews Bioengineering is titled as an application of large language models in medicine. Both are linked, the versions run two through six under one arXiv identifier, and the repository's homepage points at the preprint.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ai-in-health-medllmspracticalguide)