medspaCy: clinical NLP rules on top of a spaCy pipeline
Library for clinical NLP with spaCy.
At a glance
- What is it?
- medspaCy bundles clinical sentence splitting, section detection, target matching and ConText assertion into a spaCy v3 pipeline. It is rule-based, MIT licensed, and most of its rules exist only in English.
- Who is it for?
- Adopt medSpaCy if you already run spaCy and need negation, uncertainty and section context around rule-matched clinical concepts, and if your text is English. Do not adopt it expecting pre-trained clinical models or non-English rule coverage: the README lists ConText rules only for English, French and Dutch, and calls future work on pre-trained clinical models unfinished.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 128 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What medSpaCy adds on top of a plain spaCy pipeline
A general spaCy model will tokenize a discharge summary and tag parts of speech, but it will not tell you that "no evidence of pneumonia" means the pneumonia finding is negated, or that the sentence sits under an Assessment and Plan header. medSpaCy is a collection of components for exactly that layer. The README describes it as "a library of tools for performing clinical NLP and text processing tasks" built on spaCy, and lists its parts: preprocess, sentence_splitter, ner, context, section_detection, postprocess, io and visualization, plus SpacyQuickUMLS for UMLS concept extraction.
The audience is narrow and specific. You need Python, you need spaCy v3, and you need clinical text where the interesting signal is the modifier rather than the entity string. If your task is generic entity recognition on news text, none of this applies. If your task is finding every medication mention in a note and knowing which ones are historical versus current, medSpaCy's ConText implementation is the reason to look at it.
One framing point from the README matters before anything else: "MedSpaCy is currently in beta." The project also states that future work will include I/O, relations extraction and pre-trained clinical models. I/O has since shipped as a module, but the pre-trained model gap has not closed, and that shapes everything about how you would use it.
How the pipeline is assembled and where each component sits
medSpaCy does not ship a model in the machine-learning sense. It ships pipeline components that you add to a spaCy Language object. The README's load examples show three entry points: medspacy.load() with no arguments, medspacy.load(nlp) to wrap an existing model such as en_core_web_sm with the ner component disabled, or medspacy.load("en_core_web_sm", disable={"ner"}).
The data flow is rule-driven. Text enters spaCy, gets tokenized, then passes through medSpaCy components in pipe order. The sentence_splitter uses PyRuSH (listed as a required dependency) to break clinical text into sentences, which matters because clinical notes are full of line breaks, numbered lists and abbreviations that a general sentence segmenter mishandles. Section detection labels spans by header, so a later component can ask which section an entity belongs to. The ner module's target matcher applies TargetRule objects, which are either literal phrases or spaCy token patterns. The context module then applies ConText, the algorithm cited from the 2009 Journal of Biomedical Informatics paper, to attach attributes such as negation and uncertainty to matched entities. postprocess is described as a "flexible framework for modifying and removing extracted entities", which is where you would drop false positives or rewrite categories.
The design consequence is that recall is bounded by your rules. A target matcher with no rule for a concept will never find that concept, no matter how much text you feed it. That is the trade-off the project makes, and it is worth being honest that it is a trade-off rather than a shortcoming: rules are auditable, and in a clinical setting an auditable false negative is often easier to defend than an unexplainable false positive.
Installing medSpaCy and running a first rule-based pass
The README gives two installation paths. The pip route is the normal one, and the setup.py route is documented alongside it. Note the README's own example for the spaCy 2 line is written as `pip install medspacy==medspacy 0.1.0.2`, which as printed is not a valid command; treat that line as documentation of intent rather than something to paste.
pip install medspacyAfter install, two dependencies come along automatically: spaCy v3 and pyrush. If you want a spaCy model underneath, install it separately and pass it in. The default load with no arguments is the shortest path to a working pipeline.
import medspacy
nlp = medspacy.load()
print(nlp.pipe_names)The README's basic usage example then adds target rules and visualizes the result. The rules are the interesting part: the same concept can carry several patterns, so "atrial fibrillation" and "afib" both map to the PROBLEM label, and the diabetes rule uses a token pattern with an optional trailing token.
from medspacy.ner import TargetRule
from medspacy.visualization import visualize_ent
target_matcher = nlp.get_pipe("medspacy_target_matcher")
target_rules = [
TargetRule("atrial fibrillation", "PROBLEM"),
TargetRule("atrial fibrillation", "PROBLEM", pattern=[{"LOWER": "afib"}]),
TargetRule("pneumonia", "PROBLEM"),
TargetRule("warfarin", "MEDICATION"),
]
target_matcher.add(target_rules)
doc = nlp(text)
visualize_ent(doc)What you should see is an annotated document: matched spans labelled PROBLEM or MEDICATION, plus whatever context attributes the ConText component assigns. The README's sample text includes "There is no evidence of pneumonia", which is the sentence the example is built to demonstrate. If the visualization renders entities but no negation attribute, the context component is not in the pipeline. The README points to the notebooks folder for per-component detail, and that is where you should go next rather than guessing at pipe configuration.
Language coverage is the sharpest limitation
The README's language table is the single most useful page in the documentation, and it is also the most discouraging. English has ConText rules, section rules and a QuickUMLS sample. French has ConText rules but "very few" section rules. Dutch has ConText rules and no section rules. Spanish has very few section rules and no ConText. Polish has nothing in any column. Portuguese, Italian and German have no ConText and no section rules, though some have a QuickUMLS sample and unit test.
The README is unusually direct about why: "some languages are effectively 'empty' from the perspective of medspacy", because the team works primarily in English, and it asks readers to contribute rules. That is an honest statement of a real constraint. If your corpus is German clinical text, installing medSpaCy gives you a pipeline skeleton and nothing to put in it. You would be writing your own ConText rules from scratch, and the project offers help integrating them but no existing rules to start from.
There is a second, quieter limitation in the same area. The QuickUMLS component depends on a fork of QuickUMLS and on UMLS resources, and the README notes that generating resources beyond the small UMLS sample is covered in a notebook rather than in the README itself. UMLS itself carries licensing terms that are separate from medSpaCy's MIT licence, so a deployment that uses QuickUMLS has a second compliance question that the MIT badge on this repository does not answer.
medSpaCy versus scispaCy and a hand-rolled spaCy pipeline
The natural comparison is scispaCy, which also builds on spaCy and also targets biomedical text. The difference in approach is what each one assumes about where the knowledge lives. scispaCy centers on pre-trained models and embeddings trained on biomedical literature. medSpaCy centers on rules and pipeline components, with the README explicitly listing pre-trained clinical models as future work. If your text is published abstracts and you want a model that already knows biomedical vocabulary, scispaCy's premise fits better. If your text is a clinical note with "Afib" in a Past Medical History section and "no evidence of pneumonia" in an Assessment section, the section and assertion structure is the problem, and that is what medSpaCy's components address.
The other alternative is doing it yourself on top of spaCy. You can write a PhraseMatcher, add a custom attribute for negation with a handful of trigger words, and split on section headers with a regex. That is a legitimate choice for a narrow task, and it avoids a dependency. What you give up is the accumulated edge cases: PyRuSH's sentence splitting for clinical text, the ConText implementation with its published algorithm behind it, and the postprocess framework for entity cleanup. The honest way to decide is to count how many of those three you would end up rebuilding.
Maintenance, releases and what the MIT licence does and does not cover
The last push to the repository was on 2026-06-04. The most recent releases listed are 1.3.1 and 1.3.0, both dated 2024-11-21, and 1.2.0 from 2024-05-14. The gap between the last release and the last push is worth noting: there is activity on the branch that has not been cut into a release. The 1.3.1 release notes describe batched SQLite writes for database I/O, updated dependencies supporting spaCy up to 3.8.2, dropped support for Python 3.6 and 3.7, and an option to add sentence boundaries to section headers. The pyproject.toml sets requires-python to ">=3.8".
Upgrade cost is dominated by the spaCy version, not by medSpaCy itself. The 1.3.1 notes state dependencies were reconfigured to support spaCy up to 3.8.2, which reads as a ceiling rather than a floor. If your environment moves to a newer spaCy, that compatibility line is the thing to check before anything else, and the CHANGELOG.md at the repository root is where the project records these shifts. The README also notes that as of version 0.2.0.0 medSpaCy supports spaCy v3, so anyone still on a spaCy 2 codebase is on a different branch of the project's history entirely.
The licence is MIT, stated in the README badge, the LICENSE file and the pyproject classifier. MIT is permissive: it allows commercial and closed-source use with attribution and the licence text included. Two things sit outside that grant. First, the README says the project asks to be cited if you use it in your work, which is a request rather than a licence term. Second, UMLS resources pulled in through QuickUMLS are governed by their own terms, not by medSpaCy's MIT licence. I am not a lawyer and this is not legal advice; if you are shipping a product, the UMLS terms are the item to have reviewed.
Editorial conclusion
Adopt medSpaCy if you already run spaCy and need negation, uncertainty and section context around rule-matched clinical concepts, and if your text is English. Do not adopt it expecting pre-trained clinical models or non-English rule coverage: the README lists ConText rules only for English, French and Dutch, and calls future work on pre-trained clinical models unfinished. Before committing, check the language table against your corpus, confirm your spaCy version fits the 3.8.2 ceiling noted in the 1.3.1 release notes, and read the notebooks folder for the component you plan to use.
Frequently asked questions
What is medSpaCy?
medSpaCy is a library for clinical NLP built on spaCy. It bundles components for clinical sentence segmentation, section detection, target concept matching, ConText-based assertion and negation, entity postprocessing, I/O and visualization, and the README describes it as currently in beta.
What does medSpaCy do?
It adds clinical-domain pipeline components to a spaCy v3 model. The README lists modules for preprocessing, sentence splitting, NER utilities, ConText contextual analysis, section detection, postprocessing, I/O and visualization, each usable independently or as part of the pipeline.
How do I use medSpaCy?
Install it, call medspacy.load() to get a pipeline, then add TargetRule objects to the medspacy_target_matcher pipe and process text through nlp(). The README's basic usage example adds rules for atrial fibrillation, pneumonia and warfarin, then calls visualize_ent on the processed document.
What can spaCy be used for?
spaCy is the general-purpose NLP framework medSpaCy is built on, and medSpaCy requires spaCy v3. On its own it handles tokenization and linguistic annotation; medSpaCy adds the clinical components, which is why the README shows loading an existing spaCy model with the ner component disabled and passing it into medspacy.load.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/medspacy-medspacy)