OntoGPT: LLM Extraction That Grounds Every Identifier
LLM-based ontological extraction tools, including SPIRES
At a glance
- What is it?
- OntoGPT turns text into schema-conformant structured data by asking a model for names, not identifiers, then grounding those names against real ontologies. It is a good fit for batch curation and a poor fit for anyone who wants an agent to decide the workflow.
- Who is it for?
- Adopt OntoGPT if you have a corpus of text and a LinkML schema or a bundled template that already matches the objects you need, and you want the same extraction procedure to run over one abstract or thousands of papers. Do not adopt it if you want an autonomous agent that plans its own steps, or if your extraction target has no ontology behind it, because the grounding step is the whole point.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The identifier problem OntoGPT was built around
Ask a language model for an ontology identifier and it will often produce one that looks right and does not exist, or attach a real identifier to the wrong term. OntoGPT's answer is to never ask for the identifier at all. The README states the pipeline asks for names, grounds each name against the actual ontology through OAK annotators, and then validates every grounded identifier against its source ontology. A name that cannot be grounded is marked with the AUTO: prefix rather than guessed, so a reader can tell at a glance which parts of an output are anchored and which are not.
The audience follows from that design. This is for curators, biocuration teams and data modelers who need structured records from prose and cannot accept invented terms in the result. It is also for people who already have a LinkML schema and want to fill it from text without loading the whole schema into a prompt.
How SPIRES walks a schema instead of stuffing it
The second mechanism is the one that makes large data models workable. A big LinkML schema does not fit comfortably in a context window, and an agent that searches the model field by field for every document is slow and expensive. SPIRES walks the schema recursively instead. The model is prompted for one class at a time, with only that class's fields visible, so each individual prompt stays small while the assembled output still conforms to the full schema.
Around that core sit the parts that make it a tool rather than a script. A template plus a model gives the same procedure for every document, run from the command line or from Python, over one abstract or thousands of papers. The README describes prompt caching and output as YAML, JSON, RDF or OWL, with no agent in the loop deciding what to do next. Dozens of bundled templates cover diseases, phenotypes, drugs, genes, GO terms and environmental samples, each already wired to the right ontologies. A new template is a LinkML schema with a few annotations.
Model access goes through LiteLLM, which is why the provider list is broad: OpenAI, Azure OpenAI, Anthropic, Mistral, Groq, Cohere, Vertex AI and Replicate among others. The README recommends a provider-qualified model name such as openai/gpt-5.5 or anthropic/claude-sonnet-5, and notes that without --model, OntoGPT uses gpt-5.5. One detail worth knowing before you tune anything: reasoning models, including the GPT-5 family and Claude Sonnet 5 and Opus 5, accept only their default temperature, and OntoGPT logs a warning and retries without it if --temperature is set.
Installing OntoGPT and running a first extraction
OntoGPT needs Python 3.10 or greater. The README installs it from PyPI with pip, and the package metadata in pyproject.toml pins requires-python to >=3.10,<3.15, so the upper bound is worth checking against your interpreter before you start.
pip install ontogptCredentials come next. The README's quick start uses runoak set-apikey, and the package also reads standard LiteLLM environment variables such as OPENAI_API_KEY, ANTHROPIC_API_KEY, GROQ_API_KEY, MISTRAL_API_KEY, AZURE_API_KEY, AZURE_API_BASE and AZURE_API_VERSION directly. LiteLLM handling takes priority; Oaklib credentials are checked afterward for backward compatibility and passed through when the corresponding provider settings are missing.
runoak set-apikey -e openai <your openai api key>With a key in place, list the commands to confirm the install:
ontogpt --helpNow a real extraction. The README writes a one-line file and runs the drug template against it.
echo "One treatment for high blood pressure is carvedilol." > example.txt
ontogpt extract -i example.txt -t drugThe README says OntoGPT retrieves the necessary ontologies and prints results to the command line. Expect two things in that output. The extracted objects appear under the heading extracted_object, and a validation section reports whether each grounded identifier exists in its ontology and matches its label. Invalid identifiers are replaced when a valid term can be found. That validation block is the part to read first, because it tells you whether grounding actually worked on your input.
There is also a minimal web interface. Installing the extra and starting it are two commands, and the README is explicit that it should not be hosted publicly without authentication.
pip install ontogpt[web]
web-ontogptWhere grounding fails and where OntoGPT is the wrong tool
The AUTO: prefix is honest, but it is also a failure marker. Any name the grounder cannot resolve against the loaded ontology comes back ungrounded, and the README does not describe a fallback that recovers it. If your text is full of terms outside the ontology you selected, you will spend your time reviewing AUTO: entries rather than using the output.
The template is the second constraint. Extraction is bound to a LinkML schema, so an entity type with no template and no ontology behind it has nowhere to go. Writing a template is described as a LinkML schema with a few annotations, which is a low bar for someone who already works in LinkML and a real barrier for anyone who does not.
The design choice that rules out a whole class of use is the absence of an agent. The README frames this as a feature: no agent in the loop deciding what to do next. That gives repeatability and batchability, and it also means OntoGPT will not chase a follow-up question, re-read a table, or decide that a document needs a different template. If your task is exploratory and the steps are not known in advance, this is the wrong shape of tool.
Cost and rate limits are the practical ceiling on batch runs. Every document means model calls, and the README does not state a per-document token or price figure, so the only way to size a corpus run is to measure one yourself. The bundled Agent Skills exist for the opposite situation, teaching an external agent to call OntoGPT for the extraction and grounding step rather than reimplementing it in the agent's own context.
OntoGPT compared with an agent that calls the model directly
The alternative most teams reach for first is an agent framework with a tool that calls an LLM, plus a separate lookup against an ontology service. The difference is where the grounding happens. In the agent approach, the model is typically asked for an identifier and the agent then checks it, or the agent searches the ontology and pastes a result back into context. OntoGPT inverts the order: names first, grounding through OAK annotators second, validation of every identifier against its source ontology third.
The second difference is schema handling. An agent that walks a large data model field by field does so inside its own context, document after document. SPIRES walks the LinkML schema recursively and prompts for one class at a time, so the per-prompt payload stays small. For a wide schema over a large corpus, that is the distinction that decides whether the run is affordable.
The trade is flexibility. An agent can adapt mid-document; OntoGPT cannot. The README resolves this by offering both directions: OntoGPT as a command-line and Python extractor for repeatable runs, and Agent Skills that let an agent delegate the extraction and grounding step to OntoGPT instead of rebuilding it.
Maintenance, licensing and what an upgrade costs
The repository is not archived, and the last push was on 2026-09-09, the same day as the v1.2.0 release. The release history shows v1.1.0 on 2026-03-25, v1.1.1 on 2026-04-07, and v1.2.0 on 2026-09-09, so the cadence is uneven rather than monthly. The primary language listed for the repository is Jupyter Notebook, which is worth knowing if you intend to read the source: the notebooks directory and the templates under src/ontogpt/templates are where a lot of the working material lives, alongside the Python package.
That release history matters for upgrades. Version 1.2.0 is the release that moved credential handling to LiteLLM first, with Oaklib credentials kept for backward compatibility. If you are upgrading from 1.1.x and your deployment relies on runoak set-apikey, the README says those credentials are still checked and passed through when the corresponding provider settings are missing, but the ordering has changed. Verify your provider configuration against the setup documentation rather than assuming the old path is still primary.
The dependency surface is the other upgrade cost. The pinned list includes linkml, oaklib, litellm with the caching extra, pydantic, pymupdf, scipy and numpy, and the litellm constraint differs by Python version: >=1.95.0,<1.98 on Python below 3.11, and >=1.95.0 with no upper bound on 3.11 and above. Ontology retrieval and grounding depend on OAK, so network access to ontology sources is part of the runtime, not an optional extra.
On licensing, the package declares BSD-3-Clause in the metadata and ships a LICENSE file at the repository root. That is a permissive licence, but it covers the OntoGPT code, not the ontologies you ground against or the model providers you call. Those carry their own terms, and the README does not analyze them. This is not legal advice; check the terms of the specific ontologies and providers in your pipeline.
Editorial conclusion
Adopt OntoGPT if you have a corpus of text and a LinkML schema or a bundled template that already matches the objects you need, and you want the same extraction procedure to run over one abstract or thousands of papers. Do not adopt it if you want an autonomous agent that plans its own steps, or if your extraction target has no ontology behind it, because the grounding step is the whole point. Before committing, verify that the specific bundled template covers your entity type, that your chosen provider and model name appear in ontogpt list-models output, and that the validation section of a small sample run reports identifiers as valid rather than AUTO:.
Frequently asked questions
What model does OntoGPT use if I do not pass --model?
The README states that without --model, OntoGPT uses gpt-5.5. You can see the available names by running ontogpt list-models and using the name in the first column with the --model option.
Which providers can OntoGPT work with besides OpenAI?
OntoGPT uses LiteLLM to interface with LLMs, so any provider and model supported by the installed LiteLLM version will generally work. The README lists OpenAI, Azure OpenAI, Anthropic, Mistral, Groq, Cohere, Vertex AI and Replicate among others.
How do I install OntoGPT?
The README installs it with pip install ontogpt, after ensuring Python 3.10 or greater is present. The package metadata pins requires-python to >=3.10,<3.15.
What does the validation section in OntoGPT output tell me?
The README says the validation section reports whether each grounded identifier exists in its ontology and matches its label, and that invalid ones are replaced when a valid term can be found. Names that cannot be grounded are marked with the AUTO: prefix.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/monarch-initiative-ontogpt)