Open-source project
inception-project/inception avatar
inception-project/inception

INCEpTION: A Self-Hosted Annotation Platform With Knowledge-Base Linking and Live Recommenders

INCEpTION provides a semantic annotation platform offering intelligent annotation assistance and knowledge management.

719 stars171 forksJavaApache-2.0

At a glance

What is it?
INCEpTION is a Java web application for building annotated text corpora against your own ontology, with recommenders that train on annotations as you make them. This review covers what it does, how to install it, and where it stops being the right tool.
Who is it for?
Adopt INCEpTION if you are running a multi-annotator corpus project that needs typed layers, an ontology you control, and inter-annotator agreement before you call anything gold. Do not adopt it if you only need to label a few hundred sentences in a spreadsheet-style tool, or if you cannot run a Java server or desktop installer on your own hardware.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The corpus-building problem INCEpTION targets

Most annotation tools make one of two bets. Either they ship a fixed scheme (tokens, POS tags, named entities) and you adapt your research question to it, or they hand you a blank canvas and you encode your scheme in a config file that lives outside the tool. INCEpTION takes a third position: the annotation scheme is defined in the browser, and the same text can carry entities, relations, coreference, syntax, frames and document labels at once, each with typed features and slots. That matters when your labels are not flat. A relation between two entities needs both spans to exist before the relation can be drawn, and a coreference chain needs a layer that survives across sentences. The README describes stacking as many layers as the scheme needs over the same text.

The second bet is about knowledge. If your entities are supposed to point at something in the world, you usually end up maintaining a lookup table by hand. INCEpTION loads RDF, OWL, OBO, SKOS or Turtle, or queries a remote SPARQL endpoint live, and the README lists ready profiles for Wikidata, SNOMED CT, the Gene Ontology, the Human Phenotype Ontology and GND. The audience is therefore narrow and specific: computational linguists, clinical NLP groups, digital humanities projects and anyone building a training set where the label set is an ontology rather than a tag list.

How the annotation, recommender and curation layers fit together

The architecture visible from the outside has three moving parts. The first is the annotation editor, where a project's layers, features and slots are configured and where annotators work. The second is the recommender system. According to the README, recommenders train on what has been annotated so far and active learning asks about the cases they are least sure of. The important detail is the sentence that follows: nothing is stored until you accept it. Suggestions are proposals in the interface, not silent edits to the corpus, which keeps the gold standard free of machine output unless a human put it there.

The third part is curation. Several annotators can be curated into a gold standard, inter-annotator agreement can be measured, and the Explorer charts what was collected. This is the part that decides whether the corpus is usable, and it is why INCEpTION is a project tool rather than a personal one. A single annotator produces annotations; a curated set of annotators produces a dataset you can defend.

There is also an extension surface. The README points to a REST API, webhooks, and external recommenders hosted in a separate repository, inception-project/inception-external-recommender. That is the escape hatch when the built-in recommenders are not the model you want: you serve your own, and INCEpTION calls it. Import covers plain text, PDF, HTML and TEI; export covers UIMA CAS XMI and JSON with custom layers intact, or CoNLL-U. The UIMA lineage is visible in the topics and in the repository layout, which keeps a single inception/ module tree under the Maven root.

Installing INCEpTION and annotating your first document

The README does not list command-line install steps. It points to the Admin Guide for installing and running the platform for a group of users, and describes two deployment shapes: a desktop installer for one person, or a server deployment for a whole institution on your own hardware with your own single sign-on. The Developer Guide covers building and extending it. So the honest first step is to read the Admin Guide and pick the shape that matches your situation.

If you want to look before you install, the README links a demo server. The URL is given directly, so you can open it and click through the editor without any setup:

bash
https://morbo.ukp.informatik.tu-darmstadt.de/demo

If you are building from source, the repository is a Maven multi-module project rooted at pom.xml with the code under inception/. The Developer Guide is the reference for the build; the Maven wrapper is present in the repository layout as .mvn/, so a build does not require a separately installed Maven. The README's contributing section shows the fork-and-branch flow it expects for changes:

bash
git checkout -b my-feature

Once you have a running instance, the Getting Started Guide is the path through the first project, and the tutorial videos are linked from the README. The user-facing workflow the README implies is: create a project, define your layers and features in the browser, import a document, annotate, then export. If you intend to use an ontology, load it at project setup time from a file or point the project at a SPARQL endpoint, because entity linking depends on it being there. Expect the first session to be configuration rather than annotation.

Where INCEpTION is the wrong tool

The recommender is the most oversold part of any annotation platform, and INCEpTION's own description is more careful than most: recommenders train on what you have annotated so far. That means the first documents in a new project have nothing to learn from. If your project is a few hundred sentences, you will finish before the recommender has enough signal to be worth accepting, and you will have paid the setup cost of a multi-layer scheme for nothing.

The second limit is operational. INCEpTION is a Java application with a server deployment mode aimed at institutions. If you have no one to run it, the desktop installer covers a single person, but the curation and agreement features assume multiple annotators feeding one project. Groups that cannot host a server and cannot standardise on a desktop install lose exactly the features that distinguish INCEpTION from a simpler editor.

The third is format fidelity. The README lists what imports and what exports, but not what survives each conversion. TEI in and CoNLL-U out is not a round trip, and custom layers are only preserved on the UIMA CAS XMI/JSON path. If your pipeline is built around a format that is not on that list, the export step is where the project will cost you time. The README is silent on rollback and on versioning of annotation schemes after annotators have started, so treat scheme changes mid-project as a risk to verify in the Admin Guide rather than an assumption.

INCEpTION compared with brat and Prodigy

brat is the closest historical comparison: a web-based annotation tool with a configuration language for defining entity and relation types. The difference is where the scheme lives and what it can reference. brat's configuration is text files you write and deploy; INCEpTION's layers, features and slots are defined in the browser, and its entity types can be backed by an ontology loaded from RDF, OWL, OBO, SKOS or Turtle or queried from a live SPARQL endpoint. brat has no equivalent of the recommender loop or the curation and agreement workflow described in the README.

Prodigy is the opposite trade. It is a commercial, script-driven tool built around a Python workflow, and its active learning is central rather than optional. INCEpTION is Apache-2.0, self-hosted, and its automation is one part of a platform that also does ontology linking and multi-annotator curation. If your work is a Python pipeline where a human labels a stream of examples and the model updates immediately, Prodigy's shape fits better. If your work is a corpus with a scheme you must document and annotators you must reconcile, INCEpTION's shape fits better. Neither is a drop-in for the other, and the export formats differ, so the choice is hard to reverse once annotation has started.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-02. Releases are frequent: inception-41.5 on 2026-09-01, inception-41.4 on 2026-08-18 and inception-41.3 on 2026-08-04. A roughly two-week cadence between point releases means you should expect to upgrade, and the version numbering (41.x) suggests the project does not treat each release as a long-term support line. Nothing in the README describes an LTS branch, a migration tool or a rollback procedure, so the upgrade cost is a real unknown that the Admin Guide is the place to check before you put a production corpus on an older release.

Licensing is Apache-2.0, stated in the README and present as LICENSE.txt at the repository root alongside NOTICE.txt. Apache-2.0 is permissive and includes a patent grant, which is usually what institutions want. Two things to check yourself rather than assume: the NOTICE.txt file, which records attribution obligations you may need to carry forward, and the licence of any ontology you load. Loading SNOMED CT or GND is a separate licensing question from the licence of the software, and the README's profiles make those ontologies easy to load without making them free to redistribute. This is not legal advice; it is a pointer to the two files and the one external dependency that most often surprise people.

Editorial conclusion

Adopt INCEpTION if you are running a multi-annotator corpus project that needs typed layers, an ontology you control, and inter-annotator agreement before you call anything gold. Do not adopt it if you only need to label a few hundred sentences in a spreadsheet-style tool, or if you cannot run a Java server or desktop installer on your own hardware. Before committing, verify three things against your own data: that the import path you need (plain text, PDF, HTML, TEI) preserves the structure you care about, that your ontology loads either from a file in RDF, OWL, OBO, SKOS or Turtle or from your SPARQL endpoint, and that the CoNLL-U or UIMA CAS XMI/JSON export keeps the custom layers your downstream pipeline expects. The recommender only becomes useful after you have annotated enough to train it, so plan the first annotation round without it.

Frequently asked questions

How do I install INCEpTION?

The README does not give command-line install steps. It points to the Admin Guide for installing and running the platform for a group of users, and describes a desktop installer for one person or a server deployment for an institution. The Developer Guide covers building and extending it from source.

How does INCEpTION work?

It is a web-based annotation platform where you define layers, features and slots in the browser and annotate the same text with entities, relations, coreference, syntax, frames and document labels. Recommenders train on what you have annotated so far, and nothing is stored until you accept a suggestion.

How do I use INCEpTION for the first time?

The README recommends working through the Getting Started Guide, watching the tutorial videos, and trying the demo server at the linked URL before installing anything. Once a project exists, you define the scheme, import a document, annotate, and export.

What is INCEpTION?

INCEpTION is a semantic annotation platform offering intelligent annotation assistance and knowledge management, written in Java and released under Apache-2.0. It is aimed at building annotated text corpora against your own annotation scheme and ontology.

Official sources

  1. inception-project/inception on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/inception-project-inception.svg)](https://hysenlabs.com/projects/inception-project-inception)