# doccano: a self-hosted text annotation tool for classification, NER and summarization data

> doccano is an MIT-licensed annotation server for text classification, sequence labeling and sequence to sequence tasks. It installs from pip or runs as a Docker container, and the last push to the repository was on 2026-04-14.

**doccano/doccano** — Open source annotation tool for machine learning practitioners.

- Repository: https://github.com/doccano/doccano
- Stars: 10,770 · Forks: 1,822
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/doccano-doccano

## What doccano is for and who ends up running it

doccano solves the data production problem that sits in front of most supervised NLP work. The README lists three task families: text classification, sequence labeling, and sequence to sequence. Sentiment analysis, named entity recognition and text summarization are the examples it gives. The pitch is deliberately narrow. You create a project, upload data, and start annotating, and the README claims you can build a dataset in hours.

The intended user is a machine learning practitioner who has raw text and no labeled corpus. That usually means a small team: a few annotators working in a browser, one engineer who owns the deployment. The feature list is written for that group, naming collaborative annotation, multi-language support, mobile support, a dark theme and a RESTful API. The REST API matters more than it looks. It means the labeling step can be scripted rather than driven entirely through the web interface, which is how most pipelines actually consume the output.

Where doccano does not fit is equally clear from the feature list. There is no mention of image, audio or video annotation. This is a text tool with a browser front end, and the topics on the repository (text-annotation, natural-language-processing) confirm the scope. If your labeling task is bounding boxes on photographs, you are looking at the wrong project.

## The architecture: a Django backend, a Nuxt frontend and a separate task worker

The repository layout tells you most of what you need to know before reading any documentation. There is a backend/ directory and a frontend/ directory, plus docker/, cloud/, docs/ and tools/. The frontend is built with Nuxt and Vue, which the repository topics confirm. The backend is Python, consistent with the pip installation path.

The operational detail that catches people out is the process split. The pip instructions start a web server with doccano webserver --port 8000, and then, in another terminal, run doccano task. The README describes that second process as the task queue that handles file upload and download. So the system is not a single process. The web server serves the UI and the API; the task process does the asynchronous work of moving files in and out.

That split has a direct consequence. If you start only the web server, the interface will come up and you can log in, but import and export jobs will not complete, because nothing is consuming the queue. Anyone who deploys doccano behind a process manager needs to supervise two commands, not one. The Docker Compose path in the repository adds RabbitMQ and PostgreSQL services on top, which reflects the same design: a queue in front of background work, and a real database behind the application.

## Installing doccano with pip and annotating your first record

The pip route requires Python 3.8 or later. Install the package, then initialize the database, create a superuser, and start the two processes. The commands below are the ones the README gives, in the order it gives them.

```bash
pip install doccano

# Initialize database.
doccano init
# Create a super user.
doccano createuser --username admin --password pass
# Start a web server.
doccano webserver --port 8000
```

In a second terminal, start the task queue:

```bash
doccano task
```

After that, the README says to go to http://127.0.0.1:8000/ and log in with the credentials you just created. The default database is SQLite 3. If you want PostgreSQL instead, install the extra dependency set and point the application at your server:

```bash
pip install 'doccano[postgresql]'
DATABASE_URL="postgres://${POSTGRES_USER}:${POSTGRES_PASSWORD}@${POSTGRES_HOST}:${POSTGRES_PORT}/${POSTGRES_DB}?sslmode=disable"
```

The DATABASE_URL string is copied verbatim from the README, including the sslmode=disable parameter, which is worth noticing: if your PostgreSQL instance requires TLS, that default in the example is not what you want.

If you would rather not manage Python at all, the Docker path is shorter. The README creates the container once with environment variables for the admin account and a named volume for persistence, then starts it:

```bash
docker pull doccano/doccano
docker container create --name doccano \
  -e "ADMIN_USERNAME=admin" \
  -e "ADMIN_EMAIL=admin@example.com" \
  -e "ADMIN_PASSWORD=password" \
  -v doccano-db:/data \
  -p 8000:8000 doccano/doccano
docker container start doccano
```

Data written inside the container persists across restarts through the doccano-db volume, and the README states that stopping the container is done with docker container stop doccano -t 5. There is also a nightly tag for people who want unreleased features, which is a reasonable thing to avoid on a machine that holds your only copy of a labeled dataset.

## Running doccano as a multi-user service with Docker Compose

The single-container Docker setup is fine for one person. For a team, the README points at Docker Compose, which needs Git and a clone of the repository:

```bash
git clone https://github.com/doccano/doccano.git
cd doccano
```

You then create a .env file. The README shows the format and links to docker/.env.example in the repository. It groups the variables into platform settings (ADMIN_USERNAME, ADMIN_PASSWORD, ADMIN_EMAIL), RabbitMQ settings (RABBITMQ_DEFAULT_USER, RABBITMQ_DEFAULT_PASS) and database settings (POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB). Note that the Compose path exposes the service on port 80, not 8000:

```bash
docker-compose -f docker/docker-compose.prod.yml --env-file .env up
```

The README says to access http://127.0.0.1/ afterwards. The presence of RabbitMQ here, alongside the separate doccano task process in the pip path, is the strongest signal that background work is a first-class part of the design rather than an afterthought.

The README also carries a Windows warning that is easy to skim past. It says to configure git to handle line endings correctly or you may encounter status code 127 errors when running services later, and it gives an alternative clone command:

```bash
git clone https://github.com/doccano/doccano.git --config core.autocrlf=input
```

That is a real failure mode with a real fix, and it is documented, which is more than can be said for some of the operational questions below.

## Where doccano gets in the way

The documentation is thin on several things an operator will want to know. The README does not document rollback of an import, and it does not describe how to back up a running instance beyond the persistence note for the Docker volume. There is no discussion in the README of horizontal scaling, of what happens to in-flight annotation work when the task process dies, or of how the queue behaves under load. The documentation site is linked, and it is where the FAQ entries for creating a user, adding a user to a project and changing a password live, but the README itself does not answer these operational questions.

The release history is the second constraint. v1.8.5 was released on 2026-01-11. The release before it, v1.8.4, was on 2023-07-20, and v1.8.3 on 2022-12-08. That is a gap of roughly two and a half years between v1.8.4 and v1.8.5. The repository is not archived and the last push was on 2026-04-14, so work has happened since the release, but a team that needs a fix shipped in a tagged version should plan around that cadence rather than assume it.

The third limitation is scope. Sequence to sequence annotation in a browser is a slower and more error-prone activity than classification, and doccano does not change that; it gives you an interface for it. If your labeling task is large and highly repetitive, the tool is a starting point, not an answer to the cost of the work itself.

## doccano against Label Studio, and what the difference actually is

The comparison people reach for is Label Studio, and the difference in approach is worth stating precisely. Label Studio is built as a general labeling platform with a configuration language that lets you compose an interface for many data types, including images, audio and video. doccano is narrower: text classification, sequence labeling and sequence to sequence, with a fixed interface per task type.

That narrowness is the trade-off. A fixed interface is faster to learn and harder to get wrong, and the README's promise of a dataset in hours depends on it. A configurable interface covers more task types but pushes setup work onto the person doing the labeling configuration. If you only ever label text, doccano's constraint costs you nothing. If your next project involves images, it costs you a migration.

The other difference is the deployment surface. doccano's pip path needs a Python environment and two long-running processes; its Compose path needs PostgreSQL and RabbitMQ. Label Studio's deployment story differs, and anyone choosing between them should compare the operational burden rather than the feature list, because the feature lists overlap heavily for text tasks. Both are open source, and doccano's licence is MIT, which is permissive and imposes few conditions on how you use or redistribute it.

## Maintenance cost, licence and what to check before you commit

The MIT licence is the least complicated part of adopting doccano. It permits commercial use, modification and redistribution with the licence and copyright notice preserved. It does not, of course, come with any warranty, and nothing here is legal advice; if your organisation has a policy on open source, run it through that.

The maintenance cost is the part to think about. You are running a web application, a database and a queue. With SQLite and the pip install, that is one machine and two processes, which is manageable. With PostgreSQL and RabbitMQ, you have three services to keep alive, back up and upgrade. Upgrades are the sharp edge: the release history shows long quiet periods punctuated by releases, so an upgrade is not a routine monthly chore you can do on autopilot. Pin the version you deploy, and read the release notes before moving.

The things to verify first are the ones the README leaves open. Confirm the export format your training pipeline consumes, because the README describes the annotation features and the REST API but does not walk through export schemas. Confirm that your deployment keeps the task process supervised alongside the web server. And if you are on Windows, apply the git line-ending configuration before you clone, not after you hit the status code 127 error.

## Conclusion

Adopt doccano if you need a self-hosted, MIT-licensed text annotation server for classification, sequence labeling or seq2seq work and you are comfortable running a Python web service plus a separate task queue process. Do not adopt it if you need image, audio or video annotation, or if you expect the project to ship fixes on a fast cadence: the last push was on 2026-04-14 and the previous release before v1.8.5 dates from 2023. Before committing, verify the export format your training pipeline expects, confirm the PostgreSQL and RabbitMQ setup you will use for multi-user work, and check the documentation for the annotation workflow of your task type.

## FAQ

### How do I install doccano?

The README gives three options: pip (Python 3.8+), Docker, and Docker Compose. With pip you run pip install doccano, then doccano init, doccano createuser, doccano webserver --port 8000, and doccano task in a second terminal. With Docker you pull doccano/doccano and create a container with the ADMIN_USERNAME, ADMIN_EMAIL and ADMIN_PASSWORD environment variables.

### Is doccano free?

Yes. doccano is released under the MIT licence, which permits commercial use, modification and redistribution as long as the licence and copyright notice are preserved. The repository is public and the README describes no paid tier or licence key.

### How do I use doccano?

The README describes the flow as creating a project, uploading data, and annotating. It supports text classification, sequence labeling and sequence to sequence tasks, and exposes a RESTful API alongside the web interface. The README says you can build a dataset in hours, and points to the documentation site for details.

### What is doccano used for?

doccano is a text annotation tool for humans, aimed at machine learning practitioners. The README names sentiment analysis, named entity recognition and text summarization as examples of the labeled data you can produce with it.

## Sources

- [doccano/doccano on GitHub](https://github.com/doccano/doccano)
- [Issues](https://github.com/doccano/doccano/issues)
- [License: MIT](https://github.com/doccano/doccano/blob/master/LICENSE)
- [README](https://github.com/doccano/doccano/blob/master/README.md)
- [Releases](https://github.com/doccano/doccano/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/doccano-doccano
