# adbar/German-NLP: a curated list of German datasets, corpora and tools

> German-NLP is a curated list of open-access German-language resources, from historical corpora to transformer models. It is an index, not a library, and that shapes both its value and its limits.

**adbar/German-NLP** — Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German

- Repository: https://github.com/adbar/German-NLP
- Stars: 535 · Forks: 66
- Language: Unknown
- License: not declared
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/adbar-german-nlp

## What German-NLP actually is, and who it is for

German-NLP is a curated list, in the Awesome-list style, of open-access, open-source and off-the-shelf resources and tools built with a particular focus on German. The README states the selection is deliberately biased toward usability and user-friendliness, and that resources which can be used off the shelf or with minor adjustments and which are currently maintained are primarily chosen. That sentence is the whole editorial policy, and it is worth taking at face value: the list optimises for things you can pick up today, not for exhaustive coverage.

The audience is narrower than the topic list suggests. Someone doing computational linguistics, corpus linguistics or text mining on German text will recognise the shape of the table of contents immediately: text corpora split into general-purpose, historical, specialised and word lists; generic resources such as frameworks, treebanks, deep learning models and transformers, annotation tools and standards; then linguistic processing stages from preprocessing and tokenization through stemming, lemmatization, morphological analysis, normalization, phonology, POS-tagging, parsing and named entity recognition, with sections for industry applications and evaluation. Semantic analysis follows, covering datasets, word embeddings and senses, sentiment datasets and polarity clues, sentiment detection, GermEval, coreference resolution, summarization and simplification, and psycholinguistics. Speech NLP, machine translation, large language models, teaching resources and tutorials, and further lists round it out.

If you are an application developer who wants one German NLP library to import, this is not that. If you are starting a German-language project and do not yet know which corpus, tagger or embedding set exists, the structure alone saves you a day of searching.

## How the list is organised and how entries are chosen

There is no code, no build step and no runtime. The mechanism is a single README.md plus a contributing.md, and the only structure is Markdown headings and bullet lists. Each entry is a name, a link, and occasionally a short parenthetical gloss such as (CommonCrawl) or (14th-16th centuries).

That parenthetical gloss carries most of the editorial signal. Historical corpora are annotated with date ranges, so Anselm is labelled 14th-16th centuries, GerManC 1650-1800, Referenzkorpus Altdeutsch 750-1050, Referenzkorpus Mittelhochdeutsch 1050-1350, Referenzkorpus Frühneuhochdeutsch 1350-1650, and Referenzkorpus Mittelniederdeutsch/Niederrheinisch 1200-1650. For a corpus linguist that is the single most useful field in the list, because period coverage determines whether a corpus is usable for a given research question. General-purpose entries are mostly bare links: Araneum Germanicum, DWDS, the IDS corpora, the Leipzig Corpora Collection, SdeWaC, COW, CEHugeWebCorpus, GC4 Corpus.

Specialised corpora show the same pattern with domain labels: legal material appears as AGB-DE, Legal Entity Recognition and Open Legal Data Corpus; parliamentary material as GermaParl (Bundestag) and German Parliamentary Corpus (GerParCor); social and media material as the One Million Posts Corpus, Pegida Facebook Comments and the Dortmunder Chat Korpus. The bias toward usability shows up in what is absent: no dead links are advertised as dead, no entry carries a maintenance date, and nothing in the README records when an individual resource was last touched. The list's own currency and the currency of its entries are two different things, and only the first is visible.

## Reading the list for a German pipeline

There is nothing to install. The project is a README, so the practical first step is to read the table of contents and follow the links that match your task. The README points contributors to contributing.md for the guidelines on adding or correcting an entry.

Because the whole list is one Markdown file, the only concrete workflow the repository supports is opening that file and searching it. The headings in the table of contents are the search keys: Text corpora, Generic resources, Linguistic processing, Semantic analysis, Speech NLP, Machine Translation, Large Language Models, Teaching resources and tutorials, and More lists. Within Linguistic processing the subheadings run from Preprocessing and Tokenization / Sentence boundary detection through Stemming, Lemmatization, Morphological analysis, Normalization, Phonology, POS-tagging, Syntactical parsing and Named Entity Recognition to Industry/Applications and Evaluation.

A sensible order of work is to fix the task first, then the section. If you need sentiment labels for German reviews, the Semantic analysis section holds both Sentiment analysis datasets / polarity clues and Sentiment detection, and the GermEval subsection points at shared-task material. If you need to parse historical text, the Historical corpus list and the Syntactical parsing entries are the relevant places, and the period labels tell you quickly whether a corpus reaches your text. Everything after that happens outside this repository: open the linked project, read its own README, and check its licence and install instructions before you commit to it.

## Where German-NLP stops being the right tool

The most important limitation is that this is an index with no guarantee attached to any entry. The README says maintained resources are primarily chosen, but it does not state when each linked resource was last checked, and it carries no per-entry status field. A link that worked when it was added can rot without the list changing. Treat every entry as a lead, not a verified dependency.

The second limitation is licensing. The repository shows no licence file among its top-level entries; only README.md and contributing.md are present. Even if the list text itself were permissively licensed, that would say nothing about the corpora and tools it links to, which carry their own terms. Corpora in particular range from openly downloadable to registration-gated to research-only, and the README does not summarise those terms. Anyone planning to ship a product trained on an entry from this list has to check the source project's licence directly.

The third limitation is scope drift. The list is deliberately biased toward usability, which means it is not a systematic survey. A resource that is excellent but awkward to install may simply be missing, and absence from the list is not evidence that nothing exists. For a literature review or a coverage claim, this is the wrong instrument; for a shortlist of things to try this week, it is a reasonable one.

Finally, the list is only as current as its last push, which was on 2026-09-11. The README itself notes that community support is needed to keep the list up to date and that pull requests and suggestions are welcome, which is an acknowledgement that staleness is an ongoing risk rather than a solved problem.

## German-NLP compared with a framework such as spaCy or Stanza

The natural alternative depends on what you actually wanted. If you came for German NLP capability rather than for a reading list, the alternative is a processing framework that ships German models, such as spaCy or Stanza, or a German-specific toolkit reached through this list's linguistic processing sections. The difference in approach is fundamental: a framework is executable code with a version number, a model download and an API you call in a pipeline, while German-NLP is a document that tells you which frameworks, corpora and models exist.

Concretely, a framework gives you a tokenizer, a tagger and a parser behind one interface and one release cycle, and it will tell you what German accuracy to expect on its own evaluation sets. German-NLP gives you pointers to tokenization and sentence boundary detection, stemming, lemmatization, morphological analysis, normalization, POS-tagging, parsing and named entity recognition resources, each maintained separately by different groups, each with its own install path and licence. The trade-off is coverage against integration cost. A framework constrains you to what its maintainers have built and kept working; the list exposes you to resources a framework may never include, such as the historical corpora or the specialised legal and parliamentary collections, at the price of assembling and vetting them yourself.

There is also an evaluation angle. The list has an Evaluation section and a GermEval section, which is useful precisely because German NLP benchmarks are spread across shared tasks rather than concentrated in one leaderboard. A framework will quote its own numbers; the list is where you find out which shared-task datasets those numbers might be measured against.

## Maintenance, licence and the cost of depending on a list

Upgrade cost here is close to zero in the software sense and non-trivial in the editorial sense. There are no dependencies to bump, no migrations and no breaking changes, because there is no code. What you pay instead is the cost of re-checking links and licences yourself, and the cost of noticing when an entry you adopted has moved or been abandoned. The last push to the repository was on 2026-09-11, and the README asks for pull requests and suggestions to keep the list current, so the maintenance model is community contribution rather than a release schedule.

On licensing, the repository as presented contains README.md and contributing.md and no licence file, so the terms under which the list text itself may be reused are not stated. That matters if you plan to mirror or republish the list. It matters far more for the linked resources, because a curated list of open-access and open-source material does not and cannot relicense what it points to. Each corpus, treebank, embedding set and model keeps its own terms, and some of the corpora listed are hosted by universities or national infrastructure with their own access conditions. This is a description of what the repository contains, not legal advice; check the licence of the specific resource you intend to use.

## Conclusion

Use German-NLP when you need to find German corpora, taggers or models and want the search narrowed before you start. Skip it if you need a maintained library with an API, versioned releases and support: this is a README, and its last push was on 2026-09-11. Before relying on any entry, open the linked repository and check its licence and whether it still builds, because German-NLP itself carries no licence file and does not verify the terms of what it points to.

## FAQ

### What is the best tool to learn German?

German-NLP does not recommend language-learning products. It lists resources for processing German text, such as corpora, taggers and models, and it does include a section on teaching resources and tutorials, which is the closest thing to learning material in the list.

### What does NLP stand for?

NLP stands for natural language processing, the field the list covers. The repository topics include natural-language-processing and nlp alongside german-language and text-mining.

### Is there a free German AI tool available?

The list is explicitly a curated collection of open-access, open-source and off-the-shelf resources and tools focused on German, so it is a reasonable place to look for freely usable options. It links to each project rather than hosting anything, so the terms of any individual tool have to be checked at the source.

### Can I learn German in 3 months?

The available information does not address language-learning timelines, and German-NLP is aimed at computational processing of German text rather than at learners.

## Sources

- [adbar/German-NLP on GitHub](https://github.com/adbar/German-NLP)
- [Issues](https://github.com/adbar/German-NLP/issues)
- [README](https://github.com/adbar/German-NLP/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/adbar-german-nlp
