MorphoRuEval-2017: a shared task repository for Russian morphological analysis
morphoRuEval-2017
At a glance
- What is it?
- This repository packages the data, rules, and links for the MorphoRuEval-2017 shared task on Russian morphological parsing. It is a reference point for teams building or evaluating taggers, not a ready-to-run tool.
- Who is it for?
- Adopt this repository if you are building or testing a Russian morphological tagger and need the official shared-task data, the Universal Dependencies standard, and the closed lists for determiners and pronouns. Skip it if you want a maintained Python library or a reproducible pipeline, because the repository is a static archive of links and raw data, with no code to run and no license declared.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Probably not. The repository last received commits 108 months ago, on November 20, 2017.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What this repository actually is
MorphoRuEval-2017 is not a software project. It is a collection of materials for a shared task held at the Dialogue conference in 2017. The repository holds links to training corpora, a morphological standard, closed word lists, and external tools. The README points to a Google Sheets table with the official team results and to a Google Drive folder with the test set and extraction scripts. If you expect a pip-installable package or a command-line tool, this is not it. The primary language is Python, but the repository contains no Python code in the visible root. It is a data hub and a historical record.
The problem it solves: a common benchmark for Russian morphology
Before this shared task, Russian morphological analyzers were hard to compare because each system used a different tagset and different training data. MorphoRuEval-2017 gave teams a single evaluation framework. The repository preserves that framework. It defines the morphological standard, provides training corpora in Universal Dependencies format, and lists the open-source tools that participated. For a researcher, this is the place to find the exact data and rules that produced the published results. For an engineer, it is a way to see which approaches existed in 2017 and where to find their source code.
The morphological standard and the closed lists
The repository includes a folder called morphostandard that defines the annotation rules. Two plain-text files, DET.txt and PRON.txt, contain closed lists of all determiners and pronouns for Russian in Universal Dependencies format. These lists are useful because Russian determiners are a small, finite set, and a tagger can use them as a hard constraint. The file illustration.txt shows examples of the format and tagging. The README also links to kmike/dialog2017, a separate set of scripts to unify data formats into JSON or CoNLL-U. That separation matters: the repository itself does not convert formats, so you must pull in the external scripts.
Training data: three corpora with different access conditions
The repository links to three training corpora. The General Internet-Corpus of Russian (GIKRYA) is in a RAR archive. The Russian National Corpus (RNC) is also in a RAR archive but requires a signed license, which is provided as a PDF in the repository. OpenCorpora is the third corpus. Each is already in Universal Dependencies format, which means you can feed it directly to a UD-compatible tagger. The README also points to plain-text branches: Live Journal with 30 million words, Librusec with 300 million words (password protected), and social media with 50 million words. These are raw text, not annotated, so they are useful for unsupervised pretraining or domain adaptation.
How to get it running: what you actually do
There is no install command. To use the data, you clone the repository, download the RAR archives from the links in the README, and extract them. For the RNC corpus, you must first read and sign the license PDF. Then you need the external scripts from kmike/dialog2017 to convert the data to JSON or CoNLL-U if your tool requires that. The test set and the official extraction scripts are in a Google Drive folder, not in the repository. The README gives no command-line examples, so you must figure out the extraction and conversion flow yourself. This is a manual process, not a turnkey setup.
Limitations and failure modes
The most obvious limitation is that the repository is static. The last push is unknown, and the license is marked NOASSERTION, so you cannot rely on a clear open-source license for the data or the rules. The training corpora are large RAR archives, which may fail to download or extract on systems without the right tools. The RNC license requirement is a real barrier: if you cannot sign it, you lose one third of the annotated training data. The test set is on Google Drive, which can change or disappear, and the repository does not mirror it. Also, the data is from 2017. Russian morphology has not changed, but the evaluation methodology and the UD guidelines may have evolved, so results from this benchmark are not directly comparable to newer shared tasks.
Alternatives and how they differ
The README lists several open-source tools from the participating teams, such as XMorphy, rnnmorph, and MorphoBabushka. These are actual software projects with code you can run. They differ from this repository in that they are implementations, not just data. For example, rnnmorph is a neural tagger, while XMorphy is a rule-based analyzer. If your goal is to deploy a tagger, you should go to those repositories directly. If your goal is to evaluate a tagger on the official benchmark, you still need this repository for the data and the standard. The alternative approach is to use a modern UD treebank for Russian, such as the ones on the Universal Dependencies website, which are continuously maintained and have a clear license, but they are not the exact MorphoRuEval-2017 test set.
Maintenance and licensing concerns
The repository has no recent releases and no license declaration. The README does not state who maintains it or whether updates are planned. The data archives are hosted on GitHub, but the test set and scripts live on Google Drive, which is a maintenance risk. For the RNC corpus, you must sign a license, so the data is not freely redistributable. The other corpora may have their own terms, but the repository does not spell them out. Before using this data in a commercial product, you should contact the organizers or check the original Dialogue evaluation page. The lack of a license means you cannot assume you have the right to reuse the morphological standard or the closed lists without permission.
Editorial conclusion
Adopt this repository if you are building or testing a Russian morphological tagger and need the official shared-task data, the Universal Dependencies standard, and the closed lists for determiners and pronouns. Skip it if you want a maintained Python library or a reproducible pipeline, because the repository is a static archive of links and raw data, with no code to run and no license declared. Before using the RNC corpus, verify that your use case fits the license you must sign, and check the Google Drive folder for the test set and scripts, since those are not version-controlled here.
Frequently asked questions
What does morphological mean in MorphoRuEval-2017?
It refers to Russian morphological annotation. The repository commits a morphological standard and rules as a directory, an illustration file with examples of format and tagging, and closed word lists for determiners and pronouns in Universal Dependencies format, plus four further lists the README leaves unexplained.
What does morphology mean in simple terms in the MorphoRuEval-2017 materials?
In this repository morphology means the label assigned to a word token, and the labels are constrained by closed lists rather than free annotation. DET.txt holds all Russian determiners and PRON.txt all Russian pronouns, both in Universal Dependencies format, with ADP, CONJ, H and PART also committed at the root.
What training corpora does MorphoRuEval-2017 provide?
Three Universal Dependencies corpora: the General Internet-Corpus of Russian as a rar archive, the Russian National Corpus with a one million subset whose licence PDF has to be signed first, and OpenCorpora as a rar archive.
How are the MorphoRuEval-2017 plain text archives protected?
Inconsistently and deliberately. The Live Journal set at 30 million words and the social network set at 50 million words, covering Twitter, VKontakte and Facebook, have no password. Librusec at 300 million words is protected by a password given in the README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dialogue-evaluation-morphorueval-2017)