PyGlossary: Converting Offline Dictionaries Between StarDict, DSL, MDict and More
A tool for converting dictionary files aka glossaries. Mainly to help use our offline glossaries in any Open Source dictionary we like on any operating system / device.
At a glance
- What is it?
- PyGlossary is a Python 3.12+ dictionary converter with read and write plugins for StarDict, Babylon BGL, DSL, MDict, Kobo, Zim and dozens more. It is a format bridge, not a dictionary reader, and its behaviour depends on which plugins you install.
- Who is it for?
- PyGlossary fits people who already own a glossary file and need it in a format their reader or e-reader accepts: install it with pip, check the format table for whether the direction you need is marked as readable, writable, or both, and run a conversion on one file before committing to a batch. It is the wrong tool if you want a dictionary application to look words up in, or if your source format is one the table marks as read-only and your target is not listed at all.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap PyGlossary fills between dictionary formats
Offline dictionaries are stored in a long tail of incompatible formats. StarDict uses a directory of .ifo, .idx and .dict files. ABBYY Lingvo uses .dsl text. Octopus MDict uses a binary .mdx. Kobo e-readers use their own archive, PocketBook another, Aard 2 a .slob file. The README states the primary purpose plainly: to use "our offline glossaries in any Open Source dictionary we like on any OS/device." So the audience is not people who want a dictionary app. It is people who already have a glossary file in one format and need it in another, usually because the reader or e-reader they prefer only accepts one shape.
The scope is deliberately uneven. The README says there are countless formats and the author's time is limited, so formats are implemented when they seem useful to the author or to the open source community, with language diversity taken into account. That explains why the table includes Japanese, Korean, Chinese, Arabic and Russian sources alongside generic ones, and why some cells say a format will not be supported at all. Treat the format table as the specification of what the project is for.
Read plugins, write plugins and what the format table actually promises
Every conversion is a pair: a read plugin parses the source into an internal glossary, and a write plugin serialises it. The table marks each format with a check or a cross for read and write separately, and the crosses are not all the same. Some are simply not implemented yet; others carry the note that they will not be supported. AppleDict Binary, Dict.cc translation export, FreeDict, Wiktextract, XDXF, Zim and EPWING are read-only in the table. EPUB-2, Kobo, Mobipocket, PocketBook SDIC, SQL, DIKT JSON and HTML Directory are write-only. StarDict, Babylon BGL, CSV, Tabfile, Lingoes Source, QuickDic and Yomitan are bidirectional.
That asymmetry is the first thing to check before installing. If your source is an .mdx and your target is a Kobo dictionary, both directions exist and the conversion is plausible. If your source is a Zim archive and your target is Zim, there is nothing to do. The table also flags SQLite-backed formats such as Almaany, Dict.cc, DigitalNK and cc-kedict with a database icon and a warning: they are not detected by the .db extension, so you must select the format explicitly with the UI or the --read-format flag. The README adds a second warning not to confuse those SQLite-based formats with SQLite mode, which is a separate feature.
Installing PyGlossary and converting a StarDict file
The package requires Python 3.12 or higher and declares no mandatory runtime dependencies, so a plain pip install gives you the command-line interface. The README points at ./main.py --help or pyglossary --help as the entry points.
pip install pyglossary
pyglossary --helpThe optional extras are where the heavier format support lives. The wxPython wizard installs as pyglossary[wx], and the project also declares qt6, slint, all, tk-wizard-dnd and dict-cc-source extras. The all extra pulls PyICU, lxml and beautifulsoup4. The repository also ships a Dockerfile based on bitnami/minideb that copies the tree to /opt/pyglossary, runs scripts/docker-deb-setup.sh, and starts with python3 /opt/pyglossary/main.py --cmd, with a run-with-docker.sh helper at the top level.
Once installed, the command-line interface is what you drive a conversion with. The README gives pyglossary --help as the way to see the available options, and the format table tells you which read and write formats your source and target correspond to. The bundled pyglossary-diff script, declared in pyproject.toml, compares two glossaries and is the quickest way to sanity-check that a conversion kept the entries you expected.
Where the conversion model breaks down
A glossary format is not just a container. DSL, XDXF and TMX carry markup, article structure and sometimes morphology or frequency data. Tabfile and CSV carry almost none. Converting from a rich format to a flat one is lossy in a way the tool cannot warn you about, because the writer simply has nowhere to put the extra structure. The direction that loses least is text-to-text between formats with comparable markup; the direction that loses most is anything into Tabfile.
Binary formats are the other failure mode. MDict .mdx, Babylon .bgl, PocketBook .dic and AppleDict .dictionary are read or written by parsers that have to cope with files produced by tools the project does not control. The README does not document rollback, partial-write recovery, or what happens when a write plugin fails halfway through a large glossary, so a batch conversion should be run on one file first and the output opened in the target reader before the rest are processed.
Finally, PyGlossary is not a dictionary reader. It has GUI front ends and a web interface, and it ships pyglossary-view, but the stated purpose is conversion. If you want to look words up, you want the dictionary application you are converting for, not this.
PyGlossary compared with format-specific converters
The obvious alternative is the single-format converter: a script or utility that turns one format into one other format. Those exist for popular pairs, and they have an advantage PyGlossary does not. A dedicated StarDict-to-Kobo converter can be written against the two specifications alone, with no plugin registry, no format detection heuristics and no optional dependency tree. It will be smaller and easier to audit.
The difference in approach matters at the edges. PyGlossary puts every format behind the same read/write interface and routes everything through one internal glossary representation, which is why a conversion from Lingoes Source to Yomitan needs no new code, only the two plugins. The cost is that the internal representation has to be general enough for all of them, and generality is exactly where markup gets flattened. A dedicated converter can preserve every field of the one format it knows about. PyGlossary can move between any two formats in its table, and the table is the honest statement of which moves are possible.
Licence, packaging and the cost of keeping up
PyGlossary is GPL-3.0-or-later, stated in pyproject.toml and shipped as a LICENSE file at the repository root. That matters if you plan to embed the conversion code in a closed product rather than call the command-line tool; the licence dialog and about files in the repository are part of how the project presents this, but the terms themselves are the usual GPL ones and this is not legal advice. The dependency list in requirements.txt is broad, including PyICU, lxml, marisa-trie, libzim and pymorphy3, but those are not mandatory at install time, which keeps the base install small and pushes the cost into whichever plugins you actually use.
Upgrade cost depends on how many formats you touch. The project has shipped 5.4.0, 5.4.1 and 5.4.2 between 2026-05-26 and 2026-06-30, and the last push to the default branch was on 2026-09-18, so the plugin surface is still moving. Version 5.4.2 is the current release. Because formats are plugins, a change to one plugin does not force a migration of your whole workflow, but it does mean a pinned version is worth considering if you convert the same file type on a schedule. The setup.py reads the version from git describe unless NO_GIT_VERSION is set, so installs from a git checkout and installs from PyPI can report different version strings.
Editorial conclusion
PyGlossary fits people who already own a glossary file and need it in a format their reader or e-reader accepts: install it with pip, check the format table for whether the direction you need is marked as readable, writable, or both, and run a conversion on one file before committing to a batch. It is the wrong tool if you want a dictionary application to look words up in, or if your source format is one the table marks as read-only and your target is not listed at all. Verify first that the specific plugin pair you need exists, since the table shows several formats that can be read but never written, and that your Python is at least 3.12.
Frequently asked questions
How do I install PyGlossary?
Install it with pip, which requires Python 3.12 or higher. The base package has no mandatory dependencies; GUI front ends and some formats come from extras such as pyglossary[wx], qt6 or all.
How do I use PyGlossary to convert a dictionary?
Run the command-line interface and pass it the options shown by pyglossary --help, selecting the read and write formats from the table. For a StarDict directory, the input is the .ifo file that anchors it.
Which dictionary formats can PyGlossary read and write?
The README's format table lists read and write support separately. StarDict, Babylon BGL, CSV, Tabfile, Lingoes Source, QuickDic and Yomitan are bidirectional, while formats such as MDict, Zim, XDXF and EPWING are read-only and EPUB-2, Kobo, Mobipocket, PocketBook SDIC and SQL are write-only.
Why does PyGlossary not detect my .db dictionary file?
SQLite-based formats such as Almaany, Dict.cc, DigitalNK and cc-kedict are not detected by the .db extension, so the README says you must select the format with the UI or the --read-format flag. It also warns not to confuse these with SQLite mode.
Does PyGlossary have a graphical interface?
Yes. The README lists Gtk3, Gtk4, Qt6, wxPython and Tkinter interfaces plus a web interface, alongside the command-line and interactive command-line modes. Gtk3 is described as the best option for most Linux users.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ilius-pyglossary)