Open-source project
skishore/makemeahanzi avatar
skishore/makemeahanzi

Make Me a Hanzi: stroke-order vector data for 9000 Chinese characters

Free, open-source Chinese character data

2,702 stars615 forksJavaScriptNOASSERTION

At a glance

What is it?
Two newline-separated JSON files holding stroke paths, medians, decompositions and stroke-to-component mappings for simplified and traditional characters. Where the data comes from, and what the licensing actually says.
Who is it for?
Make Me a Hanzi is a data project rather than a code project, and treating it that way is the fastest way to use it well. dictionary.txt holds everything linguistic, graphics.txt holds everything geometric, and the character field joins them in a guaranteed identical order. The strokes use a coordinate system whose y-axis runs downwards, which is the single detail most likely to break a first implementation if you skip the transform.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 7 months ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the project actually is

Make Me a Hanzi provides dictionary and graphical data for over 9000 of the most common simplified and traditional Chinese characters. The headline feature is stroke-order vector graphics for all of them, which means you get the geometry of each stroke as path data rather than as a bitmap.

That framing matters, because the repository contains almost no application code. The tree at the top level holds COPYING, a dictionary.txt, a graphics.txt, an LGPL file, a stats.py script, directories for stroke caps and for two SVG sets, a README and an example GIF. The demonstration of what the data enables lives at a separate demo site where you can look up a character by drawing it, and the README points to a project site for general information and updates.

The repository is in JavaScript as its language, which reflects the ecosystem it feeds rather than a large codebase in this tree, and there are no tagged releases, so consumers take the files from the repository directly. Stars sit at 2,684 with 616 forks, and 71 issues are open. The most recent push was on 2026-03-08.

Two files split by licence, joined by character

The most consequential design decision is that the data is split into dictionary.txt and graphics.txt, and the reason given is licensing: the files are derived from sources with different licenses. A third artefact, an experimental tarball of animated SVGs called svgs.tar.gz, is licensed the same way as graphics.txt.

That is the reason to read this section before downloading anything. Attributing the whole project to a single licence is wrong here, because the two files are not interchangeable under the same terms. The COPYING file holds the details, and the README is explicit that the Sources section plus COPYING together carry the licensing information.

The join between the two files is simple. Both are newline-separated lists of lines where each line is a JSON object, they differ in which keys are present, and the shared character key joins them. You can also rely on the two files always coming in the same order, which is a stronger guarantee than the shared key alone: streaming them in parallel works without buffering one file to index the other.

The practical consequence for an implementer is that you can ship the graphics without the dictionary if your licence situation requires it, and a reader who wants pinyin and definitions pulls dictionary.txt independently.

Where the data comes from, and why it exists at all

dictionary.txt is derived from Unihan and from CJKlib. graphics.txt and the SVG tarball are derived from two free fonts, Arphic PL KaitiM GB and Arphic PL UKai. The README credits Arphic Technology, a Taiwanese font forge, with releasing their work under a permissive license in 1999, and says plainly that the project would not have been possible without that generosity.

That chain explains both the strength and the limit of the dataset. Deriving stroke outlines from high-quality fonts gives clean, consistent geometry across tens of thousands of glyphs, which is very hard to produce by hand. It also means the outlines inherit whatever the fonts' designers chose for the standard printed form of each character, so this is data about how characters are written in a typeface tradition rather than an independent record of all handwriting.

The README also credits Gábor Ugray for advice on the project and for verifying stroke data for most of the traditional characters in the two data sets, and notes that he maintains Zydeo, a free and open-source Chinese dictionary. Human verification of stroke order against machine-extracted outlines is the part of this dataset that cannot be automated away, and it is why a reader should treat stroke order as reviewed data rather than as derived data.

The fields that make decomposition usable

dictionary.txt has six documented keys and graphics.txt has three. In dictionary.txt, character is required and holds the Unicode character itself. definition is an optional string aimed at second-language learners. pinyin is required but may be empty, holding a comma-separated list of pronunciations. radical is the Unicode primary radical and is required.

decomposition is the interesting one, holding an Ideograph Description Sequence. It is required but invalid if it starts with a full-width question mark, and the README adds a subtlety that a naive parser will trip over: even when the first character is a proper IDS symbol, any component inside the decomposition may itself be a wide question mark, as when a character is recognised as a top and bottom structure but only the top component is known, giving something like a top component followed by a question mark.

etymology may be null, and when present always carries a type of ideographic, pictographic or pictophonetic. The first two types include a hint field explaining the formation. The pictophonetic type adds hint, phonetic and semantic fields, each a string that may be null, and the README gives the reading template explicitly: the semantic part provides the meaning while the phonetic part provides the pronunciation.

matches is the field with the least obvious purpose and the most potential. It maps strokes of the character to strokes of its components, indexed through the decomposition tree, and individual entries may be null. The README walks through a worked example, using a character whose decomposition splits it into a side-by-side pair where the right half is itself a top-and-bottom pair, and shows that its third stroke belongs to the lower component of that inner pair with a match of [1, 0], derived by taking the second child of the root and the first child of that node. The stated use is generating visualisations that mark each component within a character, with the README conceding this could serve more exotic purposes.

Rendering strokes with an inverted y-axis

graphics.txt holds character, strokes and medians. strokes is a list of SVG path data per stroke, ordered by proper stroke order, laid out on a 1024 by 1024 coordinate system where the upper-left corner sits at position (0, 900) and the lower-right at (1024, -124). The y-axis therefore decreases as you move downwards, which the README flags with a note that this is strange and that you should render the paths with a transform rather than drawing them directly.

The recommended rendering is a group element carrying a scale and translate, one path per stroke:

html
<svg viewBox="0 0 1024 1024">
  <g transform="scale(1, -1) translate(0, -900)">
    <path d="STROKE[0] DATA GOES HERE"></path>
    <path d="STROKE[1] DATA GOES HERE"></path>
    ...
  </g>
</svg>

Get that transform wrong and you will see correctly shaped characters mirrored vertically, which is the most common first-run symptom with this dataset.

medians is a list of stroke medians in the same coordinate system as the paths, each median being a list of pairs of integers, and as long as the strokes list. The README is candid about their use: they can produce a rough stroke-order animation, although it is a bit tricky. Medians are the line from stroke start to stroke end, and animating a stroke along its median, then revealing the path, is what a handwriting app needs and what a static dictionary site does not.

The experimental SVG tarball and the named files

The animated SVGs are described under future work as an experimental next step: one image per character, named by the Unicode codepoint of the character, with the codepoint obtainable in JavaScript by calling x.charCodeAt(0) on a character. The minimal embedding example in the README is a single element:

html
<body><embed src="31119.svg" width="200px" height="200px"/></body>

The word experimental is doing real work there. The README's stated reason is that it is still tricky to work with these images beyond the basic example, and it gives a specific unresolved case: it is not clear how to embed two of these images side by side. The tree contains both an svgs directory and a svgs-still directory, which suggests the still frames and the animated versions are maintained separately.

For most uses you do not need the tarball at all, since the medians field in graphics.txt is enough to build your own animation and gives you far more control over timing and interaction. A small stroke_caps directory is also present, which is consistent with hand-drawn rendering where stroke ends need specific treatment.

The open issue count of 71 against 616 forks suggests an active user base filing questions about these details, and the stats.py script at the top level is the kind of utility that would answer questions about dataset coverage.

Editorial conclusion

Make Me a Hanzi is a data project rather than a code project, and treating it that way is the fastest way to use it well. dictionary.txt holds everything linguistic, graphics.txt holds everything geometric, and the character field joins them in a guaranteed identical order. The strokes use a coordinate system whose y-axis runs downwards, which is the single detail most likely to break a first implementation if you skip the transform. On licensing, read COPYING rather than trusting the repository metadata, which reports no detectable licence while the README explains that the two files carry different terms inherited from their sources. For an ink-based writing app, the medians and matches fields are what turn static paths into stroke-order animation and component highlighting; for anything else, the raw paths are enough.

Frequently asked questions

How do you join dictionary.txt and graphics.txt?

Both files are newline-separated JSON objects sharing a character key, and the README states that the two files always come in the same order. You can therefore join them on the character field, or stream them in parallel line by line without building an index of one to look up the other.

Why do my rendered characters appear upside down?

The stroke paths use a 1024 by 1024 coordinate system whose upper-left corner is at (0, 900) and lower-right at (1024, -124), meaning the y-axis decreases as you move downwards. The README says to wrap the paths in a group with transform scale(1, -1) translate(0, -900) to render them upright.

What licence is the Make Me a Hanzi data under?

It depends on which file you use. The project splits the data in two because the sources have different licences: dictionary.txt derives from Unihan and CJKlib, while graphics.txt and the SVG tarball derive from the Arphic PL KaitiM GB and Arphic PL UKai fonts, released under a permissive licence by Arphic Technology in 1999. The COPYING file carries the details.

What are the medians in graphics.txt used for?

A median is a list of integer point pairs in the same coordinate system as the stroke paths, one per stroke, giving the centre line of each stroke. The README says they can be used to produce a rough stroke-order animation, while noting that doing so is a bit tricky.

Official sources

  1. Issues
  2. Project website
  3. README
  4. skishore/makemeahanzi on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/skishore-makemeahanzi.svg)](https://hysenlabs.com/projects/skishore-makemeahanzi)