# Prosodic's metadata claims Apache-2.0 while the repository reports GPL-3.0

> Prosodic parses poems into a linguistic hierarchy, runs a constraint-satisfaction metrical parser over it and reports stress patterns, foot schemes and named rhyme forms, for English and Finnish, with a hosted web app on top. The engineering is careful, with a lazy tree over a per-syllable DataFrame and two parser tiers you can choose between. The packaging is not: the licence recorded in the distribution metadata and the licence recorded by the repository do not match, and neither does the homepage.

**quadrismegistus/prosodic** — Prosodic: a metrical-phonological parser, written in Python. For English and Finnish, with flexible language support.

- Repository: https://github.com/quadrismegistus/prosodic
- Website: http://quadrismegistus.github.io/prosodic/
- Stars: 303 · Forks: 47
- Language: Python
- License: GPL-3.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/quadrismegistus-prosodic

## Two licences and three homepages, none of them reconciled

Start with the thing that will block a legal review. The repository metadata reports the project under a copyleft licence, and the root contains a licence file with a markdown extension rather than a plain one. But the packaging metadata says something different: the project table declares the licence as the text of the Apache License, Version 2.0, and the classifier list includes the OSI approval line for the Apache Software License. That is what a pip install publishes about itself, and it does not match what the repository says about itself. A second discrepancy sits alongside it. There are three distinct project locations in circulation: the repository metadata names a GitHub Pages site, the README links a hosted application and its documentation at a different domain, and the packaging metadata lists the GitHub repository as the homepage and yet another site as the project home. None of this is fatal, but a library that ships conflicting licence metadata is a library nobody can clear without reading the source.

## Thirty unpinned dependencies behind a Beta classifier

The runtime dependency file is a flat list of about thirty names with no version constraints on any of them, the single exception being one hashing library pinned to a major range. The list tells you what the project actually does. Text cleaning and encoding repair, language detection, a dataframe library, an IPA phonetic G2P, a multiset container, the natural language toolkit, a phonemizer, a fast JSON codec, a distance metric, a logging library, a notebook kernel, an ASGI web stack with multipart support, SciPy, requests, a progress bar, Arrow, tabulation, and a grammar-of-graphics plotting library. Two commented-out entries and a commented dependency block show the same candour as elsewhere: one phonology toolkit is deliberately left out and another is commented with a git URL. Meanwhile the packaging metadata classifies the project at development status four, Beta, and requires Python 3.10 or newer with the version read dynamically out of a separate file. A Beta library with unpinned dependencies is fine for research and painful for anything you intend to run twice.

## Two parser tiers, and the heavy one pulls in torch

The optional dependency groups are the clearest piece of design in the packaging, because each one is labelled with what it costs. The first extra is a syntax parser based on spaCy, described in a comment as the default phrasal-stress engine and marked light. The second extra is a constituency parser based on Stanza, described as the faithful engine and marked heavy, with an explicit note that it pulls in torch and a large constituency model. The documentation goes further and gives the combination command for someone who wants both. So there is a deliberate accuracy-versus-footprint choice, and it is opt-in rather than forced. The caching story is handled with the same care. The hashing library's best extra adds two storage engines for the constituency parse cache, described as single-file and fast, and the comment states that the cache degrades to a dependency-free tree-plus-compression alternative when those are absent. A library that documents how its cache gets worse without an optional package is behaving well.

## The tree is lazy and the truth is a per-syllable DataFrame

The data model is stated in one sentence that tells you how to read the whole library: Prosodic organises text into a tree of linguistic entities, from text through stanza, line and word down to syllable and phoneme, and the children are constructed lazily on first access because the underlying source of truth is a per-syllable DataFrame. So the object hierarchy you traverse in Python is a convenience layer over a table, not the other way round. The printed row for a word token shows how much indexing is carried on every entity: paragraph, line, sentence and line-part numbers, the word token number, the token text and its type, and a language code. That last field is the one that matters for a multilingual parser, since each token records which language it was tagged as rather than assuming one language for the whole document. An attribute shortcut is documented too, so the first line is reachable without indexing. Both pandas and Arrow appear in the dependency list, which is consistent with a tabular core.

## Stress is a leading apostrophe in the phoneme string

The syllable table in the quickstart is the most instructive output in the README, because it shows how stress is represented and it is not obvious. Take one word at the start of a line. It produces two rows, not one, distinguished by a word-form number. Both rows record the pronunciation as coming from a dictionary. In the first row the syllable's IPA string is the bare phonetic form. In the second row the same phonetic form appears with a leading apostrophe. That apostrophe is the stress marker, so stressed and unstressed realisations of an identical syllable are stored as two distinct rows rather than as one row with a flag. It also means the phoneme string itself is doing two jobs, carrying the phone and carrying the prominence, which is compact and slightly awkward for anyone who wants to query stress without stripping a character. The same table exposes a flag marking whether the token is punctuation, which is how the foot counts in the scansion avoid counting commas as syllables.

## The quickstart sonnet is cut off and the form is inferred anyway

Look closely at the worked example and you find a small demonstration of what the parser actually claims to do. The text passed in is the first thirteen lines of a Shakespearean sonnet plus the opening of the fourteenth, so the closing couplet is missing. The printed scan shows fourteen lines, with rhyme letters running a, b, a, b, then a gap, c, a gap, c, then d, d, e, e, which is twelve letters. The estimated schema underneath nevertheless reports the full rhyme pattern of fourteen letters, abab cdcd efefgg. So the classifier inferred the form from partial evidence rather than from a complete poem. Two lines also visibly break the pattern: one has a foot parse unlike its neighbours, and another has six feet and eleven syllables instead of five and ten, and neither stops the classification. That is what a constraint-satisfaction parser buys you, tolerance of local irregularity inside a global form, and it is also the thing to check if you need a strict reading.

## Releases skip odd numbers and land two days apart

Three releases are visible in a four-day window in July 2026: v3.6.0 on the sixth, v3.8.0 on the eighth, and v3.10.0 on the tenth. Two patterns jump out. The first is that the minor numbers go up by two each time, so the odd minors are not published as releases and the project appears to keep a stable-only line with the odd numbers reserved for development. The second is the cadence: three releases in four days is not a slow-moving academic library, and it suggests a changelog discipline that a research codebase often lacks. The last push to the default branch, which is master rather than main, was 2026-09-24, so development continues after the July releases. The version number itself is not stored in the package configuration but read out of a dedicated version file at the root, so there is one place to change it.

## Phonemising anything outside the dictionary needs espeak installed

One system dependency is mandatory and it is documented per platform, which is friendlier than most. The library needs the espeak text-to-speech engine to phonemise words that are not in the pronunciation dictionary, so on macOS it is installed through the package manager, and on Linux the command installs the engine together with its library and development headers.

```bash
brew install espeak
```

```bash
apt-get install espeak libespeak1 libespeak-dev
```

On Windows the instruction is to download a build from the engine's own releases. That is the whole dependency story: a Python package and a system synthesiser, with the synthesiser mattering exactly when a word is out of dictionary, which is to say on proper nouns, on archaic spellings and on anything from a corpus older than the dictionary. For a poetry parser that is not an edge case, it is the main case, so on Windows this is a manual step rather than an installer dependency. The rest of the install is conventional.

```bash
pip install prosodic
```

## Conclusion

Prosodic fits a computational linguist or a literature scholar who needs scansion output they can inspect line by line rather than a black-box metre label, and who is willing to install a thirty-package dependency set and a system speech synthesiser. It does not fit anyone whose compliance process needs a single unambiguous licence, because the distribution metadata and the repository disagree about it, and it does not fit anyone who needs reproducible installs, since almost none of the runtime dependencies carry a version constraint. Before you depend on it, resolve the licence question against the actual licence file rather than either metadata source, decide whether the light parser or the heavy one is the right trade for your accuracy needs, and confirm the external text synthesiser is available on the platform you deploy to, since phonemising any word outside the pronunciation dictionary needs it.

## FAQ

### What does the prosodic library do?

It parses text into a linguistic hierarchy of text, stanza, line, word, syllable and phoneme, runs a constraint-satisfaction metrical parser over it, and identifies stress patterns, foot and syllable schemes, and named rhyme schemes. Form classification covers sonnet variants, couplet and ballad among others.

### Which languages does prosodic support?

English and Finnish are the documented targets, with flexible support for others. Each word token records the language it was tagged as, so a single document can carry tokens assigned to different languages.

### What does prosodic need installed besides the Python package?

The espeak text-to-speech engine, which phonemises words that are not in the pronunciation dictionary. It is installed through the package manager on macOS and through apt on Linux, and downloaded from the engine's releases on Windows.

### Can I use prosodic without installing a heavy parser?

Yes. The syntax extra installs a lightweight spaCy dependency parse described as the default phrasal-stress engine, while a separate constituency extra installs a Stanza parse described as the faithful engine and marked heavy because it pulls in torch and a large model. You can install both together if you want.

### What licence is prosodic released under?

The sources disagree and the packaging metadata is what an install publishes. The distribution metadata declares the Apache License, Version 2.0 and carries the matching OSI classifier, while the repository metadata reports a copyleft licence, and a licence file sits at the root of the repository. Resolve this against the licence file before use.

## Sources

- [License: GPL-3.0](https://github.com/quadrismegistus/prosodic/blob/master/LICENSE)
- [Project website](http://quadrismegistus.github.io/prosodic/)
- [quadrismegistus/prosodic on GitHub](https://github.com/quadrismegistus/prosodic)
- [README](https://github.com/quadrismegistus/prosodic/blob/master/README.md)
- [Releases](https://github.com/quadrismegistus/prosodic/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/quadrismegistus-prosodic
