GLOSSOPETRAE: a procedural language engine used as a model probe
LINGUISTIC ENGINE FOR AI
At a glance
- What is it?
- A research program that generates constructed languages from a seed, then uses them to measure how frontier models acquire vocabulary and where their tokenizers disagree.
- Who is it for?
- GLOSSOPETRAE is two projects sharing one engine, and the second one is the more interesting half. The language generator produces a working conlang with phonology, writing systems and translation in a single call, which makes it usable on its own.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 109 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Generating a language from a single seed number
The entry point is one import and one call. A seed integer selects a language, and everything downstream, phonology, morphology, syntax, lexicon, writing system, is derived from it deterministically. The project claims 500 out of 500 generated languages are functional across all of those layers.
import { Glossopetrae, PRESETS } from './src/Glossopetrae.js';
// Generate a language
const lang = Glossopetrae.quick(42);That determinism is what makes the rest of the program possible. If seed 42 always yields the same language, then a K-shot learning experiment can hold the language fixed and vary only the number of examples, which is the entire experimental design in one property. `PRESETS` sits beside the main export for the curated configurations, and the web interface at `index.html` exposes twenty tabs covering the engine, the research data and the offensive modules without a build step.
open index.htmlThe package metadata is unusually spare. There is no dependency block, no bundler configuration and no transpilation step, so the engine runs directly as an ES module in a modern browser or Node runtime. That is a deliberate constraint: with 25 modules in `src/` and no framework underneath, everything you read in the source is the actual implementation.
Translating English into a language the model has never seen
The translation API is where the conlang stops being a toy and becomes an instrument. Given a source sentence, the engine returns a structured result with a `target` field holding the constructed language output.
const result = lang.translationEngine.translateToConlang("The warrior sees the mountain.");
console.log(result.target);Because a procedurally generated language has no training data anywhere, no model has seen it. Ask a frontier model to translate into it cold and it produces nothing useful. The reported K-shot table shows accuracy at 0 percent for K of zero, one and three, then 95 to 100 percent at eight examples. The project calls this the acquisition cliff, and the shape matters more than the headline number: nothing happens for several steps and then almost everything happens at once.
Nine frontier models are listed in that table, from Opus and GPT through Gemini, Haiku, Qwen, Grok, Gemma, GLM and DeepSeek. All show the same pattern, which is the strongest evidence in the repository. A finding that appears in every family tested is more likely to be a property of in-context learning than of any one vendor's tuning.
Harder for humans turned out to be easier for models
The second experiment inverts the intuition behind the first. Language difficulty levels run from L0, which a human can read, to L3, which is opaque, and the reported task accuracy goes the wrong way from what you would guess. At L0, Opus scores 20 percent and GPT 53 percent. At L3, Opus scores 97 percent and GPT 100 percent.
Cold readers, models given no key at all, score between 17 and 40 percent, which is the control that makes the result meaningful: the difficulty is real and it does work against a human or an uninstructed model. A factorial design crossing key possession with task specification over 1800 trials puts the key possession effect at roughly 95 percent and the specification effect at null.
The practical reading is that human readability and machine usability are separate design spaces. If you are building a machine-facing protocol and you have been keeping it readable for operators, you may have been optimising for the wrong audience. That is a genuinely useful result, and the project is candid that it is the kind of finding that should not be true.
Tokenizers disagree, and the gap is a channel
The offensive half of the program follows from the acquisition result. Different model families process Unicode differently, so there exist characters one tokenizer keeps and another silently discards. Hide a message in those characters and a monitor model reading the text never receives the payload at all, because it is deleted before inference starts.
The reported Haiku to GPT tag character result is complete delivery to the receiver with zero detection by the monitor across 30 trials, McNemar p below 10 to the minus 7. The PUA channel from Gemini to Haiku runs at 93 percent delivery with zero detection. The 12-model blind spot map claims 255 exploitable instances across 49 of 66 possible model pairs, concentrated at the GPT and Claude boundary.
The detection flip is the number to remember. On a clean test with negative controls, a tokenizer-blind monitor scores Youden's J of zero, meaning chance, and silently clears 15 out of 15 hidden payloads. Give the same monitor the decode key and its J is 1.0. The monitor is not being deceived by a clever prompt, it simply cannot see the channel.
Where the project admits its own errors
A credibility section in the same README reports that the first draft of the paper contained fabricated numbers, specifically a claim of 23 blind spots that was actually 3, caused by a harness that scored empty completions as stripped. The error was found by adversarial self audit and is documented in a section of the paper rather than quietly fixed.
That disclosure is worth more than it might appear. A measurement harness that silently miscounts is the normal failure mode of this kind of project, and reporting it sets the standard for everything else on the page. The repository backs the claim with structure: `experiments/results/` holds raw JSON for 78 result files, `bench/` holds the benchmarks, `redteam/` holds the adversarial material, and `VALIDATION.md` plus `AGENT_QUICKSTART.md` sit at the root.
{
"name": "glossopetrae",
"version": "3.1.0",
"type": "module",
"private": true
}Note `private: true`. The package is not published to a registry, which is consistent with a research release rather than a library. The test script chains three harnesses in sequence, and a dozen more test files at the root cover audio, glyphs, evolution, reverse mapping, drift and translation at five different versions.
What a reader should take from a dual use release
The README labels this a research release with offense and defense capabilities and names the modules directly: SemanticStego, SteganographyEngine, TokenExploiter and LanguageAttributes, the last covering phantom and adversarial language attributes. The semantic steganography result is the one that matters most and the least resolved. It uses ordinary readable English, evades every monitor tested at 27 percent detection or below, and the stated conclusion is that no two line fix exists.
Multi construction resilience is reported for all six cells surviving Bonferroni correction, with three structural detectors refuted at J of zero and a perplexity anomaly detector at 0.58, which the project describes as its best and still unreliable.
The licence is AGPL-3.0, which is the right choice for a project like this and worth reading carefully before reuse. The language is JavaScript, the default branch is `main`, and the repository has 1041 stars, 228 forks and 4 open issues with the last push on 2026-06-20. With a claim this strong, the paper at `PAPER.md` and `PAPER.html` is where the arguments actually live, and the raw JSON is where you should check them.
Editorial conclusion
GLOSSOPETRAE is two projects sharing one engine, and the second one is the more interesting half. The language generator produces a working conlang with phonology, writing systems and translation in a single call, which makes it usable on its own. Attached to that is a measurement apparatus that asks frontier models to learn a language nobody speaks, and then exploits the fact that reading a language and understanding it come apart for machines. The honest caveat is that the offensive modules are described as research grade and dual use, so treat the repository as something to study rather than to deploy. It sits at version 3.1.0 as a private ES module package under AGPL-3.0, with a last push on 2026-06-20. Start with the bundled paper, then run `node test.mjs` and read `VALIDATION.md` before drawing conclusions from any single number.
Frequently asked questions
What is the GLOSSOPETRAE project used for?
It is a procedural language generator that builds a complete constructed language from a single seed, covering phonology, morphology, syntax, lexicon, writing systems and audio. On top of that engine it runs research experiments on how frontier models acquire unfamiliar languages in context, and on where different model tokenizers disagree.
Does GLOSSOPETRAE require a build step or any dependencies?
No. The package declares no dependencies and runs as a plain ES module, and the web interface is opened directly from index.html in a browser. The test script is a chain of plain node invocations, so nothing is compiled before you can use it.
Can I audit the numbers in the GLOSSOPETRAE paper?
The raw data is meant to be re-derived rather than taken on trust. The experiments directory holds 78 raw result JSON files, the paper documents an earlier draft error that the authors caught in an adversarial self audit, and a VALIDATION.md file at the repository root covers the correction procedure.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/elder-plinius-glossopetrae)