Model or dataset
toon-format/toon avatar
toon-format/toon

TOON: Token-Oriented Object Notation for Compact LLM Input

🎒 Token-Oriented Object Notation (TOON) – compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.

25,437 stars1,123 forksTypeScriptMIT

At a glance

What is it?
TOON is a TypeScript library and CLI that encodes JSON data into a compact, human-readable format designed to minimize the token count when the data is sent to a language model. It uses indentation for nested objects and a tabular form for arrays of uniform objects, achieving CSV-like compactness while retaining explicit field names and row counts.
Who is it for?
TOON is worth trying when your application feeds arrays of uniform objects into a language model and API token costs are a meaningful factor. The format is stable at v4.1.1 but the spec is described as an idea in progress.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 27 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TOON Solves and Who It Is For

When JSON data is sent to a language model, structural characters consume tokens without carrying information content: curly braces, brackets, colons, and repeated key names on every row of an array all add to the token count and therefore to API cost.

TOON encodes the same JSON data model using a different syntax. For arrays of objects that share the same fields (uniform objects), TOON declares the field list once in a header row and then streams one row per object, similar to CSV but with explicit structure markers. The README shows a weather forecast example where the same data takes approximately 117 tokens as JSON and approximately 66 tokens as TOON, a reduction of about 44 percent.

The format is aimed at developers who already have JSON data and want to reduce the token cost of including that data in LLM prompts without losing the ability to programmatically round-trip the data back to JSON.

The Four Forms and When They Apply

TOON defines four forms for rendering a value, chosen based on the data's shape.

The inline form places a primitive array on its header line. In the weather forecast example from the README, `alerts[2]: frost,wind` uses this form.

The tabular form is the primary compression mechanism. For an array of objects with identical keys, the field list is declared once in a header, followed by one row per element. The README shows `forecast[3]{day,temp{min,max},condition,rainChance}:` as the header, with three rows following.

The keyed tabular form handles objects whose values are uniform objects, using a colon after the length bracket `[2:]` to mark it. Each row carries its own key. The README shows an environments config map where production and staging entries are compressed to two rows under a shared field header.

The list form is the fallback for data that fits none of the above patterns: mixed types, non-uniform objects, or empty values. One `- ` item per element, or a bare `-` for an empty object.

Installation and the CLI

TOON is published to npm under the @toon-format scope. No installation is required to try the CLI:

bash
cat data.json | npx @toon-format/cli --stats

This pipes a JSON file through the TOON encoder and prints the TOON output alongside a token estimate comparison. The `--stats` flag adds the savings summary. The README shows this output for the weather forecast example:

code
Token estimates: ~117 (JSON) → ~66 (TOON)
Saved ~51 tokens (-43.6%)

For programmatic use, the TypeScript SDK is the @toon-format/toon package. The monorepo in the repository also contains the spec and benchmark packages. The overall monorepo uses pnpm workspaces with the packages directory holding the individual library packages.

Benchmark Results and Accuracy Trade-offs

The README documents two benchmark tracks. The Mixed-Structure Track tests nested and semi-uniform datasets against JSON, YAML, and XML. CSV is excluded from this track because it cannot represent nested structures without lossy flattening. The Flat-Only Track includes fully tabular datasets and adds CSV as a competitor.

The README states that TOON matches JSON's retrieval accuracy while using 42.6 percent fewer tokens, based on 244 data retrieval questions across four language models on a benchmark dataset. The benchmark directory in the repository holds the test data and results.

The README gives explicit guidance on when TOON is not the right choice: deeply nested or non-uniform structures where tabular eligibility is close to zero; semi-uniform arrays where only 40 to 60 percent of records share the same fields; purely tabular data where CSV is smaller; and latency-sensitive deployments where some models process compact JSON faster despite the higher token count. The README notes that latency measurement must be done on the user's own setup.

When TOON Is the Wrong Format

The README section titled When Not to Use TOON is direct about the format's boundaries. When the data structure is deeply nested or non-uniform, with tabular eligibility near zero percent, compact JSON often wins outright because TOON's header syntax adds overhead that the tabular rows don't offset.

For arrays where only 40 to 60 percent of elements are eligible for the tabular form, the savings shrink. The README suggests staying on JSON when the pipeline already uses it and the gains are small.

For purely tabular data, CSV is smaller. TOON adds approximately 5 to 10 percent overhead compared to CSV on flat data because it includes declared lengths, field lists, and delimiter scoping. The README frames this as a reliability trade: the structural metadata helps a language model detect truncation or malformed output, which CSV does not provide.

Latency is a case the README flags but cannot quantify generically. Some deployments with local or quantized models process compact JSON faster than TOON despite the lower token count, so end-to-end timing matters more than token counts alone in those environments.

Spec Status and Ecosystem

The TOON format spec is hosted in a separate repository at toon-format/spec. The README describes the format as stable but also as an idea in progress, explicitly noting that nothing is set in stone and inviting contributions to the spec.

The README lists an official implementations section and mentions dozens of community ports targeting one spec with a shared conformance test suite. The media type for TOON files is documented in a SPEC.md file at the root of the repository, which also covers the file extension.

The monorepo's package.json is at version 4.1.1 and uses pnpm 11.17.0. The TypeScript version is 6.0.3 as listed in the devDependencies. The automd tool is used to keep benchmark results in the README synchronized with the benchmarks directory output. The last push was on 2026-09-03, and the most recent release is v4.1.1, published on 2026-08-05.

Editorial conclusion

TOON is worth trying when your application feeds arrays of uniform objects into a language model and API token costs are a meaningful factor. The format is stable at v4.1.1 but the spec is described as an idea in progress. For deeply nested or non-uniform data, the README itself says JSON may be more efficient. Measure on your own data before committing to the format, since the savings depend on how uniform your structures are.

Frequently asked questions

Is TOON the same as CSV?

No. TOON's tabular form looks similar to CSV for flat data, but it adds declared field names, row counts, and support for nested objects. CSV cannot represent nested structures without flattening them lossy. TOON is also typically 5 to 10 percent larger than CSV on purely flat data because of this structural overhead.

What is TOON?

TOON (Token-Oriented Object Notation) is a compact, human-readable encoding of the JSON data model that reduces token count when data is sent to a language model. It uses indentation for nested objects and a tabular form for uniform arrays, and supports lossless round-trips back to JSON.

What is an alternative to JSON for LLMs?

TOON is one alternative, optimized for arrays of uniform objects. YAML is another human-readable format that reduces some of JSON's structural verbosity. The README notes that for deeply nested or non-uniform data, JSON may be more efficient than TOON itself.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. toon-format/toon on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/toon-format-toon.svg)](https://hysenlabs.com/projects/toon-format-toon)