Open-source project
FeiZhuLulu/real-api-pricing avatar
FeiZhuLulu/real-api-pricing

Real API Pricing: subscription fee divided by usable tokens

Real API pricing: subscription fee divided by usable tokens, with allowance, unit price and leaderboard Pareto charts.

826 stars34 forksHTMLMIT

At a glance

What is it?
Real API Pricing turns monthly subscription fees into a $/MTok figure, and the README is explicit that the conversion rests on one project-wide workload assumption. The data is the point; the assumption is the caveat.
Who is it for?
Adopt it if you are comparing subscription plans on a per-token basis and you accept a fixed 97.5% cache-read / 2.15% fresh input / 0.35% output workload, because that convention is what makes the numbers comparable at all. Do not adopt it if your traffic is output-heavy or cache-write-heavy: the README states cache writes are not modeled separately, so converted token allowances may be overstated where a provider charges for them.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the $/MTok figure actually answers

Most pricing pages quote a per-token rate for metered API calls. Subscription plans do not work that way. You pay a flat monthly fee and get a pool of usable tokens whose size depends on the model you pick, the plan tier, and the provider's own definition of a month. Comparing a $20 subscription against a metered API rate is therefore not a subtraction; it is a unit conversion, and the README's formula states the conversion directly: real unit price equals monthly subscription fee divided by monthly usable tokens.

The project is for people doing that comparison in public or in a procurement document. It is not a library you import and not a service you call. The repository holds a data snapshot, a set of derived point files, and a chart pipeline that renders SVG and PNG. The README points readers to an interactive site at real-api-pricing.vercel.app, and the underlying numbers are downloadable as data/adopted.csv, derived/points.csv and derived/points.json. The audience is narrow on purpose: anyone who needs a defensible $/MTok number for a subscription plan rather than a vendor's marketing page.

The standard workload is the whole method

Every conversion in the project runs through one fixed token mix, stated in the README as 97.5% cache reads, 2.15% fresh input, and 0.35% output. That mix exists because dollar pools and three-part token prices cannot be compared otherwise. A provider that gives you a credit balance and a provider that gives you separate input, output and cache rates are describing different things, and the project forces both onto the same axis by assuming a workload.

The README is unusually direct about what this is: a comparison convention, not a claim about any provider's actual workload. That sentence is the most important one in the repository. It means the ranking is valid within the convention and says nothing about your traffic. The project also declines to normalize measurements that already report total tokens. Dashboard back-calculations, local usage logs, controlled saturation tests, and official absolute-token tables are taken as they are. Where only a total and a cost-weighted percentage are available but the token-type split is unknown, the README states the observed total is retained and the limitation is recorded rather than inventing a split. That is a defensible choice, and it means the dataset is not uniformly derived.

Snapshot discipline and the v4.3 index break

The adopted snapshot is dated 2026-09-09. The README notes that AA Intelligence moved to Intelligence Index v4.3, announced 2026-09-07, while AA Coding Agent remains v1.4. It then makes a point that reviewers of this kind of data usually get wrong: the new intelligence methodology replaces the old snapshot as a whole, and lower numerical scores are not evidence of model regression across index versions. If you pull two snapshots and plot a trend line, you are plotting a methodology change.

Each row is one plan crossed with one actually served model, and the README states plainly that allowances of different models under the same plan are alternatives and must not be added together. That single rule invalidates most naive spreadsheet work. The coverage table lists 202 adopted plan-model points, of which 188 carry a monthly allowance, plus 13 metered API baselines. Score coverage is uneven by design: Code Arena and Agent Arena cover 136 and 140 points, AA Intelligence 173, AA Coding Agent 71, OpenDesign Arena 70, Terminal-Bench 4.0 70. Each chart uses scores from its named leaderboard only, and the README warns that Code Arena here means the WebDev Overall Arena Score specifically, not general coding ability.

Installing and rendering the charts locally

The repository is a static site plus a render pipeline, not a package you install into an application. package.json declares it private with the MIT license field set, one script named render that runs node scripts/render_svg.cjs, and a single dependency on sharp. requirements.txt pins matplotlib to a range of 3.8 up to 4. There is no published npm package name to install and no server component.

Clone the repository and install the Node dependency first. The README does not document a specific install command, but the dependency list in package.json is the source of truth:

bash
npm install

That resolves sharp, which the SVG render script needs. If you plan to run the Python side of the pipeline, install the pinned matplotlib range as well:

bash
pip install -r requirements.txt

With dependencies in place, the render script is invoked through the package script name declared in package.json:

bash
npm run render

Expect the script to write into the charts directory tree, which is already organized as charts/en and charts/zh with overview subfolders. For a first real use, skip rendering entirely. Open data/adopted.csv and filter to a single plan, then confirm that the rows for that plan are alternative models rather than additive allowances. That check takes a minute and prevents the most common misreading of the dataset.

Cache writes are the known hole

The README states that cache writes are not modeled separately, and that where a provider charges for them, converted token allowances may be overstated. This is not a footnote about rounding. Cache writes are a real line item for providers that bill them, and a workload that writes a large cache pays for it in a way the standard 97.5% cache-read assumption does not capture. If your workload is cache-write heavy, the project's $/MTok figure for that provider is optimistic, and the size of the optimism is not quantified anywhere in the README.

The second limitation is the convention itself. A 0.35% output share is a specific kind of workload. Code generation, long-form drafting and agentic loops are output-heavy, and under those workloads the ranking can invert, because output tokens are the expensive part of a three-part price. The project does not claim otherwise, but a reader who skips the conventions section will treat the leaderboard as universal. It is not. It is a ranking under one stated mix.

The third constraint is provenance. The README notes that for GLM Coding Plan there is still no fully specified independent V3 Pro/Max saturation test, and that the available Caijing saturation-cost test and community evidence are consistent in scale but not a substitute. That is an honest statement of a gap, and it means some points rest on weaker evidence than others.

How it differs from a metered-rate calculator

A conventional API pricing calculator takes your token counts and multiplies them by published input, output and cache rates. That works when you are buying metered API access, and it is the right tool if your usage is bursty or small, because you pay only for what you consume. Real API Pricing solves a different problem: it converts a flat subscription fee into an implied per-token rate so that a subscription can be placed on the same axis as a metered baseline. The 13 metered API baselines in the dataset exist for exactly that comparison.

The practical difference is what varies. In a metered calculator, your usage varies and the rates are fixed. Here, the fee is fixed, the allowance varies by plan and model, and the conversion factor is a project-wide constant. That makes the output stable and comparable across providers, and it also makes it insensitive to your actual traffic. If you want a number that reflects your own logs, this project will not produce it. If you want a number that lets you line up twenty plans without rebuilding a workload model for each one, this is the design that does it.

Licence, maintenance and the cost of staying current

The repository is MIT licensed, declared both in the LICENSE file and in the license field of package.json. MIT permits reuse and modification with attribution and without warranty. Nothing in the README suggests any additional restriction on the data files or charts, but the README does not state a separate licence for the data, so treat the repository-level MIT declaration as the only licence statement available. This is not legal advice; check the LICENSE file yourself if you plan to redistribute the charts.

The repository is not archived, and the last push was on 2026-09-12. The maintenance cost sits with the data, not the code. Vendors change plan definitions, the README records that Kimi's monthly pool is five times its weekly pool and that GLM is recomputed from Zhipu's official weekly credits, and the AA index version changed on 2026-09-07. Each of those is a reason to re-derive points rather than patch a number. The dated evidence lives under data/research/, with files such as the token-mix audit dated 2026-09-07, the quotas archive from 2026-09, and the GLM community review from 2026-09-07. Upgrading means re-running the pipeline against a new snapshot and accepting that scores from different index versions are not comparable.

Editorial conclusion

Adopt it if you are comparing subscription plans on a per-token basis and you accept a fixed 97.5% cache-read / 2.15% fresh input / 0.35% output workload, because that convention is what makes the numbers comparable at all. Do not adopt it if your traffic is output-heavy or cache-write-heavy: the README states cache writes are not modeled separately, so converted token allowances may be overstated where a provider charges for them. Before you quote any figure, open data/conventions.json and the token-mix audit file, and confirm which snapshot you are reading, because the adopted snapshot is 2026-09-09 and the AA Intelligence index changed to v4.3 on 2026-09-07.

Frequently asked questions

How much do APIs typically cost according to Real API Pricing?

The project does not publish a single typical figure. It converts each subscription plan into a $/MTok value by dividing the monthly fee by the monthly usable tokens, and it places all 200 subscription and API points on one comparable scale. The README notes that prices use a logarithmic axis, with cheaper points farther right.

How much do 1000 tokens cost in the Real API Pricing dataset?

The dataset reports unit prices per million tokens ($/MTok), not per thousand, so a 1000-token figure is a division of the published number rather than a value the project states. The unit price depends on the plan, the model and the standard workload the project applies.

How do you price an API with Real API Pricing?

The README gives the formula as monthly subscription fee divided by monthly usable tokens. Dollar and credit pools are converted using one project-wide workload of 97.5% cache reads, 2.15% fresh input and 0.35% output, which the README describes as a comparison convention rather than a claim about any provider's actual workload.

Can I get an API for free according to Real API Pricing?

The README does not document a free tier. It covers 202 adopted plan-model points, 188 of which carry a monthly allowance, plus 13 metered API baselines, and every one of those is tied to a plan or a metered rate rather than a no-cost option.

Official sources

  1. FeiZhuLulu/real-api-pricing on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/feizhululu-real-api-pricing.svg)](https://hysenlabs.com/projects/feizhululu-real-api-pricing)