Library / SDK
elliotgao2/toapi avatar
elliotgao2/toapi

toapi keys its JSON response by your class name, and its cache has no documented lifetime

Every web site provides APIs.

3,555 stars241 forksPythonMIT

At a glance

What is it?
toapi turns web pages into JSON APIs declaratively, with CSS selectors, on-demand fetching and an in-memory cache. The response is wrapped in a key named after your Item class, the headless browser option expects a Firefox driver path, the package ships a console script the README never mentions, and the manifest's URLs point at a different account name than the repository.
Who is it for?
toapi fits the case where you want a small read-only endpoint in front of a page that has no API and you do not want to run a crawler. Its declaration style is the reason to try it: fields are CSS selectors, routes map your clean paths to the messy source URLs, and cleaning hooks let you post-process a field without touching the fetch.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 39 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The response envelope is your class name

Look at the JSON in the quickstart and the first thing to notice is the wrapper. The class in the example is called Post, and the response is an object with a single Post key holding a list of entries, each carrying the two fields that were declared, a title and a url. So the shape of your API is derived from a Python identifier, and renaming the class changes every response you serve. The example class also carries four decorators rather than one: the site it belongs to, the selector that marks a list of items, and two routes, one static and one with a page placeholder that maps your parameterised path onto the source site's own query syntax. That is the mechanism for pagination, and it is expressed as two paths rather than as one rule.

The cache is in memory with no documented rules

The how it works section is a diagram and four numbered steps: a route maps your path to a source URL, a fetch step retrieves the page, a parse step runs your selectors, and a serve step returns JSON. The caching sits inside step two, and step four says subsequent calls hit the cache. What the documentation does not contain is anything about how long an entry lives. There is no time to live, no maximum size, no eviction policy and no way to invalidate from the README, and the cache is described as being in memory, which means it goes when the process does and is not shared between workers. For a page that changes slowly this is a feature. For one that changes hourly, or for a long-running process, the behaviour is undocumented rather than configurable, as far as this repository is concerned.

The browser option asks for a Firefox driver

The normal fetch path is a plain HTTP library, and the documentation names it: pages are fetched with the requests library. Then there is the escape hatch for pages that need JavaScript, which is described as passing a browser argument to the Api constructor, with the example value being a path to geckodriver. That is Firefox's driver, so the browser mode is a headless Firefox rather than a headless Chromium, and the path is a driver executable rather than a browser name. Meanwhile the dependency list contains two fetching libraries rather than one, the requests library named in the text and a second one whose name describes fetching HTML. Neither the browser option nor the second fetching library is discussed in the how it works section, which covers only the requests path.

A packaged command line tool nobody documents

The project manifest declares a console script. The name is the package name and it points at a module in the cli subpackage, so installing the package puts a command on your path. The repository also carries an examples directory, and one of its two entries is a directory named for the command line framework the author uses for that interface. So the tool exists, it is packaged, and it has an example, and none of that appears in the readme. The installation section covers one thing, the pip command, and the usage section covers one thing, the Flask application. If you were evaluating this for a scripting use rather than a service use, the readme gives you nothing to go on and the examples directory is where the answer is.

The manifest points its links at another account name

Three of the four top level entries in the project metadata do not use the name the repository is published under. The homepage, the repository link and the documentation link are all written against a different handle, while the author field carries the same handle with an email address. The readme's own links and the license line use the repository's handle. Since both handles appear to belong to the same person, nothing is broken, but it is the kind of thing that breaks a packaging script, a badge generator or a citation tool that derives links from metadata. Two more fields are worth noting in the same block: the version is 2.2.4 and the classifier declares the project at production and stable status, while the repository publishes no releases at all.

Eight runtime dependencies, and a line length that is not enforced

The runtime list is eight packages. Two of them fetch, one parses HTML with CSS selector support, two are the web framework and the command line framework, one detects character encodings, and one is a terminal colour library that a server-side request handler has little use for. Every one of them is a lower bound rather than a pinned version, which is the opposite of what the companion lock file in the repository root implies. The lint configuration is also worth reading closely: the line length is set to a hundred characters, and in the same file the rule that enforces line length is switched off. Tests are additionally exempted from one specific check, the one about assertions without a message.

The quickstart names no file and then runs one

The quickstart shows a fragment with no filename. It creates an Api, decorates an Item class with four decorators, declares two fields, and calls run with a host and a port. The next heading is Run it, and the command under it is python app.py. So the document assumes a filename it never told you to use. The developer workflow, by contrast, is exact, because it is four commands with comments:

bash
git clone https://github.com/elliotgao2/toapi.git
cd toapi
uv sync          # install deps into .venv
uv run pytest    # run tests
uv run ruff check .

The contributing section then repeats the two checks that have to pass, and the header itself carries four identical badges pointing at the same package page on PyPI, which is a small sign of the same copy and paste habit.

Editorial conclusion

toapi fits the case where you want a small read-only endpoint in front of a page that has no API and you do not want to run a crawler. Its declaration style is the reason to try it: fields are CSS selectors, routes map your clean paths to the messy source URLs, and cleaning hooks let you post-process a field without touching the fetch. Two things to know before you build on it. The cache is in memory with no lifetime, size limit or eviction rule written down anywhere, so it is rebuilt on every process start and grows for as long as the process lives. And the response envelope is keyed by your class name, so renaming the Item class is a change to your API contract, not a refactor. Read the examples directory before writing the first route; there is a full page example there that the README does not mention.

Frequently asked questions

What does elliotgao2/toapi do?

It turns web pages into JSON APIs declaratively. You point it at a site, declare the fields you want with CSS selectors, and it fetches and parses pages on demand with caching, so there is no crawler to babysit and no database to maintain. It is MIT licensed and requires Python 3.10 or newer.

How do I install and run toapi?

Install it with pip. Then create an Api object, decorate an Item class with the site, list selector and route decorators, declare your fields, and call run with a host and a port. The quickstart does not say what file to put that code in, yet the step after it tells you to run python app.py.

Does toapi cache anything, and for how long?

It caches pages and parsed results, and the documentation says subsequent calls hit the cache while the cache itself is in memory. Nothing in the readme states a lifetime, a size limit or an eviction rule, and nothing describes a way to invalidate an entry, so the cache is rebuilt on every process start.

Does toapi ship a command line tool?

The manifest declares one: a console script named toapi mapped to the module toapi.cli, and the repository has an examples directory for it. The readme's installation and usage sections cover only the pip install and the Flask application, so the command is undocumented in the readme.

What is the shape of a toapi JSON response?

It is keyed by the name of your Item class. In the quickstart the class is called Post and the response is an object with a Post key holding a list of entries with the fields you declared. Renaming the class therefore changes the API contract rather than being a refactor.

Official sources

  1. elliotgao2/toapi on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/elliotgao2-toapi.svg)](https://hysenlabs.com/projects/elliotgao2-toapi)