Model or dataset
kossisoroyce/timber avatar
kossisoroyce/timber

Timber: a compiler for tree models, and a container that installs the wrong copy of it

Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.

689 stars23 forksPythonNOASSERTION

At a glance

What is it?
kossisoroyce/timber turns XGBoost, LightGBM, scikit-learn, CatBoost, ONNX and URDF files into self-contained C99 inference artifacts behind an Ollama-compatible HTTP API, while its Dockerfile pip-installs the published wheel instead of the source sitting next to it and its declared dependencies never mention the C compiler it needs.
Who is it for?
Timber is doing something specific and worth doing: taking a trained model out of Python entirely, and the artifact it produces is small, deterministic, allocation-free, and readable as C rather than as an opaque binary. The Ollama-shaped API makes it cheap to try, since a client that already talks to a model server will find the same endpoints on port 11434.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 172 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Dockerfile installs the published wheel and never copies the source

The image is short enough to see the problem in full. It starts from `python:3.11-slim`, installs `gcc`, and then runs `pip install --no-cache-dir timber-compiler`. There is no `COPY` of the repository anywhere in it. The package name on PyPI is `timber-compiler`, not `timber`, so that line resolves against the registry rather than against the working tree, and `docker compose build` in a clone of the project produces an image containing the last release rather than your edits. The compose file has the same character: it builds from `.`, publishes `11434:11434`, passes a `MODEL_URL` environment variable defaulting to the breast cancer example hosted on GitHub raw, and mounts a named volume at `/data/timber`, which matches the `TIMBER_HOME` the image sets. The default command serves that remote model with `--host 0.0.0.0`, which binds every interface.

The C compiler is a runtime requirement that no dependency provides

The pipeline has five stages, and the fourth is the one that reaches outside Python. Parse reads the native model format into a framework-agnostic Timber IR, Optimize applies dead-leaf elimination, threshold quantization, constant-feature folding, and branch sorting, Emit generates deterministic portable C99 with no dynamic allocation and no recursion, Compile has `gcc` or `clang` produce a shared library which is then loaded via `ctypes`, and Serve wraps the binary in an Ollama-compatible HTTP API. So the compiler runs when you serve a model, not when you install one. The declared Python dependencies are `click`, `numpy`, `rich`, `requests`, `cryptography`, and `tomli` for Python below 3.11, and none of them is a toolchain. The image installs `gcc` for exactly this reason, and it stays in the final layer.

Port 11434 and /api/predict make it a stand-in for Ollama

The API surface is deliberately narrow. The server reports its endpoint as `http://localhost:11434`, which is Ollama's default port, and answers on `/api/predict` for both `POST` and `GET`. The request body is a model name and a matrix of rows, so a thirty-feature sample posts as an array of thirty numbers. The response shape is not identical between the two worked examples in the documentation: one returns `{"model": ..., "outputs": [[0.9971]], "n_samples": 1}` and the other returns `{"model": "fraud-detector", "outputs": [[0.031]], "latency_us": 1.8}`. A client written against the first will not see the timing field. Models are addressed by name once loaded, through two commands:

console
$ timber load fraud_model.json --name fraud-detector
$ timber serve fraud-detector

`timber serve` also accepts a URL directly, and the headline claim about that is that no pre-download step is needed: point it at any URL and it downloads, compiles, and serves in one command. Which formats that works for is bounded by what the parser reads natively rather than by conversion, and the table is specific: XGBoost from `.json` for all objectives, LightGBM from `.txt`, `.model`, or `.lgb`, scikit-learn from `.pkl` or `.pickle`.

Five optimization passes, and the tool tells you how many applied

The demo output is the most revealing part of the documentation because it reports its own work rather than asserting a result. Format detection prints `xgboost`, parsing reports `50 trees`, `30 features`, and `binary:logistic`, optimization reports `3/5 passes applied`, and emission reports `169 lines` of C99 producing a `47.9 KB` compiled binary. So of five optimization passes, three changed anything for that particular model, which is a more honest signal than a throughput number because it tells you how much of the compiler's machinery was idle. The headline figures sit in the same block: roughly 2 microseconds for single-sample inference, roughly 336 times faster than Python inference, a roughly 48 KB artifact, and zero runtime dependencies. The 336 is a ratio against a named baseline, Python XGBoost, not a general claim about model serving. The estimator list behind the scikit-learn column is longer than tree ensembles: GradientBoosting, RandomForest, ExtraTrees, DecisionTree, linear and logistic models, SVMs including OneClassSVM, IsolationForest, GaussianNB, and k-neighbours classifiers and regressors, with Gaussian process regression named alongside them in the general description.

The pyproject hard-codes 0.6.0 and still requires setuptools-scm

The build requirements are `setuptools>=68.0` and `setuptools-scm>=8.0`, and the version is written literally as `0.6.0` in the `[project]` table. A project that declares setuptools-scm in its build requirements and then pins its own version string is doing one of two things: it wants SCM-derived versions and the literal field wins, or it has migrated away from SCM and left the dependency behind. Either way, a tagged release does not produce a new version automatically. The classifier list has its own oddity, since it includes `Programming Language :: Python :: C`, which is not a real classifier value; the correct one would be `Programming Language :: C`. The project otherwise declares itself as `Development Status :: 4 - Beta`, requires Python 3.10 or newer, and lists classifiers for 3.10 through 3.12.

Diagnose scripts ship beside four pre-trained model files

The `examples/` directory is more informative than a demo usually is. It holds four models to point at immediately: `breast_cancer_model.json`, `xgb_breast_cancer.json`, `lgb_breast_cancer.txt` for LightGBM, and `sklearn_pipeline.pkl`, plus `sample_model.json` and `test_samples.csv`. Alongside them sit two diagnosis scripts, `diagnose_accuracy.py` and `diagnose_base_score.py`, and the second of those pairs with a specific note in the format table, which flags XGBoost 3.1 and later per-class `base_score` handling as a supported case. There are four quickstart clients, one each for XGBoost, LightGBM, scikit-learn, and the ctypes wrapper, plus `bench_python.py` for the comparison the speed claims rest on, and `train_and_save.py` for producing your own. Two output directories, `compiled/` and `compiled_exact/`, suggest the difference between an optimized build and a faithful one is tracked rather than discarded.

Certification reports and signing come out of the same pip install

The `timber accel` backend is described as covering a very wide surface. On the hardware side: AVX2, AVX-512, NEON, SVE, and RVV SIMD, CUDA, Metal, and OpenCL for GPU, Xilinx and Intel FPGA through HLS, and Cortex-M, ESP32, and STM32 variants for embedded. On the assurance side: WCET analysis, DO-178C, ISO-26262, and IEC-62304 certification reports, Ed25519 artifact signing, AES-256-GCM encryption, air-gapped deployment bundles, and server generators for ROS 2, PX4, and gRPC. The README states that everything ships in one `pip install`, and the declared dependencies support part of that, since `cryptography` provides the signing and encryption primitives. The certification reports are generated artefacts rather than a certificate, and the FPGA and embedded targets need toolchains that no Python wheel contains.

Apache-2.0 in the metadata, no identifier on the repository

The licence is reported two different ways. `pyproject.toml` declares `license = {text = "Apache-2.0"}`, carries the `License :: OSI Approved :: Apache Software License` classifier, and the repository root contains a `LICENSE` file. The repository's own licence field reports no recognised identifier, which is what happens when a detector cannot map the file it finds. For most consumers the package metadata is what their tooling reads, so the declaration is usable, but the mismatch is worth resolving before an automated compliance scan flags it. The root also carries a `SECURITY.md`, a `CHANGELOG.md`, a `CODE_OF_CONDUCT.md`, and a `paper/` directory holding the technical paper, alongside `timber_technical_doc.md`, `llms.txt`, `skill.md`, `mkdocs.yml`, `docs/`, and a `website/` directory, which is a lot of documentation surface for a project at 0.6.0.

Editorial conclusion

Timber is doing something specific and worth doing: taking a trained model out of Python entirely, and the artifact it produces is small, deterministic, allocation-free, and readable as C rather than as an opaque binary. The Ollama-shaped API makes it cheap to try, since a client that already talks to a model server will find the same endpoints on port 11434. Before you depend on it, check three things. That a C compiler exists wherever you run it, because the pipeline shells out to `gcc` or `clang` and no Python dependency provides one. That your Dockerfile builds what you think, since the shipped one installs the published package and will quietly ignore your working tree. And that the licence is settled in writing for your organisation, since the package metadata says Apache-2.0 while the repository reports no recognised identifier. The performance figures are the author's own; measure them against your own baseline before they reach a design document.

Frequently asked questions

What does Timber compile and what does it produce?

It takes trained XGBoost, LightGBM, scikit-learn, CatBoost, and ONNX models, plus URDF robot descriptions for forward kinematics, and emits a self-contained C99 inference artifact with zero runtime dependencies, roughly 48 KB. The package also advertises LLVM IR and WASM targets.

How do I install Timber and serve a model?

Run `pip install timber-compiler`, then either `timber serve <url-or-path>` or `timber load fraud_model.json --name fraud-detector` followed by `timber serve fraud-detector`. The server listens on `http://localhost:11434` and answers on `/api/predict` with a model name and a matrix of rows.

Does Timber need a C compiler installed?

The compiled path does. The fourth pipeline stage runs `gcc` or `clang` to produce a shared library that is then loaded via `ctypes`, and the Dockerfile installs `gcc` for that reason. The declared Python dependencies are click, numpy, rich, requests, cryptography, and tomli, so the compiler is a system requirement rather than a package dependency.

What licence is Timber released under?

The package metadata declares `Apache-2.0` with a matching OSI classifier, and a `LICENSE` file is present in the repository root, while the repository's own licence field reports no recognised identifier. The package is published on PyPI as `timber-compiler` and is classified as `Development Status :: 4 - Beta`.

Official sources

  1. Issues
  2. kossisoroyce/timber on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kossisoroyce-timber.svg)](https://hysenlabs.com/projects/kossisoroyce-timber)