Model or dataset
scaleapi/llm-engine avatar
scaleapi/llm-engine

LLM Engine: an open source engine whose newest tag is two years older than its last commit

Scale LLM Engine public repository

839 stars83 forksPythonApache-2.0

At a glance

What is it?
Scale's LLM Engine packages model serving and fine-tuning as a Python library, a CLI and a set of Helm charts, usable against Scale's hosted infrastructure or your own Kubernetes cluster. Two things complicate that story: the release train stopped at a June 2024 pre-release while commits kept landing, and the documented quick start requires a Scale account before you can run a single line.
Who is it for?
LLM Engine is a reasonable choice if you want one client surface for both Scale's hosted inference and a cluster you operate yourself, and if you are willing to be the person who writes the Kubernetes documentation. Four things to verify first.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The last release is a beta from June 2024, the last commit is from this week

The release history has three entries visible, and all three are pre-releases of a zero major version: v0.0.0beta34 published on 2024-06-04, v0.0.0beta33 on 2024-05-20 and v0.0.0beta32 on 2024-05-07. The cadence in that window was roughly fortnightly. The repository, by contrast, was last pushed on 2026-10-02. So there are more than two years of commits after the final tag, and nothing in the repository declares them. A separate `VERSION` file sits at the root, which means the version is tracked in more than one place and the two can disagree. For anyone installing by tag this matters: `pip install scale-llm-engine` gives you the beta34 line, not the current tree, and there is no release cadence left running to notice that your install has fallen behind.

The quick start requires a Scale account before anything runs

The installation step is one command.

bash
pip install scale-llm-engine

The next step is not local. The quick start sends you to create an account with Scale's hosted platform, then to the settings page to grab an API key, then to export it.

bash
export SCALE_API_KEY="[Your API key]"

There is even a troubleshooting note attached to the most common failure: if you see an invalid API key error, you probably need to re-read your shell configuration. That sequence is the documented path into the project, and it means the first request you make goes to someone else's inference capacity. The self-hosted alternative exists and is described as first class, using the Helm charts in the repository on your own Kubernetes infrastructure. The gap is not in the capability but in which half got the documentation, which the project itself concedes further down.

The Helm charts are the half the project admits it has not documented

The project's own description of the problem is accurate. Deploying foundation models to the cloud and fine-tuning them are expensive operations needing infrastructure and machine-learning expertise, and they are hard to maintain as new models and new techniques appear. The answer is offered as three things: a Python library, a CLI and a Helm chart, covering both Scale's hosted infrastructure and your own Kubernetes cluster. Then the readme lists three capabilities under features coming soon, and the first of them is the documentation for exactly that second path. The text says the team is working on documenting installation and maintenance of inference and fine-tuning on your own infrastructure, and that for now the documentation covers using the client libraries against Scale's hosted service. The other two are cold-start behaviour, scaling to zero when idle and back up within seconds, and cost optimisation. All three are the self-hosting story.

Four names for one artefact, and a licence badge pointing at the wrong branch

Tracking the pieces takes a moment. The repository is llm-engine. The top-level source directory is `model-engine/`. The import in the example is `from llmengine import Completion`, with no hyphen. The distribution you install is `scale-llm-engine`. The documentation site is published under the scaleapi organisation. The Apache licence badge at the top of the readme points at a path on the `master` branch while the repository's default branch is `main`, which means the badge resolves to whatever that path holds on a branch that is no longer the default. None of this is dangerous, and all of it is friction for someone trying to answer the question of which artefact they are actually looking at.

Four formatting configs and three CI definitions in one root

The root directory carries `.black.toml`, `.isort.cfg`, `.ruff.toml` and `.pre-commit-config.yaml`. Ruff overlaps Black and isort by design, so the three formatters are not three opinions, they are two and a half, wired through a pre-commit hook. Alongside them sit three continuous-integration definitions: a CircleCI directory with a badge in the readme, a GitLab CI file, and a GitHub directory. There is also no packaging manifest at the root at all, no `setup.py` and no `pyproject.toml`, so the build configuration lives inside the source directory rather than where you would look first. Development and documentation dependencies are split into two separate requirement files, which is a reasonable choice for a project that builds a documentation site with its own toolchain and a poor one if you are trying to reproduce a working environment from the root.

The client API looks like OpenAI's without being OpenAI's

The starter example is short and is the shape most people will copy.

python
from llmengine import Completion

response = Completion.create(
    model="falcon-7b-instruct",
    prompt="I'm opening a pancake restaurant that specializes in unique pancake shapes, colors, and flavors. List 3 quirky names I could name my restaurant.",
    max_new_tokens=100,
    temperature=0.2,
)

print(response.output.text)

The request arguments are the ones you would expect, with `max_new_tokens` rather than a max-completion-tokens name. The response is not. There is no choices array and no message object; the text sits at `response.output.text`. A sibling API named `FineTune` is documented alongside it. So any code you already have that reads a chat completion will need an adapter, and the example's target model is a hosted Falcon instruct model, which means the first call you make from this snippet is a hosted call.

The readme names model families but never names an inference backend

The key features name LLaMA, MPT and Falcon as the foundation models served out of the box, and promise that any Hugging Face model can be deployed with a single command. Inference is described as optimised, with streaming responses and dynamic batching of inputs for higher throughput and lower latency. What is never stated is what actually runs the model. No inference server, no kernel library and no serving runtime appears anywhere in the readme. That is a meaningful omission for a project whose name is engine, because it means the value being sold is the control plane and the client surface rather than the numerical work, and it means the performance characteristics you would want to evaluate are decided by charts you have to read rather than by a sentence you can read. The same applies to fine-tuning: the readme says you can fine-tune on your own data and does not say with which trainer.

A worktrees directory is checked in beside the source

One entry in the root listing is out of place in a project of this kind, and it is a good indicator of how the work is actually done. Alongside `.github/`, `charts/`, `clients/`, `docs/`, `examples/`, `integration_tests/`, `scripts/` and the source directory there is a `.worktrees/` entry. That is the layout a local git worktree leaves behind, and it is present in the repository rather than ignored. The rest of the tree is conventional for a project of this size: charts for the Helm output, a separate clients directory, integration tests distinct from unit tests, an examples directory holding two notebooks, one on downloading a fine-tuned model and one on fine-tuning Llama 2 on a science question-answering dataset. The licensing is the Apache pair, a `LICENSE` file and a separate `NOTICE`, which is the correct arrangement for a redistributed Apache project.

Editorial conclusion

LLM Engine is a reasonable choice if you want one client surface for both Scale's hosted inference and a cluster you operate yourself, and if you are willing to be the person who writes the Kubernetes documentation. Four things to verify first. Treat the package as pre-release, because every published version is a beta of a zero major version and the last tag predates two years of commits. Read the hosted and self-hosted paths separately, since only the hosted one has working documentation. Decide whether the OpenAI-shaped request and response objects will satisfy your code or whether you will write an adapter, since the field names are their own rather than the familiar choices array. And check which inference backend the charts actually deploy, because the readme names model families but never names the engine underneath.

Frequently asked questions

How do I install Scale's LLM Engine?

One command from an index: `pip install scale-llm-engine`. Note that the distribution name differs from the repository name. The source directory in the repository is `model-engine/`, while the import is `from llmengine import Completion`.

What do I need before I can send an LLM Engine request?

An account with Scale's hosted platform and an API key taken from its settings page, exported as `SCALE_API_KEY`. If you see an invalid API key error, re-read your shell configuration so the new variable is picked up. The starter example targets the hosted model `falcon-7b-instruct`.

Can LLM Engine run on my own Kubernetes cluster?

Yes, through the Helm charts in the repository, using the Python library, CLI and chart on your own infrastructure. Be aware the readme lists Kubernetes installation and maintenance documentation under features coming soon and states that current documentation covers using the client libraries against Scale's hosted infrastructure.

What does the LLM Engine completion API look like?

Import `Completion` from `llmengine` and call `Completion.create` with `model`, `prompt`, `max_new_tokens` and `temperature`. The generated text is read from `response.output.text`, not from a choices array, so it differs from the OpenAI response shape. A `FineTune` API is documented alongside it.

What are the most recent LLM Engine releases?

v0.0.0beta34 on 2024-06-04, v0.0.0beta33 on 2024-05-20 and v0.0.0beta32 on 2024-05-07. Every published version is a pre-release of a zero major version, and the repository has continued to receive commits long after the last tag.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. scaleapi/llm-engine on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scaleapi-llm-engine.svg)](https://hysenlabs.com/projects/scaleapi-llm-engine)