CLI tool
leptonai/leptonai avatar
leptonai/leptonai

leptonai/leptonai: the lep CLI and Python Client for NVIDIA DGX Cloud Lepton

A Pythonic framework to simplify AI service building

2,825 stars198 forksPythonApache-2.0

At a glance

What is it?
The leptonai Python package ships the lep command-line tool and a Client that turns deployed endpoints into callable Python functions. It is a control-plane client for NVIDIA DGX Cloud Lepton, not a model server, and the README's own commands show what it can and cannot do.
Who is it for?
Adopt leptonai if you already run workloads on NVIDIA DGX Cloud Lepton and want the lep CLI or the Python Client in front of it; the package is a control-plane client, so it is the wrong tool if you have no workspace to point it at.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What leptonai actually is, and who the package is for

The README opens with a sentence that settles most of the ambiguity: this is "the Python library and `lep` CLI for NVIDIA DGX Cloud Lepton." That phrasing matters because the repository name and the topics list (artificial-intelligence, cloud, deep-learning, gpu) suggest a model-serving framework, while the package itself is a client for a hosted platform. The library does not train models or serve inference on your machine. It talks to a Lepton workspace.

The intended user is someone who already has that workspace and wants to drive it from a terminal or from Python. The feature list is a control plane inventory: endpoints, batch jobs, dev pods, Ray and Slurm clusters, fine-tuning jobs, storage, secrets. If your day involves creating a training job on a GPU cluster and then checking whether it is still running, this package covers that loop. If your day involves writing a FastAPI handler for a model you host yourself, it does not.

There is a second audience the README addresses explicitly: agent users. The repository ships an agent skill under plugins/lepton-cli/skills/lepton-cli/SKILL.md that lets Claude Code or Codex drive the lep CLI in natural language, with per-agent manifests for Claude Code, Codex and Cursor. That is an unusual thing to find in a cloud SDK, and it tells you the maintainers expect the CLI, not the Python API, to be the common entry point.

How the lep CLI and Client are wired together

The architecture visible in the repository is thin by design. The pyproject.toml entry point maps the `lep` command to `leptonai.cli:lep`, so the CLI is the same package you import. Dependencies include click for the command surface, httpx with an exact pin of 0.27.2, fastapi, pydantic with an explicit exclusion of 2.1.0 because of a bug that breaks fastapi, and ray[default]. The ray dependency is worth noticing: it is not optional, so installing the library pulls Ray into your environment even if you only want the HTTP client.

The Client half works differently from a hand-written SDK. According to the README, `Client` reads the endpoint's OpenAPI schema and exposes each path as a method. That means method names are not fixed in the library. They come from whatever the deployed endpoint declares. The README's example calls `c.paths()` to discover available paths and `c.run.__doc__` to read the docstring for one of them, which is the documented way to find out what you can call before you call it.

Authentication sits across both halves. Endpoint creation has three modes: no `--tokens` under secure defaults, which protects the endpoint automatically and prints a generated API token; one or more repeatable `--tokens` values; or `--allow-unauthenticated-access`, which the CLI accepts with a warning. The README is explicit that `--public` only controls IP reachability and does not disable API-token authentication, and that `lep endpoint status` reports the two dimensions separately as `IP Access` and `API Token Authentication`. That separation is a good design choice, and the documentation states it clearly rather than leaving it to be discovered.

Installing leptonai and creating your first endpoint

Installation is a single pip command, and the README notes it also installs the `lep` command-line tool. The declared Python range in pyproject.toml is >=3.9,<3.14, so a 3.14 interpreter will not satisfy the requirement.

bash
pip install -U leptonai

After install, authenticate against your workspace. The README says this opens a browser to fetch credentials when you do not pass them in.

bash
lep login

Now deploy a container image as an endpoint and inspect it. The image reference is your own; the README uses a placeholder registry path.

bash
lep endpoint create -n my-endpoint --container-image my-registry/my-app:latest
lep endpoint list
lep endpoint status -n my-endpoint

What you should see: the create command prints a generated API token when secure endpoint defaults are enabled and you did not pass `--tokens`. Save that value, because clients need it to authenticate. The status command then reports `IP Access` and `API Token Authentication` as separate lines. If you would rather supply the credential yourself, `lep endpoint update -n my-endpoint --tokens MY_TOKEN` replaces the token list, and `lep endpoint update -n my-endpoint --allow-unauthenticated-access` clears it. The README states that the former hidden `update --remove-tokens` option is rejected.

For a first Python call, the README's pattern reads the endpoint schema and then invokes a path as a method. Note the `local(port=8080)` helper for something running on your own machine.

python
from leptonai.client import Client, local

c = Client("my-workspace", "my-endpoint", token="MY_TOKEN")
# ...or to something running locally:
c = Client(local(port=8080))

print(c.paths())
print(c.run.__doc__)
print(c.run(inputs="hello world"))

The two print statements are the useful part. Run them before the call and you will see which paths exist and what arguments the endpoint expects.

Token handling is the sharp edge

The most detailed passage in the README is about credentials, and that is a fair signal of where users get hurt. The central problem: a server-generated token can only be returned once, at creation. The README states that SDK callers who leave endpoint authentication unspecified should use `client.deployment.create_with_response(...)` and save the token from the returned resource, because the older `create(...)` method keeps its boolean return contract and cannot return a server-generated credential. It emits a `RuntimeWarning` for requests that may ask the server to generate one. If your integration code ignores warnings, you can create an endpoint whose token you never captured.

Updates have a matching trap. The README says SDK updates may not clear `api_tokens` by sending an empty list alone; you must set `allow_unauthenticated_access=true` in the same update, or replace the list with at least one token. So the intuitive operation, sending an empty list to mean "no tokens," does not do what it looks like.

On the read path, `lep endpoint get` redacts literal tokens by default. `--show-tokens` returns a credential-bearing response, and the README says to handle that output as a secret. This is the right default, but it means spec export workflows that need the literal token have to opt in explicitly. Treat any saved output from that flag the way you would treat a private key.

Batch jobs, dev pods and the Ray dependency

Beyond endpoints, the CLI covers two workload shapes the README demonstrates directly. A batch job takes a container image and a command, and a dev pod takes a resource shape.

bash
lep job create -n my-job --container-image my-registry/my-trainer:latest --command "python train.py"
lep pod create -n my-pod --resource-shape gpu.a10

The README names `gpu.a10` as a resource shape, and the feature list adds Ray and Slurm clusters, fine-tuning jobs, storage and secrets to the same command surface. For discovery, `lep --help` and `lep <command> --help` are the documented entry points, with full CLI references hosted at docs.nvidia.com.

The cost of this breadth is the dependency set. Ray is a required dependency, not an extra, and httpx is pinned to an exact version. If you are adding leptonai to an existing service that already pins httpx or Ray differently, expect to resolve that conflict before anything runs. The package also targets Python below 3.14, which excludes the newest interpreter at the time of writing.

Where leptonai is the wrong tool

The clearest limitation is structural: there is no offline mode. Every meaningful command targets a Lepton workspace, and `lep login` fetches credentials from the platform. If you want to run a model on your own laptop with no cloud account, this package gives you nothing. It is not a replacement for a local inference server, and nothing in the README suggests it is meant to be.

The second limitation is that the Python Client is only as good as the endpoint's OpenAPI schema. Because methods are generated from declared paths, an endpoint with a thin or inaccurate schema produces a thin or inaccurate client. The README's `c.paths()` and `c.run.__doc__` calls exist precisely because you cannot know the method surface ahead of time. Code that hardcodes method names without checking will break when the endpoint changes.

The third is documentation coverage. The README documents token redaction, token replacement and the unauthenticated opt-out, but it does not document rollback of an endpoint update, nor does it describe what happens to in-flight requests when a token list is replaced. Those gaps are in the README, not in the platform, but a reader deciding whether to automate endpoint updates should know the README is silent there. The repository also carries a developer note that early development happened in a separate mono-repo, which explains why commits from `leptonai/lepton` appear in the history.

Alternatives and how the approach differs

The honest comparison is not another Python library. It is the platform's own web console and the underlying cloud provider's tooling. A web console lets you click through endpoint creation without writing code, and it will show you the generated token in the interface. The lep CLI trades that for repeatability: the same `lep endpoint create` line can go in a script or a CI job, and `lep endpoint status` gives you a parseable answer about both IP access and token authentication. If your workflow is one-off experimentation, the console is faster. If it is repeated provisioning, the CLI is the one that survives.

Within Python, the alternative is calling the platform's HTTP API directly with httpx or requests. That gives you full control over retries and error handling, and it avoids pulling Ray into your dependency tree. What you give up is the schema-driven Client, which discovers paths and docstrings for you, and the CLI's consistent flag surface for tokens. For a small integration against one stable endpoint, raw HTTP is defensible. For a team managing many endpoints, jobs and pods, reimplementing that surface is wasted effort.

A third option is to skip the control plane entirely and run your own scheduler, for example Ray directly on your own machines. That removes the platform dependency but also removes the managed workspace, the token model and the endpoint abstraction that this library exists to talk to. The choice is really about whether you want a managed control plane at all.

Editorial conclusion

Adopt leptonai if you already run workloads on NVIDIA DGX Cloud Lepton and want the lep CLI or the Python Client in front of it; the package is a control-plane client, so it is the wrong tool if you have no workspace to point it at. Verify two things before rolling it out: that your target Python is inside the >=3.9,<3.14 range declared in pyproject.toml, and that your endpoint creation path saves the API token, because an endpoint created without --tokens under secure defaults is protected automatically and the token is printed only at create time. If your SDK code still calls create(...) rather than create_with_response(...), the RuntimeWarning is the signal that a server-generated credential may be dropped.

Frequently asked questions

What is leptonai/leptonai?

It is the Python library and lep CLI for NVIDIA DGX Cloud Lepton, released under Apache-2.0. It lets you create and manage endpoints, batch jobs, dev pods, clusters, storage and secrets from the command line or from Python, and call deployed endpoints through a Client that reads the endpoint's OpenAPI schema.

How does the Python Client turn an endpoint into callable functions?

The Client reads the endpoint's OpenAPI schema and exposes each path as a method, so method names come from the deployed endpoint rather than the library. The README shows calling c.paths() to list available paths and c.run.__doc__ to read a method's documentation before invoking it.

What happens if I create a leptonai endpoint without passing tokens?

In workspaces with secure endpoint defaults enabled, the endpoint is protected automatically and the create command prints a generated API token that you need to save. The README notes that the --public option only controls IP reachability and does not disable API-token authentication.

Why does leptonai's create(...) method emit a RuntimeWarning?

The README states that create(...) keeps its boolean return contract and cannot return a server-generated credential, so it warns for requests that may ask the server to generate one. SDK callers who leave authentication unspecified should use create_with_response(...) and save the token from the returned resource instead.

Official sources

  1. leptonai/leptonai on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/leptonai-leptonai.svg)](https://hysenlabs.com/projects/leptonai-leptonai)