Model or dataset
beam-cloud/beta9 avatar
beam-cloud/beta9

Beta9: The Open-Source Engine Behind Beam's Serverless GPU Platform

Ultrafast serverless GPU inference, sandboxes, and background jobs

1,799 stars169 forksGoAGPL-3.0

At a glance

What is it?
Beta9 is the Go-based open-source runtime that powers Beam.cloud, offering serverless GPU containers with sub-second cold starts, isolated sandboxes, and background task queues through a Python decorator SDK that works the same whether you self-host or use the managed service.
Who is it for?
Beta9 suits teams that need serverless GPU inference with a clean Python SDK and want the option to run the same engine on their own infrastructure. The AGPL-3.0 licence requires releasing modifications to anyone who receives the running service, which rules it out for teams that want to ship a modified version without open-sourcing their changes.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Beta9 Solves and How It Fits Into the Beam Ecosystem

Running GPU inference at scale involves two problems that are harder to solve together than separately: getting a container started quickly enough to avoid unacceptable latency, and not paying for idle capacity when demand drops. Beta9 addresses both with a custom container runtime and scheduler designed for cold starts under a second, and serverless-by-default execution that scales to zero when there are no pending tasks.

Beta9 is the open-source engine that powers Beam.cloud, a managed cloud platform described in the README. The relationship is explicit: users can self-host Beta9 for free or use Beam.cloud for managed hosting. This means the open-source codebase and the commercial product run the same engine. A team that starts with the managed service and later decides to self-host can do so without switching platforms.

Three capabilities are exposed through the Python SDK: serverless inference endpoints, isolated sandboxes for running LLM-generated or untrusted code, and background task queues. All three share the same container image system and the same scaling mechanism.

Architecture: Go Engine, Python SDK, and Kubernetes Runtime

The server-side engine is written in Go. The go.mod specifies Go 1.25.7 as the language version. The module is github.com/beam-cloud/beta9. The SDK used by application developers is a Python package named beam-client. The two communicate through a defined API; the repository structure separates them into a cmd/ directory for the Go service entrypoints and an sdk/ directory for the Python client.

The runtime layer uses a custom container runtime and a scheduler, with embedded caching for fast cold starts. The README states containers start in under a second. The repository includes manifests for Kubernetes deployment (manifests/ directory, with kustomize overlays). The Makefile shows the development setup uses k3d, a Kubernetes distribution that runs in Docker, with Helm for deploying the Beta9 chart.

Worker processes are released separately from the main engine, as shown in the release history: worker-0.1.771 on 2026-09-26, worker-0.1.770 on 2026-09-26, worker-0.1.769 on 2026-09-26. The fast release cadence on the worker component suggests that the execution layer is iterated independently from the core infrastructure code.

Installing the SDK and Deploying a First Endpoint

The Python SDK installs from PyPI:

shell
pip install beam-client

The README states that new users should create an account at beam.cloud and follow the Getting Started guide. The SDK uses Python decorators to define endpoints. A serverless inference endpoint is defined by applying the @endpoint decorator to a function and specifying the image, GPU type, CPU, and memory:

python
from beam import Image, endpoint
from beam import QueueDepthAutoscaler

@endpoint(
    image=Image(python_version="python3.11"),
    gpu="A10G",
    cpu=2,
    memory="16Gi",
    autoscaler=QueueDepthAutoscaler(max_containers=5, tasks_per_container=30)
)
def handler():
    return {"label": "cat", "confidence": 0.97}

The `QueueDepthAutoscaler` in this example caps the deployment at 5 containers, with each container handling up to 30 tasks concurrently. This is the autoscaling configuration for the endpoint. The GPU type (`gpu="A10G"`) is passed as a string. The README also lists 4090s and H100s as supported GPUs on the managed cloud; self-hosted deployments depend on whatever GPUs are attached to your nodes.

Deploying this function sends it to the configured backend (managed Beam.cloud or a self-hosted Beta9 cluster). The endpoint is then accessible as an HTTP endpoint.

Sandboxes for Isolated Code Execution

The sandbox feature creates isolated containers specifically for running code that the caller does not fully trust, such as LLM-generated code or user-submitted scripts. A sandbox is created with a single call:

python
from beam import Image, Sandbox

sandbox = Sandbox(image=Image()).create()
response = sandbox.process.run_code("print('I am running remotely')")

print(response.result)

The sandbox runs in an isolated container provisioned on demand. The README describes this as a use case for running LLM-generated code. Each call to `run_code` runs in the container's process space; the `response.result` field holds the output.

This capability sits alongside the inference endpoint and background task features in the same SDK. A team that uses Beta9 for inference could use sandboxes in the same deployment for safe code execution without running a separate orchestration layer. The trade-off is that the sandbox's isolation properties depend on the container runtime's security guarantees, which the README does not document in detail.

Background Task Queues and Resilience

The task queue feature wraps a function in a durable background worker with retry logic. The @task_queue decorator defines the resource requirements, an input schema, and a retry policy:

python
from beam import Image, TaskPolicy, schema, task_queue

class Input(schema.Schema):
    image_url = schema.String()

@task_queue(
    name="image-processor",
    image=Image(python_version="python3.11"),
    cpu=1,
    memory=1024,
    inputs=Input,
    task_policy=TaskPolicy(max_retries=3),
)
def my_background_task(input: Input, *, context):
    image_url = input.image_url
    print(f"Processing image: {image_url}")
    return {"image_url": image_url}

The `TaskPolicy(max_retries=3)` sets the retry ceiling for failed tasks. The README describes this as a replacement for a Celery queue. The function is invoked by calling `.put()` on the decorated function with the appropriate input, as shown in the README's example.

The README also mentions hot-reloading and webhook support alongside task queues. These features make it possible to update a running worker without redeploying the full container image, and to trigger external callbacks when a task completes.

Self-Hosting Beta9 and the AGPL-3.0 Licence Implications

Beta9 is licensed under AGPL-3.0. The AGPL adds a network use clause to the GPL: anyone who deploys a modified version of AGPL-licensed software as a service and allows users to interact with it over a network must release the source of those modifications. This is materially different from Apache-2.0 or MIT. A team that builds on Beta9 and offers it as a service while keeping internal changes private is not compliant with the AGPL.

Self-hosting requires a Kubernetes environment. The Makefile in the repository shows the development setup uses k3d, the kustomize build toolchain, and Helm:

makefile
setup:
	bash bin/setup.sh
	make k3d-up runner worker gateway

The deploy/ directory holds Helm charts and the manifests/ directory holds kustomize overlays for cluster and staging environments. The README does not provide a simplified single-server setup path outside of using Beam.cloud.

A well-known alternative in the serverless GPU space is Modal, which exposes a similar Python decorator API for GPU workloads but is a fully managed, closed-source service with no self-hosting option. Beta9's differentiator is the open-source engine and the self-hosting path under the AGPL.

Maintenance Status

The last push to the main branch was on 2026-09-27. The worker component received three releases on 2026-09-26: worker-0.1.770, worker-0.1.771, and worker-0.1.769. The release cadence on the worker builds is high. The go.mod file requires Go 1.25.7, indicating the project tracks a recent Go version.

The repository includes a CODE_OF_CONDUCT.md and a CONTRIBUTING.md, and the README has a contributions section inviting feature requests, bug reports, and pull requests. The SECURITY.md file is present. These are standard indicators of an actively maintained open-source project.

The managed Beam.cloud service adds a commercial layer on top of the open-source engine. The relationship between the open-source release cadence and the managed service's feature set is not described in the repository materials.

Editorial conclusion

Beta9 suits teams that need serverless GPU inference with a clean Python SDK and want the option to run the same engine on their own infrastructure. The AGPL-3.0 licence requires releasing modifications to anyone who receives the running service, which rules it out for teams that want to ship a modified version without open-sourcing their changes. Before self-hosting, verify that your environment can run a Kubernetes cluster with k3d, since the setup script requires it. The managed Beam.cloud service removes that requirement but introduces vendor dependency.

Frequently asked questions

Does Beta9 require Kubernetes to self-host?

The Makefile in the repository uses k3d (a Kubernetes distribution that runs in Docker) and kustomize for the local development setup. The deploy/ directory contains Helm charts for cluster deployment. The README does not describe a non-Kubernetes self-hosting path.

What does the AGPL-3.0 licence mean for Beta9 deployments?

AGPL-3.0 requires that teams who deploy a modified version of Beta9 as a network service must release the source of their modifications. This differs from permissive licences such as MIT or Apache-2.0, which do not impose this requirement.

What GPUs does Beta9 support?

The README lists 4090s and H100s as available on the managed Beam.cloud service, and states that users can bring their own GPUs. The SDK accepts the GPU type as a string in the @endpoint decorator, such as gpu="A10G".

Official sources

  1. beam-cloud/beta9 on GitHub
  2. License: AGPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/beam-cloud-beta9.svg)](https://hysenlabs.com/projects/beam-cloud-beta9)