NVIDIA NVCF: self-managed GPU inference and batch workloads on Kubernetes
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
At a glance
- What is it?
- NVIDIA Cloud Functions is a Go monorepo for routing inference, streaming and run-to-completion GPU work to worker clusters. It is aimed at platform teams that already run Kubernetes and want to operate the control plane themselves.
- Who is it for?
- Adopt NVCF if you already operate Kubernetes and GPU nodes and need one control plane to route inference, streaming and run-to-completion tasks across more than one cluster. Do not adopt it if you want a hosted endpoint with no cluster of your own, or if a single Triton or vLLM deployment behind an ingress already covers your traffic.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NVCF is for, and who ends up operating it
NVCF is a platform for deploying, managing and running GPU-accelerated workloads at scale. The README frames the problem in one line: it routes inference, streaming and other GPU work to worker clusters, so a team scales demanding workloads with less infrastructure to run itself. That phrasing is honest about the trade. You still run infrastructure. What you stop writing is the routing, lifecycle and multi-cluster glue around it.
The intended user is a platform or infrastructure team that already has Kubernetes and GPU nodes, possibly in more than one region, and wants one control plane in front of them. The repository is a monorepo holding service code, deployment assets, documentation, examples, CLI code, agent skills and validation tooling, which tells you the project expects to be installed and operated rather than consumed as a library. If your mental model of an inference endpoint is a single container behind a load balancer, NVCF is a layer above that, and the operational surface is correspondingly larger.
Control plane, invocation plane, compute plane: how the pieces connect
The README splits NVCF into services rather than a single binary. The control plane exposes the NVCF API, manages function and deployment state, handles secret management and coordinates platform operations. The invocation plane receives HTTP, streaming and gRPC requests, applies routing and rate limiting, and sends work to running function workloads. GPU clusters connect through the NVIDIA Cluster Agent, called NVCA, which registers GPU resources and manages workload execution on GPU nodes. Function artifacts live in registries the deployment can access, and observability, dashboards and runbooks sit alongside for operators.
That separation is the interesting design decision. Routing policy lives in the invocation plane, so a slow or failing worker can be taken out of rotation without touching the workload itself. Cluster registration is a separate concern handled by NVCA, which is what allows the deployment to span regions and GPU clusters, as the architecture diagram in the README shows. The repository map mirrors this: src/control-plane-services, src/invocation-plane-services and src/compute-plane-services are distinct trees, and deploy/ plus migrations/ hold the Helm charts, stack installation and datastore migrations. The cost of the split is that a working deployment means running several services, not one.
Functions versus tasks: picking the right workload type
NVCF draws a line between two workload shapes, and getting it wrong is the most common way to make the platform feel heavier than it is. Functions are long-running, invokable workloads. Use one when a client needs an endpoint for inference, streaming or another service-style GPU workflow. Tasks are asynchronous and run to completion. Use a task for batch inference, evaluation, fine-tuning or data preparation, where the job should finish and report status instead of staying online behind an invocation endpoint.
Both shapes can be packaged the same two ways. A container works when the workload is a single service with health and inference endpoints. A Helm chart works when the workload needs multiple coordinated containers, services, sidecars or other Kubernetes resources. The packaging choice is independent of the function-versus-task choice, which is a sensible split: a batch job that needs a sidecar is still a batch job.
The capability list is where the platform earns its place. Load-balanced routing balances workloads based on worker availability. Multi-cluster autoscaling scales workloads from zero to max across clusters. Mixed GPU support covers clusters with different GPU types, which matters when one workload needs a specific card and another does not. Health checks and telemetry track worker status and request latency.
Installing a self-managed deployment and running a first function
The README points installation at docs/user/installation.md and the fuller flow at docs/user/cli.md and docs/user/quickstart.md. There is no single install command in the README itself, and setup.sh at the repository root is not described there, so treat the installation document as the entry point rather than guessing.
Once a self-managed deployment is running and nvcf-cli is configured, the README gives this workflow. The init step sets up the CLI, and the api-key command generates a key for it:
nvcf-cli init
nvcf-cli api-key generateThe next steps create a function from a JSON file. The README says to update the example file with your function image before creating it, which is the step people skip:
nvcf-cli function create --input-file src/clis/nvcf-cli/examples/create-function.json
nvcf-cli function deploy create
nvcf-cli function invoke --request-body '{"message": "hello world"}'After invoke, you should get a response from the running workload. If the deployment is not healthy, the failure will surface earlier, at deploy create, because the platform waits on worker availability before routing traffic.
Building from source is a separate path. Bazel is the build, test and packaging tool across the monorepo, and the README gives a Linux quick start that fetches bazelisk into ~/.local/bin:
curl -fSL -o ~/.local/bin/bazel \
"https://github.com/bazelbuild/bazelisk/releases/download/v1.25.0/bazelisk-linux-$(dpkg --print-architecture)"
chmod +x ~/.local/bin/bazel
bazel build //src/clis/nvcf-cli:nvcf-cli
bazel test //src/clis/nvcf-cli/...On macOS the equivalent is brew install bazelisk followed by the same bazel targets. Builds read from a configured remote cache by default and do not upload local results, which is worth knowing before you spend time wondering why a build is slow or why it fails before local execution starts.
Where the build story is still uneven
The README is unusually direct about this. Native subtrees, meaning src/clis/nvcf-cli and src/libraries/go/lib, build fully under Bazel today. Phase B has landed Bazel scaffolds in upstream-owned service trees: nvcf-grpc-proxy, nvcf-ratelimiter, nvcf-nats-auth-callout-service, nvcf-cache/nvcf-unbound for dns-cache, nvcf-image-credential-helper and nvca. Their BUILD.bazel, MODULE.bazel and rules/oci files are picked up automatically when the subtrees are synced into the umbrella.
Read that carefully before planning a fork. The scaffolding means you can build, test and produce OCI images for those services from the umbrella without leaving the monorepo, but the wording distinguishes them from the native subtrees that build fully. A team that wants to modify the invocation plane and ship it has a different amount of work ahead of it than a team that only wants to build the CLI. The remote cache default adds a second wrinkle: if your network cannot reach the configured cache, the README's advice is to disable it for that build, which means your first build in a new environment may need a flag rather than working out of the box.
When NVCF is the wrong tool
The clearest case against it is scale of need. If you have one model, one cluster and steady traffic, NVCF asks you to run a control plane, an invocation plane, a compute-plane agent and the datastores behind them, then keep them upgraded. A single deployment behind an ingress does the same job with a fraction of the surface area. The platform's own capability list is the test: if you do not need multi-region routing, autoscaling from zero across clusters, or mixed GPU types, you are paying for features you will not use.
The second case is teams without Kubernetes operations experience. NVCA registers GPU resources and manages workload execution on GPU nodes, and the README's architecture section assumes familiarity with clusters, registries and Helm. Nothing in the README suggests a managed fallback for operators who would rather not own that.
The third is packaging fit. A workload that is a single long-running process with a health endpoint fits the container path cleanly. A workload that needs a bespoke scheduler, or that holds state across requests in a way that conflicts with routing across workers, will fight the model. The README does not document rollback behaviour for a function deployment, so if your release process depends on fast, documented rollback, that is a gap to resolve before adoption rather than after.
How NVCF compares with running Triton or vLLM directly
The natural alternative is to skip the platform and run an inference server such as NVIDIA Triton directly on Kubernetes, with your own autoscaling and your own routing. The difference is where the intelligence sits. Triton is a model server: it loads models and answers requests on one deployment. NVCF is the layer that decides which worker cluster answers, balances based on worker availability, applies rate limiting at the invocation plane, and scales workloads from zero to max across clusters. Running Triton directly means you write or assemble that layer yourself, typically with a service mesh, an HPA and some custom routing.
The second alternative is the hosted path. The README links build.nvidia.com as powered by NVCF, which is the same platform operated for you. Choosing that removes the install, the Helm charts and the cluster agent from your plate, at the cost of control over where workloads run and how the control plane is configured. The repository itself is the self-managed option, and the two are not interchangeable: one is a deployment you operate, the other is a service you call.
A third comparison is packaging. Because NVCF accepts a Helm chart as a workload, a team that has already packaged a multi-container inference stack does not need to rewrite it as a single container to move onto the platform. That is a lower migration cost than a system that only accepts a container image.
Licence, maintenance and the upgrade surface
The repository is licensed under Apache-2.0, and the tree carries a NOTICE file, an .allowed-licenses.txt and a license-compliance.md, which suggests dependency licence checking is part of the project's own tooling under tools/. Apache-2.0 permits commercial use and modification with the usual notice and patent terms; how that interacts with your own distribution obligations is a question for your legal team, not something the repository answers. Note that the licence covers the code in this monorepo, not the NVIDIA hardware or any hosted service you might pair it with.
On maintenance: the repository is not archived, and the last push was on 2026-08-28. Recent releases on that date include src/uis/nvcf-ui/v0.1.1 and src/libraries/rust/stargate/v0.14.2, with v0.14.1 earlier the same day. The presence of a Rust library alongside Go services is worth noting if you planned to assume a single-language codebase.
Upgrade cost is the part the README does not settle. The deployment tree holds Helm charts and migrations, and the README mentions datastore migrations under migrations/, so an upgrade is not just a new image tag: schema changes are part of the release. The README does not document a supported upgrade path or version compatibility matrix, and it does not document rollback. Verify both against docs/user/installation.md and the release notes before you put a production deployment on a version you cannot leave.
Editorial conclusion
Adopt NVCF if you already operate Kubernetes and GPU nodes and need one control plane to route inference, streaming and run-to-completion tasks across more than one cluster. Do not adopt it if you want a hosted endpoint with no cluster of your own, or if a single Triton or vLLM deployment behind an ingress already covers your traffic. Before committing, read docs/user/installation.md and docs/user/quickstart.md, confirm which of the services under src/control-plane-services, src/invocation-plane-services and src/compute-plane-services your team can operate, and check the Bazel scaffolds: only src/clis/nvcf-cli and src/libraries/go/lib are described as building fully under Bazel today, so plan for the rest of the tree to need work.
Frequently asked questions
What is NVIDIA NVCF?
NVIDIA Cloud Functions is a platform for deploying, managing and running GPU-accelerated workloads at scale, routing inference, streaming and other GPU work to worker clusters. It runs as Kubernetes services split into a control plane, an invocation plane and a compute plane, with GPU clusters connecting through the NVIDIA Cluster Agent (NVCA).
Is NVCF open source?
The repository is public and licensed under Apache-2.0, and it contains service code, deployment assets, documentation, examples, CLI code and validation tooling. The licence covers the monorepo code, not any NVIDIA hardware or hosted service.
How do I install NVCF?
The README points to docs/user/installation.md for installation and to docs/user/cli.md and docs/user/quickstart.md for the full setup, cleanup and configuration flow. There is no single install command documented in the README itself.
How do I invoke a function with nvcf-cli?
After initializing the CLI and generating an API key, you create a function from a JSON file, deploy it, then invoke it. The README's example runs nvcf-cli function invoke with a request body of {"message": "hello world"}.
What is the difference between an NVCF function and an NVCF task?
Functions are long-running, invokable workloads that give a client an endpoint for inference or streaming. Tasks are asynchronous, run-to-completion workloads suited to batch inference, evaluation, fine-tuning or data preparation, and they report status instead of staying online.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nvcf)