# LUPINE: Attaching Remote GPUs to CPU-Only Machines Over IP

> LUPINE is an Apache-2.0 C++ bridge that exposes GPUs on remote machines to CPU-only hosts through a CUDA shim, so a Mac or a GPU-less container can run CUDA code against a server elsewhere. Here is how the mechanism works, where it breaks, and what to check before adopting it.

**lupinemachines/lupine** — LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

- Repository: https://github.com/lupinemachines/lupine
- Website: https://lupine.sh
- Stars: 2,435 · Forks: 136
- Language: C++
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/lupinemachines-lupine

## What LUPINE solves, and for whom

A CPU-only machine cannot run CUDA code, because CUDA needs a driver and a device. The usual workarounds are to move the workload to the GPU host, or to rewrite it against a remote-execution API. LUPINE takes a third route: it puts a shim library in front of the CUDA driver API on the client, so the program still links against libcuda.so.1 and still calls cuLaunchKernel, but those calls travel over the network to a server process sitting next to a real GPU. The README describes the project as "a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines."

The intended audience is narrow and fairly clear. The README's Mac demo spins up a container with a virtual GPU on an arm64 Darwin host and runs a PyTorch-style tensor example that reports `device: lupine:0` and `gpu: NVIDIA GeForce RTX 4090`. That is the core use case: a developer on hardware that will never take a discrete NVIDIA card, pointing at a Linux box that has one. The second case is a CI or build container that has no GPU passthrough but can reach a GPU host on the network. Both are about avoiding a rewrite, not about sharing a GPU between many tenants, and the documentation never claims the latter.

## The shim, the HTTP/2 stream and the LZ4 body encoding

Inside the client container, `LD_LIBRARY_PATH=/opt/lupine/lib` is already set. CUDA driver users pick up the LUPINE libcuda.so.1 shim, and NVML users such as nvidia-smi pick up the LUPINE libnvidia-ml.so.1 shim. That is the whole trick on the client side: no source changes, no different API, just a library that intercepts the driver entry points and forwards them.

The transport is one long-lived TCP stream per client/server connection, carrying HTTP/2. RPC request and response bodies require `content-encoding: lz4`. Compression is applied transparently as one LZ4 frame per HTTP/2 body, and the README is explicit that peers do not negotiate or fall back to another encoding. That is a design decision worth noting: it removes negotiation complexity, but it means a peer that cannot do LZ4 cannot talk to LUPINE at all.

The repository layout backs this up. There are separate translation units for the client and server halves of each API surface (`cuda_client.cpp`, `cuda_server.cpp`, `cudart_client.cpp`, `cudart_server.cpp`, `hip_client.cpp`, `hip_server.cpp`), a `h2.cpp` for the HTTP/2 layer, and `dispatch.cpp` for routing. The presence of both CUDA and HIP client/server pairs in the tree is the clearest signal of where the project is heading, though the README's documented examples are all CUDA.

## Installing it and running nvidia-smi against a remote GPU

The README's Quick Start uses the published GHCR images rather than a source build. The examples pin CUDA 13.3.1 on Ubuntu 24.04, and other published tags use the same `cuda-<cuda-version>-ubuntu<ubuntu-version>` format. On the machine that actually has the GPU, start the server:

```bash
docker run --rm --gpus all -p 14833:14833 \
  ghcr.io/lupinemachines/lupine-server:cuda-13.3.1-ubuntu24.04
```

The `--gpus all` flag hands the container the host's devices, and port 14833 is the default LUPINE port (also the value in the repository's `.env.example`, alongside `LUPINE_SERVER=0.0.0.0`). On the CPU-only machine, point a client at it:

```bash
docker run --rm -it \
  -e LUPINE_SERVER=<server>:14833 \
  ghcr.io/lupinemachines/lupine-client:cuda-13.3.1-ubuntu24.04 \
  nvidia-smi
```

Replace `<server>` with the address of the GPU host. What you should see is the familiar nvidia-smi table, but with the remote card's name and driver version. The README shows a real run against an RTX 4090 with Driver Version 590.48.01 and CUDA Version 13.1, which is a useful reminder that the client image tag and the server-side driver version are independent numbers; the tag describes the CUDA toolkit bundled in the image, not the driver on the GPU host.

If you do not have a GPU host yet, the README offers a hosted demo with a T4 attached, run as a client-only container:

```bash
docker run --rm \
  -e LUPINE_SERVER=demo.lupinemachines.com:14833 \
  ghcr.io/lupinemachines/lupine-client:cuda-13.3.1-ubuntu24.04 \
  nvidia-smi -L
```

The README warns that this can take a while if no GPU is currently provisioned. To see LUPINE from a Python program rather than a shell command, the README's Mac demo runs a bundled example directly from the raw GitHub URL:

```bash
uv run https://raw.githubusercontent.com/lupinemachines/lupine/main/python/examples/tensor.py
```

It prompts for the LUPINE server host and port, then prints `cuda available: True`, `device: lupine:0`, the GPU name, and a small tensor result. The same example is the cheapest way to check that your client can reach your server before you try a real workload.

## Multi-GPU across servers, and what the ordinal list hides

The client accepts a comma-separated `LUPINE_SERVER` list. Devices are exposed as one local ordinal list in server order: all GPUs from the first server, then all GPUs from the next, and so on. The README's example starts a server on `gpu-host-a` and another on `gpu-host-b`, and the client then sees a flat device list.

This is convenient and also the least transparent part of the design. Nothing in the README describes a way to query which physical host a given ordinal belongs to at runtime, so a program that picks `cuda:3` is picking it by position in a list that depends on the order you typed the servers. Add a server to the front of the list, or lose one to a restart, and ordinals shift under the application. For a fixed job that always uses the same hosts this is fine. For anything that discovers devices dynamically, it is a footgun worth planning around, and the README does not document a stable mapping between ordinal and host.

## Connection stability, and the retry policy that is deliberately absent

The README is unusually candid about a real failure mode. Each client/server connection is a single long-lived TCP stream. Long-running workloads sit idle between training steps, during host-side data loading, or inside long kernels, and stateful middleboxes such as cloud load balancers, NAT gateways, conntrack tables and firewalls reap idle flows far sooner than the kernel's default 2-hour keepalive. The next RPC then fails fatally.

LUPINE's answer is TCP keepalive on every connection, client and server, with a 60s idle interval, 15s between probes, and 3 unanswered probes before giving up. Probes are sent only while idle, so active transfers pay no latency cost, and a dead peer is detected in roughly 105s instead of hanging on the TCP retransmit timer. There is also a connect retry with exponential backoff, each attempt bounded by a deadline, so a packet-filtered port is detected quickly rather than blocking for the full SYN-retransmit window.

The important negative is stated plainly: LUPINE does not retry RPCs, because retrying would break CUDA semantics. That is the correct call for a driver shim, but it means a dropped connection mid-kernel is not something the bridge papers over. If your workload runs for hours across flaky links, the keepalive tuning reduces the chance of a silent reap; it does not make the connection resumable. Socket buffer sizes are left to the OS, which auto-tunes on modern kernels, so there is no LUPINE-level tuning knob there either.

## Checkpointing is Linux-only and provider-supplied

Graceful shutdown is the feature most likely to be misread. On Linux, SIGTERM stops the server from accepting connections, asks every connection child to finish its in-flight CUDA calls, and waits for those children to exit. The README states this drain happens in the open-source server with no extra runtime dependency. That part is genuinely self-contained.

Actual checkpointing is not. Each connection child looks for `liblupinecr.so.0`, then `liblupinecr.so`, and uses the versioned provider ABI in `checkpoint_provider.h`. A missing or incompatible provider is a no-op, and the server still drains and exits normally. The provider is loaded before the child's first CUDA call so it can observe RM/UVM activity needed to discover allocations. Set `LUPINE_SESSION` in the client to attach a stable connection identifier, which the provider receives to restore the connection before its first CUDA RPC and checkpoint it after shutdown has drained. For an unkeyed connection, restore is skipped and checkpoint receives a null identifier. `LUPINE_CHECKPOINT_LIBRARY` can override the provider library path for a private deployment.

The boundaries are worth stating without softening. Providers own storage configuration, file layout, and any fallback policy for unkeyed connections; Lupine does not select a checkpoint directory. The README does not document a rollback procedure, nor does it ship a provider. So the open-source server gives you a correct drain and a hook, and the durable part is entirely yours to write. Anyone reading "graceful server checkpoints" as a working save/restore feature will be disappointed.

## Tracing, device printf, and the cost of opaque fatbins

Set `LUPINE_TRACE` on the client, server, or both to enable trace logging. `LUPINE_TRACE=0` or an unset value disables tracing, `LUPINE_TRACE=1` writes to stdout, `LUPINE_TRACE=2` writes to stderr, and any other non-empty value is treated as a file path opened in append mode. One variable controls both sides; `LUPINE_SERVER_TRACE` is no longer used, which is a rename that will silently do nothing for anyone upgrading from an older setup.

Device printf forwarding is the more interesting mechanism. LUPINE inspects uploaded PTX and cubin symbol data for `vprintf`, the CUDA device printf implementation. Until an image that may use device stdout is loaded, synchronization avoids stdout redirection and its process-global lock, allowing independent RPC lanes to synchronize concurrently. After a device-output-capable image is loaded, context, stream and event synchronization captures server fd 1 and forwards the bounded CUDA printf buffer to the client's stdout, with capture staying process-global so output from concurrent synchronization lanes is not misattributed.

The trade-off is explicit in the README: fully opaque compressed fatbins are treated conservatively as potentially using device stdout. In other words, if LUPINE cannot read the symbols, it assumes the worst case and takes the slower synchronization path. That is the right default for correctness, but it means a stripped or encrypted kernel image can cost you concurrency you did not expect to lose.

## Where LUPINE is the wrong tool, and what to use instead

LUPINE is not a GPU-sharing or multi-tenancy layer. Nothing in the README describes per-client isolation, quotas, or scheduling between clients on one server. A different approach to the same underlying problem is rCUDA, which also presents remote GPUs to a local application through a driver-level interception layer, but which is built around a client/server GPU virtualization stack with its own resource management and deployment model rather than a single long-lived HTTP/2 stream per connection. The practical difference for an evaluator is that LUPINE's transport and failure semantics are documented in the README in unusual detail (keepalive timings, no RPC retry, LZ4-only bodies), while the multi-client story simply is not there. If you need many users sharing one GPU with fair scheduling, look at container-level time-slicing or MIG on the GPU host, or at a remote-execution framework where the job is submitted rather than the driver call intercepted.

A second wrong-tool case is a network you do not control. The connection is a single TCP stream on port 14833, and the README's own framing is that stateful middleboxes reap idle flows. If your path crosses a corporate firewall that resets long-lived streams, keepalive probes will not save you, and there is no documented fallback transport. A batch system that moves data to the GPU host and runs the job there has none of these properties, at the cost of moving the code and the data.

## Licence, maintenance and upgrade cost

LUPINE is Apache-2.0, and the repository carries a NOTICE file alongside the LICENSE, which is the standard Apache-2.0 arrangement: you can use it commercially, modify it, and redistribute it, provided you keep the licence and notice and state significant changes. Apache-2.0 also includes an express patent grant, which matters for a project that reimplements driver-facing interfaces. This is a description of the licence text, not legal advice; if you are shipping LUPINE inside a product, have counsel review the NOTICE and any bundled third-party components.

The maintenance signal is good but should be read precisely. The repository is not archived, and the last push was on 2026-08-11, which is the same timestamp as the v1.0.0 release. v0.5.0 landed on 2026-08-04 and v0.4.0 on 2026-07-31, so the project moved through three releases in under two weeks before the v1.0.0 tag. That is a burst of activity rather than a long, steady cadence, and a 1.0 tag on a project whose checkpoint provider ABI is versioned and whose README still says `LUPINE_SERVER_TRACE` is no longer used suggests interface churn is not finished.

The upgrade cost is dominated by the image tags, not the source. Client and server images are published as `cuda-<cuda-version>-ubuntu<ubuntu-version>`, and the README's own example pairs a CUDA 13.1 driver version with a 13.3.1 image tag, so the tag is a toolkit version rather than a guarantee about the host driver. Expect to pin tags explicitly and to re-test after any bump, particularly if you use the checkpoint provider, because the provider ABI in `checkpoint_provider.h` is loaded by exact soname (`liblupinecr.so.0` then `liblupinecr.so`) and an incompatible provider silently becomes a no-op rather than an error.

## Conclusion

Adopt LUPINE if you have a CPU-only client (a Mac, a laptop, a container without a GPU) and a reachable Linux GPU host, and you want CUDA and NVML calls to work without rewriting the program. Do not adopt it if you need multi-tenant isolation between clients, a documented security model, or checkpointing on a platform other than Linux. Before committing, verify three things: that the client image tag matches your CUDA version in the cuda-<cuda-version>-ubuntu<ubuntu-version> format, that your middleboxes tolerate a long-lived HTTP/2 stream on port 14833, and that you can supply your own liblupinecr.so.0 if you intend to use the checkpoint provider, because Lupine does not select a checkpoint directory for you.

## FAQ

### What is LUPINE used for?

It attaches GPUs on remote machines to CPU-only machines, so a program can make CUDA and NVML calls locally while the work runs on a GPU host elsewhere. The README's examples show a Mac running a tensor example against a remote RTX 4090, and a client container running nvidia-smi against a server.

### How do I install LUPINE?

The README's Quick Start uses the published GHCR images rather than a source build. You run the server on the GPU machine with `--gpus all -p 14833:14833`, then run the client with `LUPINE_SERVER` set to that host and port. Image tags follow the `cuda-<cuda-version>-ubuntu<ubuntu-version>` format.

### Does LUPINE work on a Mac?

The README includes a Mac demo on Darwin arm64 that runs a Python tensor example and reports `cuda available: True` with `device: lupine:0` pointing at a remote RTX 4090. The Mac acts as the client; the GPU still has to live on a Linux server.

### Does LUPINE retry a failed RPC?

No. The README states that LUPINE keeps connections alive without retrying RPCs, because retrying would break CUDA semantics. It relies on TCP keepalive with a 60s idle interval, 15s between probes and 3 unanswered probes, so a dead peer is detected in roughly 105s.

### Can I use LUPINE with more than one GPU server?

Yes. The client accepts a comma-separated `LUPINE_SERVER` list, and devices are exposed as one local ordinal list in server order: all GPUs from the first server, then all GPUs from the next. The README does not document a way to map an ordinal back to a specific host at runtime.

## Sources

- [Official documentation](https://lupine.sh)
- [Official README](https://github.com/lupinemachines/lupine#readme)
- [Project repository](https://github.com/lupinemachines/lupine)
- [Release notes](https://github.com/lupinemachines/lupine/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lupinemachines-lupine
