Surogate: a C++/CUDA training and serving toolkit for NVIDIA GPUs
Train and serve LLMs at extreme speed and massive throughput.
At a glance
- What is it?
- Surogate combines a native training engine and an OpenAI-compatible serving engine in one Apache-2.0 repository. It is aimed at teams with NVIDIA hardware who want control over memory and precision, and the README's own numbers show where that control pays off and where it does not.
- Who is it for?
- Surogate fits teams with NVIDIA GPUs who want training and serving in one Apache-2.0 codebase and are willing to build C++/CUDA from source, especially single-card or small-card setups where the README's throughput gap over Unsloth is widest. Teams without CUDA hardware, or those who need a documented Windows or macOS path, should not start here.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Surogate solves, and who it is actually for
Most teams that fine-tune an LLM end up stitching two stacks together: a Python training library for the adapter work and a separate inference server for the deployed model. Surogate's pitch is that both halves live in one repository, with C++/CUDA engines underneath and Python configuration on top. The README describes it as putting "training and serving in one toolkit," and the repository layout backs that up: csrc/ holds the native engines, surogate/ holds the Python package, and the pyproject.toml declares a console script named surogate.
The audience is narrower than the tagline suggests. The classifiers in pyproject.toml list Linux only, with CUDA 12 and CUDA 13 environments, and the topic list is CUDA, NVIDIA GPU, fine-tuning, llama, qwen, SFT. There is no CPU path described and no Windows or macOS classifier. If you do not have an NVIDIA card, this project is not aimed at you. If you do, the interesting question is whether the native engine buys enough over the Python-first alternatives to justify a source build.
Two engines, one CLI: how the pieces connect
The architecture splits cleanly. The training engine handles pretraining, full fine-tuning, LoRA and QLoRA adapters, GRPO reinforcement learning, DPO preference training, and knowledge distillation. The serving engine is a native C++/CUDA HTTP server exposing OpenAI-compatible Chat Completions, Completions and Responses endpoints plus Anthropic-compatible Messages. Configuration for training is Python; execution is native.
A few mechanisms are worth calling out because they are not standard fare. For models larger than one card, the README says pipeline parallelism streams frozen weights for LoRA across PCIe GPUs, including systems without NVLink or GPU-to-GPU P2P. That is a deliberate trade: streaming weights costs bandwidth per step, but it lets a workstation with several consumer cards train adapters on a model that would not otherwise fit. Separately, the training engine supports expert parallelism for MoE models, with load balancing, routing metrics and expert imbalance detection. The extensibility story is a Python DSL with ahead-of-time automatic differentiation and native kernel dispatch, which is a real commitment to a custom graph format rather than a wrapper around an existing framework.
The Makefile reveals the build's real coupling. It explicitly links the extension against the virtual environment's pip-installed NCCL (2.29.x) rather than the system libnccl (2.28.x), because loading the system copy first breaks torch's libtorch_cuda with an undefined ncclCommResume symbol. That is a concrete integration constraint, and it means the build expects a .venv at a specific path with the nvidia/nccl package installed.
Installing Surogate and running a first fine-tune
The repository ships an install.sh at the top level and two Dockerfiles, Dockerfile.cu128 and Dockerfile.cu130, matching the CUDA 12 and CUDA 13 classifiers. A Python entry point also exists: pyproject.toml declares scripts = { "surogate" = "surogate.cli.main:cli_main" }, so once the package is installed you invoke the surogate command directly. The package requires Python 3.12 or newer.
For a source build, the Makefile wraps CMake. The default target is build, and CUDA_HOME is resolved with realpath so that a version-switching symlink like /usr/local/cuda does not leave a stale CMake cache pointing at a different toolkit. The Makefile also defines wheel, wheel-cu128 and wheel-cu130 targets.
make buildIf ccache is on the PATH, the Makefile wires it into the C, C++ and CUDA compiler launchers automatically, and exports CCACHE_CUDA_PATHS. You do not pass a flag for that; the presence of the binary is enough.
After the build, the CLI is the entry point. The repository keeps runnable examples under examples/, organized by task: examples/sft/, examples/dpo/, examples/grpo/, examples/pt/, examples/distillation/, examples/serve/ and examples/ruler/. Those directories are the place to copy a working configuration from, because the README points at them rather than printing a full config inline. The README does not reproduce the CLI's help output, so treat the examples directory as the authoritative starting point for a real run rather than guessing at flags.
Where the numbers are strong, and where they are not
The README's benchmark tables are unusually honest about the ceiling. Serving Qwen3.5-0.8B to a single user on one RTX 5090 reaches 802 tok/s against vLLM's 346, a 2.32x gain. At 8 users on the same card it is 2,765 tok/s against llama.cpp's 683, or 4.05x. But at 100 users on Qwen3.5-4B, the gap collapses to 5,345 tok/s against vLLM's 4,481, a 1.19x difference. The advantage is largest when concurrency is low or the baseline is llama.cpp, and it narrows sharply against vLLM under heavy load. If your workload is many concurrent users on a mid-size dense model, the headline multipliers do not describe your case.
Time to first token tells a similar story. The README reports 40 ms median for Qwen3.5-4B at 100 users versus 230 ms for vLLM, and 1.59 seconds for GLM-5.3-Flash at 16 users versus 55.13 seconds for llama.cpp. Those are latency claims from the repository's own docs, not independent measurements, and they should be reproduced on your hardware before they inform a deployment decision.
On training, the pattern is reversed: the smallest models show the biggest wins. Qwen3-0.6B on one H100 reaches 53,900 tok/s in BF16 against Unsloth's 21,300, a 2.53x gain. On a single RTX 5090 the same model gives 30,100 against 22,100, only 1.36x. Qwen3-8B on one RTX 5090 reaches 6,900 tok/s in FP4 against 3,500 in BF16, which is a precision comparison rather than a like-for-like one. The takeaway: the throughput advantage is real in the README's tables but is not uniform, and it depends on model size, precision and card count.
Limitations the README states plainly
Two constraints are documented rather than hidden. GRPO with shared-weight single-GPU BF16 LoRA is described as working across supported training families except Nemotron, so that specific configuration is out for one model family. And the README does not document rollback or downgrade steps between releases; the release list shows v1.4.6 through v1.4.8 landing within about three weeks of each other, which is a fast cadence to track if you pin a version without a rollback plan.
The build itself is the larger cost. This is C++/CUDA compiled through scikit-build-core and nanobind, and the Makefile carries a comment about toolkit ABI mismatches between cudaDeviceProp definitions when a symlinked CUDA_HOME changes underneath an existing CMake cache. That is a class of failure a pure-Python package never produces. The NCCL linking note adds another: the extension expects the pip NCCL 2.29.x that torch expects, not the system 2.28.x, and getting that wrong surfaces as an undefined ncclCommResume symbol at import time.
Finally, the README does not describe a Windows or macOS installation path, and pyproject.toml lists only POSIX Linux. Treat the toolkit as Linux and NVIDIA only.
How it differs from Unsloth and vLLM
The README benchmarks Surogate against Unsloth for training and against vLLM and llama.cpp for serving, so those are the honest comparisons. Unsloth is a Python library that patches and optimizes existing training code; Surogate replaces the execution layer with native C++/CUDA kernels and a compiled training graph. That difference is why the build is heavier and why the throughput gap is largest on the smallest models, where kernel overhead dominates. If you value a pip install and a notebook over compiled throughput, Unsloth is the lower-friction choice.
On serving, vLLM is the closer architectural peer, both being high-concurrency servers. Surogate's own table shows it ahead by 1.19x at 100 users on a 4B model, which is a modest margin for swapping out a mature server. The clearer distinction is native GGUF decoding: the README reports 802 tok/s for a single user on one RTX 5090, and llama.cpp is the baseline it beats there by 4.05x at 8 users. If your deployment is GGUF-based and lightly loaded, that is the case where Surogate's design shows up most. If you are already running vLLM at high concurrency, the README's numbers do not establish a reason to move.
Licence and the cost of keeping up
The project is Apache-2.0, and pyproject.toml sets license = "Apache-2.0" with license-files listing LICENSE alongside two bundled third-party licences: surogate/kernels/triton/fla_kda/LICENSE and csrc/src/serve/family/impl/sglang-LICENSE. The second one matters if you redistribute the serving engine, because a file named sglang-LICENSE indicates code derived from SGLang is present in the serving family implementations. Apache-2.0 is permissive, but the bundled notices are the ones to read before shipping a modified binary. This is a description of what the repository declares, not legal advice.
The upgrade cost is the build. Each release is a compiled artifact, so moving from v1.4.6 to v1.4.8 means recompiling the C++/CUDA extension against your toolkit, not just bumping a Python dependency. The Makefile's ccache integration softens that for repeated builds, but the first build on a new CUDA version is a full compile. The release cadence, three versions in roughly three weeks, means a team that pins versions will spend real time on rebuilds if it wants to stay current. The repository does not document a supported-version policy or a rollback procedure.
Editorial conclusion
Surogate fits teams with NVIDIA GPUs who want training and serving in one Apache-2.0 codebase and are willing to build C++/CUDA from source, especially single-card or small-card setups where the README's throughput gap over Unsloth is widest. Teams without CUDA hardware, or those who need a documented Windows or macOS path, should not start here. Before committing, verify three things: that your CUDA toolkit and driver match Dockerfile.cu128 or Dockerfile.cu130, that your target model family appears in the supported-models list, and that the serving API fields you depend on are listed in docs/inference/api.md rather than assumed from OpenAI compatibility.
Frequently asked questions
What does Surogate do?
Surogate is a native C++/CUDA toolkit that puts LLM training and serving in one repository. The training engine covers pretraining, full fine-tuning, LoRA and QLoRA, GRPO, DPO and distillation, while the serving engine exposes OpenAI-compatible and Anthropic-compatible HTTP APIs.
Which operating systems and GPUs does Surogate support?
The pyproject.toml classifiers list POSIX Linux with NVIDIA CUDA 12 or CUDA 13, and the repository ships Dockerfile.cu128 and Dockerfile.cu130 for those two toolkit versions. No Windows or macOS path is described.
How do I install Surogate?
The repository provides install.sh plus two CUDA-specific Dockerfiles, and the Makefile wraps the CMake build with a default build target. Once installed, pyproject.toml declares a console script named surogate that maps to surogate.cli.main:cli_main.
What Python version does Surogate require?
The pyproject.toml sets requires-python to ">=3.12", and the Makefile's NCCL linking logic references a .venv/lib/python3.12 path, so Python 3.12 is the version the build is written against.
Does Surogate support GGUF models for serving?
Yes. The README lists GGUF alongside NVFP4, BF16 and FP8, and its serving benchmarks describe native GGUF decoding at 802 tokens per second for a single user on one RTX 5090.
Is Surogate free to use?
It is released under Apache-2.0, and pyproject.toml lists LICENSE together with two bundled third-party licence files, one under surogate/kernels/triton/fla_kda and one named sglang-LICENSE under csrc/src/serve/family/impl. Read those notices if you redistribute a modified build.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/invergent-ai-surogate)