# FlashRT: Realtime AI Inference Engine for VLA Control and Small-Batch LLM Serving

> FlashRT is a C++ inference system designed for small-batch, latency-sensitive AI workloads: Vision-Language-Action (VLA) control on robotics hardware, LLM and VLM serving at single-stream or small concurrency, and video or audio generation paths. Unlike general-purpose serving frameworks, it builds model-specific dataflows with hand-written kernels and no ONNX export or per-driver compilation step, targeting the sub-50ms decision latency requirements of physical robot control.

**flashrt-project/FlashRT** — FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B

- Repository: https://github.com/flashrt-project/FlashRT
- Stars: 608 · Forks: 85
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/flashrt-project-flashrt

## The Small-Batch Realtime Cell That General-Purpose Serving Misses

AI inference tooling organizes around two common shapes. TensorRT optimizes for throughput on known workloads by searching for the fastest kernel tactics at compilation time and freezing them into an engine, which works well when the batch size and input shape are fixed. vLLM and SGLang maximize token throughput for LLM serving by batching requests from many concurrent users. Neither fits the small-batch realtime case: a robotics policy running on a Jetson edge device at 30 Hz with a batch size of one, or an LLM serving a single interactive stream where time-to-first-token matters more than aggregate throughput.

FlashRT targets that gap. Its README describes the design target as "the small-batch realtime cell with hand-tuned kernels and no compile step." The tool does not invoke TensorRT's tactic search or ONNX export. Instead, it assembles model-specific dataflows from captured graphs, system-level quantization plans, cross-layer fusion schemes, and hand-written or adapted kernels, and those dataflows run directly from the Python package after a CMake build of the CUDA extensions.

## Model-Specific Dataflows and How They Reduce Latency

FlashRT's architecture is described in the README as combining "a latency-first static execution pipeline, system-level quantization and calibration, cross-layer fusion plans, captured graphs over stable buffers, and hand-written or explicitly adapted kernels." The composition pattern addresses different latency contributors at different layers of the stack.

Captured graphs over stable buffers eliminate per-call overhead from Python dispatch and CUDA stream management. Cross-layer fusion plans merge operations that would otherwise require separate kernel launches and intermediate memory writes. System-level quantization reduces memory bandwidth, which is the bottleneck for small-batch inference on NVIDIA hardware. Hand-written kernels handle the workload-specific operations (flow-matching denoisers, attention with precise cache management, expert routing in mixture-of-experts models) where general-purpose libraries make sub-optimal trade-offs for this batch size regime.

The README notes that the composition pattern is hardware-agnostic in concept, with the current NVIDIA implementations spanning Jetson AGX Thor and Jetson Orin on the edge side through A100 and RTX 4090 and 5090 on the workstation and server side. AMD GPU, NPU, and other hardware support are described as actively expanding.

## Installing FlashRT and Building CUDA Kernels

FlashRT installs as a Python package with C++ extensions that must be built separately. The setup.py documents the workflow:

```bash
pip install -e ".[torch]"
```

Then build the CUDA kernel extensions with CMake:

```bash
cmake -B build -S .
cmake --build build -j
```

The CMake build drops the compiled .so files directly into flash_rt/ without a separate install step, so the editable pip install immediately sees the new kernels. The setup.py lists optional dependency groups: [torch] for the PyTorch frontend, [jax] for the JAX frontend, [server] for FastAPI serving, and [all] for everything. The legacy upstream attention path and the GROOT backend additionally require the flash-attn wheel; the default RTX Pi0 and Pi0.5 path uses a vendored flash_rt_fa2 extension and does not require flash-attn, allowing installation in environments where a prebuilt flash-attn wheel is not available.

The examples/ directory contains quickstart files for each supported model family, including examples/quickstart.py for the general case and model-specific files such as examples/cosmos3_video_quickstart.py and examples/higgs_audio_v3_quickstart.py.

## VLA Control: Production Frontends for Pi0, GROOT, and Pi0-FAST

The primary stated use case for FlashRT is VLA control: running vision-language-action models at decision rates fast enough for physical robot control. The README reports demos for Pi0.5 on a Jetson AGX Thor reducing per-decision latency from 266.0 and 303.6 ms (two reference implementations) to 36.1 ms and then to 29.2 ms, moving from 3.8 Hz to 34.2 Hz of policy decisions. The README attributes this to the same checkpoint executing through the staged pipeline rather than through the original host framework.

For GR00T N1.7 on an RTX 5090, the README reports per-decision latency moving from 44.4 ms to 15.8 ms, changing the effective decision frequency from 22.5 Hz to 63.1 Hz. All four arms of the LIBERO benchmark task completion are cited as validating correctness alongside the latency improvement.

The LIBERO benchmark results are the stated validation method. The README notes that the two independent hosting frameworks (Isaac-GR00T and LeRobot) produce cosine-similar outputs (0.999995 on one decision from a byte-identical observation), which supports the claim that the optimized pipeline preserves model accuracy. The quickstart examples for each model family are in the examples/ directory.

## LLM and VLM Support in the Realtime Frame

Beyond VLA control, FlashRT provides single-stream LLM and VLM inference paths. The README demonstrates Qwen3-VL-8B on the unmodified transformers host, where attaching FlashRT's structure patterns to the existing host moves decode throughput from 65.6 to 145.7 tokens per second, a 1.76x increase over the host's own compiled form, with TTFT around 30 ms across all configurations.

For mixture-of-experts LLM inference, the README describes a Qwen3.6-35B-A3B model (67 GB of BF16 weights) running on a single 32 GB card after quantize_on_adopt regrids the expert banks to 22.3 GiB before the model reaches the device. Decode throughput reaches 284.9 tokens per second in the third configuration using the checkpoint's own draft head.

FlashRT also supports attaching to existing serving engines without forking them. The README shows Qwen3-8B inside both vLLM and SGLang at 144 seats bound by a hook that fires after the engine loads; vLLM moves from 99.1 to 145.3 tok/s and SGLang from 101.2 to 203.0 tok/s. This attachment pattern means teams using vLLM or SGLang can try FlashRT kernels inside their existing setup without migrating the full serving stack.

## FlashRT vs TensorRT and vLLM

TensorRT optimizes deep learning inference through tactic-search compilation: it runs many candidate kernel implementations on the target hardware, measures each, and freezes the winners into an engine. This produces highly optimized engines but requires recompilation when the model, input shape, or CUDA driver changes. It is the right tool for stable, high-throughput deployment of fixed models at fixed batch sizes.

FlashRT avoids compilation entirely. Its dataflows assemble at load time from pre-tuned kernel libraries, which means no upfront compilation cost and no rebuild when the driver updates. The trade-off is that the optimization work goes into the hand-written kernels rather than into automated search, which requires more engineering effort per supported model and means unsupported models cannot be adapted through configuration alone.

vLLM and SGLang are designed for high-batch LLM serving where maximizing tokens per second across many concurrent users is the metric. They use continuous batching, paged attention, and iteration-level scheduling to keep GPUs saturated. FlashRT targets the opposite regime: one or a few simultaneous requests where latency per token matters more than aggregate throughput. The README is specific: "FlashRT targets the small-batch realtime cell with hand-tuned kernels and no compile step."

For robotics applications, neither TensorRT nor vLLM addresses the sub-50ms per-decision latency constraint at single batch size, which is the gap FlashRT occupies.

## Limitations and Maintenance

FlashRT supports a specific set of model checkpoints with production-validated frontends. Models not in that set require new dataflow implementations, which the README describes as engineering work rather than configuration. The USAGE.md in the repository documents the supported paths; teams with models outside that list cannot simply point FlashRT at an arbitrary checkpoint and expect performance gains.

The build process requires CMake and a CUDA-compatible toolkit, adding setup complexity compared to pure Python packages. The optional flash-attn wheel introduces an additional dependency for GROOT and some other backends, and building it from source requires matching CUDA and PyTorch versions.

The last push to the repository was on 2026-09-25, and the most recent release is v0.2.0, published on 2026-09-11. The Apache 2.0 license allows commercial use and modification with attribution. The repository includes a NOTICE file alongside the license, and the README links to docs/COMMUNITY_AND_CREDIT.md, which describes the project's expectations for community contributions and credit attribution.

## Conclusion

FlashRT suits engineers who need sub-50ms per-decision latency for VLA control on Jetson or RTX hardware and who are willing to build CUDA kernels from source. Teams serving large-batch LLM workloads should use vLLM or SGLang, which are designed for that problem. Before adopting FlashRT, verify that your model checkpoint is among the supported production frontends, confirm your CUDA toolkit version meets the build requirements, and review docs/COMMUNITY_AND_CREDIT.md to understand the collaboration expectations the project sets for reuse of its designs.

## FAQ

### What AI models does FlashRT support for VLA control?

The README lists Pi0, Pi0.5, GROOT N1.6, GROOT N1.7, and Pi0-FAST as production VLA frontends, with validation on LIBERO where applicable. The same kernel set also supports BAGEL world-model research paths, Higgs Audio v3 TTS, Wan2.2 video-policy paths, and single-stream LLM inference with Qwen3.6-27B.

### Does FlashRT require ONNX export or TensorRT compilation?

No. FlashRT assembles model-specific dataflows from captured graphs and hand-written kernels at load time, without ONNX export, engine compilation, or per-driver rebuild. The setup.py documents a pip install plus a separate CMake build for the CUDA kernel extensions.

### What hardware does FlashRT support?

The README describes NVIDIA implementations spanning Jetson AGX Thor and Jetson Orin through A100 and RTX 4090 and 5090. AMD GPU, NPU, and other hardware support are described as actively expanding.

### Can FlashRT attach to an existing vLLM or SGLang deployment?

Yes, according to the README. FlashRT can attach inside vLLM or SGLang using a hook that fires after the engine loads, without forking the serving engine. The README reports throughput gains for Qwen3-8B inside both engines at 144 concurrent seats.

## Sources

- [flashrt-project/FlashRT on GitHub](https://github.com/flashrt-project/FlashRT)
- [Issues](https://github.com/flashrt-project/FlashRT/issues)
- [License: Apache-2.0](https://github.com/flashrt-project/FlashRT/blob/main/LICENSE)
- [README](https://github.com/flashrt-project/FlashRT/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/flashrt-project-flashrt
