What is LLM inference server?
An LLM inference server is a process that loads a language model into memory and answers generation requests over an API, usually HTTP. It sits between an application and the model weights, handling batching, memory and scheduling so callers do not have to.
How an inference server actually works
A model file on disk is just weights and a graph description. An inference server turns that into a long-running service. At startup it reads the weights, allocates device memory (GPU VRAM, unified memory on Apple Silicon, or plain RAM), and builds the runtime that executes the forward pass. After that it listens on a socket, typically an HTTP endpoint that mimics the OpenAI chat completions shape, so existing client libraries work without changes.
The interesting part is what happens between the request and the token. Generation is autoregressive: each new token depends on every token before it. A naive server decodes one request at a time and leaves the GPU mostly idle, because the expensive matrix multiplications are memory-bandwidth bound at batch size one. Two mechanisms fix this. Continuous batching admits new requests into the running batch as soon as a sequence finishes, instead of waiting for the whole batch to drain. Paged attention splits the KV cache, the per-token key and value tensors that make generation fast, into fixed-size blocks that can be allocated and freed like virtual memory pages. vLLM's README and our earlier analysis describe exactly this pairing: paged KV-cache memory management plus continuous batching, in exchange for a heavy CUDA build.
A request therefore passes through several stages: tokenisation, admission to a scheduler queue, prefill (processing the prompt in one pass), decode (one token per step), and detokenisation of the streamed output. The scheduler decides how many sequences share a step, how much cache each may hold, and which requests get preempted when memory runs short. Preemption is where servers differ most in practice. Some recompute the prefill later; some swap blocks to host memory; some simply queue. Each choice shows up as latency under load.
Servers also expose knobs that change behaviour: maximum context length, tensor parallelism across GPUs, quantisation format, and cache eviction policy. These are not cosmetic. Setting the context window larger than the KV cache can hold forces eviction or rejection, and the failure mode is usually a truncated answer or an out-of-memory error rather than a clean message.
When you need a server, and when you do not
You need an inference server when more than one client shares a model, when requests arrive concurrently, or when you want a stable API instead of a script. A single developer running one prompt in a notebook does not need one. A batch job that processes a fixed list of texts can call the model directly and skip the HTTP layer entirely.
The line is concurrency plus reuse. If the model takes 30 seconds to load and you call it once an hour, the load time dominates and a server just adds a process to babysit. If ten users hit the same model at once, a server earns its keep through batching and cache reuse. The same logic applies to memory: a 70B model in 16-bit precision needs roughly 140 GB just for weights, so serving it on one device is a hardware question before it is a software one.
There is also a hardware boundary. WebLLM runs quantised models inside the browser through WebGPU and removes the server entirely, which is the right trade when the data must not leave the client. The cost is stated in its own documentation: browsers without WebGPU have no fallback. On the other end, Dynamo coordinates many GPUs and nodes and explicitly does not replace your inference engine, so it is only relevant when the bottleneck is cluster-level rather than single-GPU.
A useful test: if you can describe your workload as one model, one process, one machine, and low concurrency, start with the engine's own offline API. Add a server when you can point to queueing, shared cache, or multiple callers as the actual problem.
Common pitfalls and limits
The first pitfall is treating the server as free. Every layer adds latency and a failure mode. HTTP parsing, tokenisation, scheduling and detokenisation all cost time, and a server that batches aggressively can increase time-to-first-token for a single user. Continuous batching improves throughput, not necessarily the latency of any one request.
The second is memory arithmetic done late. KV cache size scales with batch size, context length, layer count and head dimension. A configuration that fits one long conversation may fail with eight short ones. Servers differ in how they handle that pressure: some preempt and recompute, some swap to CPU, some reject. The README of a given project is the place to check, and several are silent on the exact policy.
The third is portability. vLLM's strength, paged attention, is tied to a CUDA build, and our analysis notes that this is the trade it makes. MNN targets ARM phones, PCs and IoT boards and is a poor fit if you want a Python-first workflow or a wide catalogue of prebuilt desktop wheels. PowerInfer depends on ReLU-sparse models and a single consumer GPU, and its own README limits how far that idea travels. oMLX is honest about its own limits: custom kernels need a full Xcode install, and some model families fall back to much slower generic paths. None of these are defects in the abstract; they are constraints you inherit.
The fourth is operational. Triton Inference Server supports TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL across GPU, x86 and ARM CPU and AWS Inferentia, and the cost is a container release cadence you have to track. A serving stack is a dependency with its own upgrade schedule.
Finally, packaging. FreeToken's engine is credible, but our analysis says its packaging story is not yet finished. That gap matters more than benchmark numbers when you are trying to deploy.
How it shows up in open-source projects
The projects below are not interchangeable. Some are servers, some are engines, some sit above or beside the server.
vllm-project/vllm is the clearest example of the concept itself: an Apache-2.0 inference and serving engine whose README and our analysis centre on paged KV-cache memory management and continuous batching. It is the reference point for high-throughput serving on CUDA hardware.
jundot/omlx is a Python inference server for Apple Silicon. It keeps KV cache in a hot memory tier and a cold SSD tier, exposes an OpenAI-compatible endpoint on port 8000, and is managed from the macOS menu bar. Its documented limits are the useful part: custom kernels need full Xcode, and some model families fall back to much slower generic paths.
antirez/ds4 (DwarfStar) is a C inference engine built for a handful of models rather than every GGUF file. It targets DeepSeek V4 Flash first, adds GLM 5.2, and asks for 96 GB of unified memory or more on Macs. That is an engine, not a general server.
mlc-ai/web-llm runs quantised models client-side through WebGPU with an OpenAI-shaped API. It removes the server, and with it any fallback for browsers that lack WebGPU.
alibaba/MNN is a C++ inference and training framework for phones, PCs and IoT boards, with an LLM runtime and a Stable Diffusion runtime layered on top. It fits local ARM deployment and is the wrong choice for Python-first workflows.
Tiiny-AI/PowerInfer splits inference between GPU and CPU by preloading frequently activated neurons and computing the rest on the CPU. It targets ReLU-sparse models on a single consumer GPU.
FlashML-org/FreeToken is an Apache-2.0 Python runtime that treats GPU, CPU and host memory as one elastic platform for frontier-scale Mixture-of-Experts models on laptops and desktops. The engine is credible; packaging is unfinished.
triton-inference-server/server is NVIDIA's model server for TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL across GPU, x86 and ARM CPU and AWS Inferentia. It pays off with several frameworks and dynamic batching, and costs a container release cadence.
ai-dynamo/dynamo is a Rust-and-Python serving stack that coordinates multiple GPUs and nodes rather than replacing your engine. It matters only when the bottleneck is cluster-level.
Two projects sit adjacent rather than inside the category. zylon-ai/private-gpt is not a model runner: it is an API layer in front of any OpenAI-compatible inference server, supplying retrieval, tools, MCP and database access. That makes it a consumer of servers like the ones above.
In practice
An inference server is worth adding when concurrency, shared memory or a stable API is the actual problem, and it is worth skipping when one process, one model and one caller will do. Read the engine's README for its batching and cache policy before you commit, since that is where the real limits live. If you are starting today, try vLLM on CUDA hardware or oMLX on an Apple Silicon Mac, and compare time-to-first-token under concurrent load rather than single-request speed.