Model or dataset
patchy631/time-to-first-token avatar
patchy631/time-to-first-token

patchy631/time-to-first-token: a 10-week roadmap that ships one inference service

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

965 stars108 forksHTMLApache-2.0

At a glance

What is it?
The repository is a 50-session, 25-hour plan for learning LLM inference serving by building a single OpenAI-compatible service on rented GPUs. It is study material, not a serving framework, and its value depends on whether you actually rent the hardware.
Who is it for?
Adopt this roadmap if you already write Python and understand transformer architecture, have a few hundred dollars of GPU rental budget, and want one instrumented service instead of a folder of notebooks. Skip it if you need a working serving stack today, if you cannot rent a 24GB card, or if you want a reference implementation rather than a curriculum.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 47 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the roadmap actually asks you to build

The README opens with a specific claim: fifty sessions, each feeding one artifact. That artifact is an OpenAI-compatible inference service you deploy on a rented GPU, instrument, load test past 1000 concurrent requests, optimize with quantization and speculative decoding, and put behind a cost-aware router. The repository itself contains an index.html, a progress.md, an assets folder, a LICENSE and the README. There is no server code to clone. The deliverable is something you write while following the sessions.

The audience is narrow and stated plainly. You should be comfortable with Python, transformers at the architecture level, and the command line. Prior serving, Kubernetes or CUDA experience is not required. That combination is unusual for a serving curriculum, most of which assume you have already deployed something and want to tune it. Here the assumption is the reverse: you understand attention and KV cache shapes on paper, and you have never watched a queue depth graph during a load test.

The cost estimate is the part most readers will skim past. A 7-8B model needs roughly one 24GB GPU, rented by the half hour, plus two H100 sessions. The README's own arithmetic puts the whole thing at 25 hours across 10 weeks, with two buffer days per week. Buffer days are for catching up, not new material, which is a realistic admission that a 10-week plan will slip.

Why the roofline model comes before any tuning knob

The ordering section is the most opinionated part of the README, and it differs from how most inference reading lists are written. The roofline model lands in week 1, before vLLM, before benchmarking, before any optimization. The stated reason is that every later optimization is a move on the same plot. Quantization attacks memory-bandwidth-bound decode. Continuous batching raises arithmetic intensity toward the compute roof. Speculative decoding spends FLOPs, cheap in the memory-bound regime, to cut sequential memory loads. Disaggregation exists because prefill is compute-bound and decode is bandwidth-bound and they contend for one GPU.

That framing has a practical consequence. If you cannot say which phase of your workload is bandwidth-bound, the tuning sessions become knob-turning. The week 1 deliverable is a hand-derived arithmetic intensity figure for the model you will serve, not a summary of the reading. Friday is marked build and skim, with the note that the value is producing your own figure.

Measurement follows the same logic. Instrumentation lands in week 3 and load testing in week 5, both before the tuning knobs. The README's justification is blunt: nothing after them is verifiable without them, and a public benchmark is mostly a credible measurement harness with a model attached. That is a defensible position and it explains why the roadmap front-loads Grafana dashboards and a load-test harness rather than jumping to FP8.

Install and first session: vLLM, GuideLLM, SGLang

The setup section lists three packages and one container runtime. All three install from PyPI; the README gives no version pins, which matters because vLLM's scheduler and metrics surface change between minor releases.

bash
pip install vllm
pip install guidellm
pip install sglang

Docker is needed for the Prometheus and Grafana stack, and kind or a small managed cluster for week 8. The README also lists GPU providers with a short note on each: RunPod for per-second billing and quick pods, Modal for serverless benchmark sweeps, Lambda and vast.ai for cheap on-demand and marketplace GPUs, and Colab for pure-Python sessions with no serving.

Before renting anything, read the access note carefully. All of week 1, most of week 3, the week 6 lecture days, and all of week 9 reading need no GPU. One H100 is needed for exactly two sessions: the 1000-concurrent test in week 5, and disaggregation in week 8 if you run it for real. Batching rentals around the build days is the difference between a cheap course and an expensive one.

The first real use is week 2: stand up an OpenAI-compatible endpoint serving a 7-8B model, then read the block manager and scheduler until the paging design is obvious from the code. The README marks PagedAttention as text-first because the paper is more precise, and the V1 architecture as video-first. The deliverable is the running endpoint, not notes about it. Expect to spend the GPU time on the deploy and the reading time off the clock.

What the roadmap deliberately does not build

Two of the most commonly reimplemented components are excluded on purpose. Paged attention and continuous batching are not build exercises, because vLLM implements both and chunked prefill is the default scheduling strategy in vLLM and SGLang. The README argues that reimplementing them teaches less than reading the block manager and the scheduler until you can explain from the code why memory waste drops under 4 percent and why the GPU stops idling between requests.

That is a real trade-off, and it cuts both ways. You finish the roadmap able to read vLLM's internals and reason about its behavior, but you will not have written a paged attention kernel or a scheduler. If your goal is to contribute patches to a serving engine rather than operate one, this roadmap leaves a gap you will have to fill elsewhere. The README does not claim otherwise.

Edge deployment gets the same treatment for a different reason. ONNX Runtime, TensorRT-LLM and WebLLM are placed last and marked optional, because they are client-side and embedded runtimes that share almost no operational surface with the datacenter stack the other nine weeks build. Including them earlier would dilute the single-service thread. The cost is that browser and on-device inference, which is where a lot of production latency work now happens, is out of scope.

The router week and why unit economics sit next to it

Week 8 pairs the cost/latency/quality router with unit economics, and the README explains why. A router that picks a backend by cost, latency and quality is the one component that forces a dollar figure and a latency budget into code. Studying economics the same week turns the reading into a routing policy instead of a blog post you agreed with.

That is the sharpest design decision in the plan. Most curricula treat cost as a sidebar. Here it is a required input to a component you have to write, with per-request token budgeting as a stated end goal. The router also forces a decision the rest of the roadmap can defer: which variant of your model serves which traffic. FP16, FP8, INT4, speculative decoding and KV eviction are all benchmarked as variants, and the router is where those measurements become a policy rather than a table.

The weak point is that the README does not describe the router's interface, the backends it can dispatch to, or how quality is scored. Those details presumably live in the week 8 sessions inside index.html, which the README does not reproduce. If you are evaluating the roadmap for a team, read that week before assuming the router is more than a sketch.

Alternatives and what they trade away

The obvious alternative is the vLLM Quickstart plus the PagedAttention paper, read directly. That path gets you a running endpoint in an afternoon and a precise explanation of the memory layout, and it costs nothing. What it does not give you is a measurement harness, a benchmark methodology, or a reason to keep going past the first deployment. The roadmap's own framing is that seventeen disconnected experiments spend most of their time on setup, and one service that keeps growing gets the same coverage and leaves you with something to show. Whether that is true depends on your discipline: a self-directed engineer with a load-testing habit can assemble an equivalent path from the same primary sources.

A second alternative is a managed inference platform, where you never touch the scheduler or the KV cache. That removes the entire subject of the roadmap. It is the right choice if you need throughput next week and not expertise next quarter.

A third is a structured course with graded assignments and a cohort. The roadmap has no grading, no cohort and no deadlines beyond the ones you set. The buffer days acknowledge that self-paced plans slip. If external accountability is what makes you finish things, this format will not supply it.

Maintenance, licence and what to check before starting

The repository is not archived, and the last push was on 2026-08-14. That is recent enough that the plan reflects the current vLLM and SGLang surface, but there are no releases, so there is no versioned snapshot to pin against. You track main.

The practical upgrade risk sits in the linked material rather than the roadmap text. The sessions point at external posts, papers and lectures, and those links can rot or be revised independently of this repository. The README also gives no version pins for vllm, guidellm or sglang, so a scheduler change in vLLM can invalidate a session's expected output. If you hit that, the fix is to install an older vLLM rather than to skip the session, but the README does not document rollback or a compatibility matrix.

The licence is Apache-2.0, which permits commercial use, modification and redistribution with the usual conditions around notices and patent grants. The repository is a curriculum that links to third-party papers, videos and blog posts; those carry their own terms, and Apache-2.0 covers this repository's own files only. Nothing here is legal advice, and if you plan to reuse the material inside a commercial training program, check how the linked sources are licensed separately.

Editorial conclusion

Adopt this roadmap if you already write Python and understand transformer architecture, have a few hundred dollars of GPU rental budget, and want one instrumented service instead of a folder of notebooks. Skip it if you need a working serving stack today, if you cannot rent a 24GB card, or if you want a reference implementation rather than a curriculum. Before committing, open index.html in a browser and confirm the week 2 and week 5 sessions match your current vLLM version, since the troubleshooting notes are written against pinned versions the README does not list.

Frequently asked questions

What does time to first token mean in this roadmap?

TTFT is one of the metrics the roadmap has you instrument, alongside inter-token latency, throughput, queue depth and cost per request. The README lists Grafana dashboards showing TTFT as part of the finished artifact. The repository does not define the metric itself; it treats it as something you measure on your own service.

How do I measure time to first token while following patchy631/time-to-first-token?

Instrumentation lands in week 3 and load testing in week 5, both before the tuning sessions. The README states that nothing after them is verifiable without them. The exact harness is not shown in the README, which points to the week 3 and week 5 sessions instead.

How do I improve time to first token in this roadmap?

The roadmap ties TTFT improvements to the roofline model: prefill is compute-bound and decode is bandwidth-bound, and disaggregation exists because they contend for one GPU. Quantization, continuous batching and speculative decoding are each framed as moves on that same plot. The README does not give TTFT targets or claim a specific reduction.

How do I reduce time to first token in this roadmap?

The roadmap places measurement before optimization: instrumentation in week 3 and load testing in week 5 come before the tuning knobs, because nothing after them is verifiable without them. The optimization sessions then work from the roofline model rather than from a list of flags. The README does not promise a specific reduction.

What affects time to first token according to this roadmap?

The roadmap's framing is that prefill is compute-bound and decode is bandwidth-bound, and that disaggregation exists because the two phases fight over one GPU. Quantization, continuous batching and speculative decoding are each described as moves on the same roofline plot. The README does not enumerate a full list of contributing factors.

What is the difference between TPOT and ITL?

The README lists inter-token latency as one of the metrics your Grafana dashboards should show, next to TTFT, throughput, queue depth and cost per request. It does not define TPOT or spell out how it relates to inter-token latency. The repository treats both as things you measure on the service you build.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. patchy631/time-to-first-token on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/patchy631-time-to-first-token.svg)](https://hysenlabs.com/projects/patchy631-time-to-first-token)