# 100 Days of LLM Inference: A Structured Engineering Challenge

> 100 Days of LLM Inference is a public learning challenge that pairs daily runnable scripts with the book "Inference Engineering" by Philip Kiely (Baseten Books, 2026), covering the full stack from CUDA kernels to multi-cloud autoscaling. The repository documents 26 days of completed work on a home-lab cluster of two NVIDIA DGX Sparks, with five additional phases planned.

**elizabetht/100-days-of-inference** — 100 days of LLM inference engineering — daily posts, experiments, and visualizations

- Repository: https://github.com/elizabetht/100-days-of-inference
- Stars: 1,798 · Forks: 232
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/elizabetht-100-days-of-inference

## What This Challenge Is and How It Works

100 Days of LLM Inference is a solo learning challenge structured as a daily deep dive into inference engineering for large language models. The repository owner describes it as built around "Inference Engineering" by Philip Kiely (Baseten Books, 2026), which the README quotes as defining three layers: Runtime, Infrastructure, and Tooling. The challenge attempts to cover all three layers systematically.

Each entry in the challenge is a runnable script. The README states that all experiments run on a home-lab cluster of two NVIDIA DGX Sparks. This hardware context matters: topics like disaggregated prefill and decode, multi-GPU tensor parallelism, and Firecracker microVM profiling assume direct access to GPU nodes, not cloud notebooks.

The repository structure reflects the daily organisation directly: day01/ through day26/ at the root, plus a copy of the Inference Engineering PDF. The challenge is organised into six phases, with Phase 1 through part of Phase 3 completed in the 26 visible day directories.

## Phase 1: Single-Instance Runtime Optimisation

Phase 1 covers 18 days on single-GPU performance. The topics map directly to chapters in the Kiely book and progress from fundamentals to specific optimisation techniques.

Days 1 through 6 cover LLM inference mechanics, model internals and tokenisation, embeddings, transformer blocks and attention, the KV cache, and ops:byte ratio with arithmetic intensity. Days 7 and 8 address CUDA kernels, kernel fusion, PyTorch, and model file formats including ONNX and TensorRT.

Days 9 through 11 cover the major inference serving frameworks: vLLM with PagedAttention and continuous batching, SGLang with RadixAttention and structured outputs, and TensorRT-LLM compilation. Day 12 covers NVIDIA Dynamo for disaggregated serving.

Days 13 through 18 address quantisation (FP8, INT8, INT4, NVFP4, GPTQ, AWQ, SmoothQuant), speculative decoding with draft-target models and EAGLE, KV cache prefix caching, and model parallelism across tensor, expert, pipeline, and data dimensions.

## Phase 2: Infrastructure Scaling Across Clusters

Phase 2 runs from day 19 through day 26 and covers multi-GPU and multi-cloud infrastructure. This phase also maps to chapters in the Kiely book.

Topics include GPU architecture (SMs, memory hierarchy, HBM), GPU generations from Hopper through Blackwell and Rubin, multi-GPU instances and Multi-Instance GPU (MIG), containerisation with Docker and NVIDIA NIMs, autoscaling with concurrency and cold start considerations, routing and load balancing with queueing, multi-cloud capacity management, and zero-downtime deployment with cost estimation.

The Phase 2 topics assume the reader understands the Phase 1 runtime concepts. Autoscaling policies, for example, are discussed in the context of TTFT and throughput tradeoffs that Phase 1 established.

## Planned Phases 3 Through 6

The README documents four more phases that extend beyond the current 26 completed days. Phase 3 covers tooling: performance benchmarking, observability metrics and tracing, and client code for streaming and async protocols.

Phase 4 is described as deep implementation: building each major concept from scratch in Python, including a BPE tokeniser, scaled dot-product attention, Flash Attention in simplified tiled form, INT8 quantisation pipelines, GPTQ-style quantisation, speculative decoding simulation, KV cache managers, and prefix caching with hash-based deduplication.

Phase 5 covers production systems: writing production Dockerfiles for vLLM, NIM-compatible containers, autoscaling simulators, load balancers, priority request queues, blue-green deployments, Prometheus metrics, Grafana dashboards, and OpenTelemetry distributed tracing.

Phase 6 extends inference engineering to modalities beyond text: vision language model inference with image preprocessing, embedding model batching, ASR, and other multi-modal topics referenced in the Kiely book.

None of Phases 3 through 6 are implemented in the repository. The day directories stop at day26.

## Limitations and the Hardware Dependency

The challenge has two structural limitations that determine who can follow it usefully.

First, the book "Inference Engineering" by Philip Kiely is a required companion. The README maps every day and phase to a specific chapter of the book. Without the book, the day topics become headings without the underlying explanations the challenge assumes the reader has already read.

Second, the experiments run on a home-lab cluster of two NVIDIA DGX Sparks. DGX Spark is a compact but high-performance GPU system aimed at AI development. Phase 5 explicitly names spark-01 and spark-02 as the deployment targets for vLLM, SGLang, and TensorRT-LLM benchmarks. Cloud GPU rentals could substitute for some experiments, but others assume persistent cluster access and multi-node configuration.

The last push was on 2026-04-30, meaning the challenge is paused or incomplete as of the review date. Following it as a live, daily resource is not possible.

## Comparison With Reading the Book Directly

The repository is a companion to the book, not a replacement. "Inference Engineering" by Philip Kiely is published by Baseten Books in 2026 and the README quotes it and maps every topic to specific chapters.

A reader who works through the book without the companion scripts gets the conceptual explanation but no hands-on implementation. The scripts in this repository provide the implementation side: profiling with `torch.profiler`, deploying vLLM, benchmarking TTFT and throughput, and writing CUDA kernels via Triton are all listed as Phase 4 and Phase 5 projects.

An alternative that serves a similar purpose is Hugging Face's model deployment and optimisation tutorials, which are freely available and regularly updated. The difference is that Hugging Face tutorials focus on the ecosystem around Transformers, while this challenge follows the specific structure and scope of the Kiely book and uses NVIDIA-specific tools throughout.

## Conclusion

This repository is best suited for ML engineers and systems programmers who want to follow a structured, practitioner-led study path through LLM inference and already have access to GPU hardware. Without the book by Philip Kiely it references, the day-by-day structure loses context. The last push was on 2026-04-30, and the repository covers only 26 of the planned days, so prospective followers should treat it as a work in progress rather than a complete course.

## FAQ

### What are some good resources for learning inference engineering?

The challenge is built around "Inference Engineering" by Philip Kiely (Baseten Books, 2026), which the README describes as the primary reference. The repository provides daily runnable scripts that implement the book's concepts, covering topics from vLLM and SGLang to multi-GPU tensor parallelism.

### How expensive is LLM inference?

The challenge includes a GPU cost model project in Phase 5: building a model for cost per token across instance types at different GPU utilisations. The Phase 2 topics also include zero-downtime deployment and cost estimation, mapping to chapter 7.4 of the Inference Engineering book.

### What GPU hardware does this challenge require?

The README states that all experiments run on a home-lab cluster of two NVIDIA DGX Sparks. Phase 5 specifically benchmarks across spark-01 and spark-02. Cloud GPU alternatives could substitute for some experiments, but the repository does not document cloud-based setup instructions.

## Sources

- [elizabetht/100-days-of-inference on GitHub](https://github.com/elizabetht/100-days-of-inference)
- [Issues](https://github.com/elizabetht/100-days-of-inference/issues)
- [README](https://github.com/elizabetht/100-days-of-inference/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elizabetht-100-days-of-inference
