# InfraTech: A Notebook Collection for AI Infrastructure Engineers

> InfraTech is a Chinese-language collection of Python Jupyter notebooks covering AI infrastructure topics including inference optimization, distributed training, attention mechanisms, and LLM framework internals. It is aimed at engineers who want annotated, runnable code alongside detailed technical explanations.

**CalvinXKY/InfraTech** — 分享AI Infra知识&代码练习：PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等

- Repository: https://github.com/CalvinXKY/InfraTech
- Stars: 4,094 · Forks: 398
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/calvinxky-infratech

## What InfraTech Is and Who It Is For

InfraTech is a personal knowledge-sharing repository maintained by CalvinXKY. The repository's stated purpose is to introduce AI infrastructure knowledge through annotated Python notebooks, covering training and inference frameworks, performance optimization, deep learning fundamentals, and hardware-level details. All prose explanations are in Chinese.

The target reader is an engineer already working in or studying the AI infrastructure domain. The content assumes familiarity with deep learning concepts; notebooks are rated by difficulty with a single lightning-bolt symbol for introductory material, two for intermediate, and three for advanced. Each notebook typically links to a longer technical article on Zhihu, the Chinese engineering blog platform, where the author elaborates on the topic.

The last push was on 2026-09-20. The repository carries no license, which means the default copyright rules apply: the code and text are owned by the author and cannot be redistributed or used commercially without explicit permission.

## Repository Layout: Three Main Directories

The repository organizes notebooks into three primary directories. The `llm_infer/` folder holds the largest collection, covering inference-specific topics from scheduling and sampling to parallelism and memory management. The `models/` directory contains notebooks on individual model components such as RoPE and MLA (Multi-head Latent Attention). The `docs/` directory holds shared reference material.

Additional directories address specific topics: `deeplearning_framework/` covers distributed training primitives, `deepseek_v3/` holds notebooks specific to the DeepSeek V3 architecture, `multi_lora/` addresses LoRA and Multi-LoRA techniques, `pytorch_vista/` covers PyTorch internals, and `rl/` addresses reinforcement learning as used in LLM training.

There is no package to install and no command-line interface. To use InfraTech, you clone the repository and open the notebooks in a Jupyter environment. The README does not document specific setup steps, but the notebooks use Python and require standard ML libraries such as PyTorch.

## Inference Notebooks: Scheduling, Sampling, and Speculative Decoding

The `llm_infer/` directory contains the most substantial cluster of notebooks. Several cover inference fundamentals that engineers encounter when optimizing LLM serving.

Chunked Prefill and Flash Decoding are explained in `chunked_prefill_and_flash_decoding.ipynb`, a two-bolt intermediate notebook. Speculative decoding, which uses a smaller draft model to propose tokens that a larger model then verifies, is covered in `speculative_decoding.ipynb`. A follow-on notebook, `dflash_dspark_principle.ipynb`, addresses parallel speculative decoding schemes called DFlash and DSpark at the two-bolt level.

The `LLM_sampling.ipynb` notebook covers sampling strategies for token generation. The `quantization.ipynb` notebook surveys quantization basics for inference at an introductory level. The `kv_cache_transfer_vs_recomputation.ipynb` notebook examines whether pooled KV cache data is faster than recomputation from scratch, and `linear_attention_kv_cache_size.ipynb` explores the storage advantages of linear attention mechanisms.

The `parallel_strategies.ipynb` notebook introduces the five main parallelism strategies for LLM inference: data parallelism, tensor parallelism, pipeline parallelism, sequence parallelism, and expert parallelism. A related notebook, `ulysses_mha_demo.ipynb`, covers the Ulysses parallelism optimization for multi-head attention.

## Framework-Specific Notebooks: vLLM and SGLang

A distinct strength of InfraTech is its framework-specific content for vLLM and SGLang, two widely-used open-source LLM serving frameworks.

For vLLM, `vllm_basic_scheduler.ipynb` walks through building a basic scheduler from scratch, rated at two bolts. The `vllm_mem_snapshot.ipynb` notebook covers GPU memory visualization and management. The `cuda_graph.ipynb` notebook addresses why vLLM does not apply CUDA graph optimization during the prefill phase, answering a specific design question from the vLLM codebase.

For SGLang, `sglang_radix_attention.ipynb` is a two-bolt notebook on implementing RadixAttention, the data structure SGLang uses to share prefix cache across requests. The `sglang_profiling_from_scratch.ipynb` notebook covers profiling data collection and analysis. The `switch_role_update_weights.ipynb` notebook covers how SGLang and vLLM implement seamless weight switching between training and inference roles on shared GPU hardware, reducing overhead in reinforcement learning workflows.

The `nano_vllm.ipynb` notebook provides an entry point using Nano-vLLM, a minimal reimplementation intended for building a conceptual understanding of serving framework internals.

## Attention Mechanisms and Model Architecture Notebooks

The `models/` and `deepseek_v3/` directories contain notebooks on attention mechanism details that are relevant for engineers working with modern transformer variants.

The `rope_principle.ipynb` notebook in `models/modules/` covers the RoPE (Rotary Position Embedding) calculation from one-dimensional to three-dimensional cases, rated as advanced with three bolts. The `MLA_diff_mode_mfu_calculation.ipynb` notebook in `deepseek_v3/` provides a detailed breakdown of MLA (Multi-head Latent Attention) computation flow and absorption matrix comparison, also rated advanced.

The `collective_operations.ipynb` notebook in `deeplearning_framework/` addresses collective communication primitives used in distributed training and inference, such as all-reduce and all-gather, rated at one bolt as an entry-level topic. The `prefix cache` notebook examines why prefix caching in attention computation can have near-zero overhead.

## Limitations for Non-Chinese Readers and Commercial Users

The most immediate limitation is language: all explanatory text in the notebooks and the README is written in Chinese. Engineers who do not read Chinese cannot use the explanations, though they could read the code itself. The linked Zhihu articles, which provide deeper context for most notebooks, are also in Chinese.

The repository carries no license. Under default copyright law, the code and prose belong to the author. Engineers who want to reuse, adapt, or redistribute any notebook content in a commercial project cannot do so without explicit permission.

There is no formal course structure. Notebooks are organized by topic and difficulty, but there is no recommended learning path, no exercises with solutions, and no progress tracking. The difficulty ratings give rough guidance, but readers are expected to navigate independently.

The official documentation and source code for vLLM and SGLang are available in English and cover the same technical territory in greater depth. InfraTech complements those sources with concise, annotated code examples rather than replacing them.

## Comparing InfraTech to Official Framework Documentation

The closest alternative reference for the same topics is the official documentation and developer guides published by the vLLM and SGLang projects themselves. Both projects maintain English-language documentation with architecture overviews, configuration references, and contribution guides. The key difference is that InfraTech provides compact, self-contained notebooks with Python code you can run and inspect directly, while official documentation tends toward reference material and conceptual explanations without executable code alongside the prose.

For engineers who learn by running code and reading annotated implementations rather than reading prose documentation, InfraTech's notebook format has a practical advantage. For engineers who need production-grade detail, version-specific behavior, or English explanations, the official sources are more complete.

The Multi-LoRA notebook in `multi_lora/LoRA_to_Multi_LoRA.ipynb` covers techniques for serving multiple LoRA adapters simultaneously, a topic covered less thoroughly in the official serving framework documentation. That makes it one of the more distinctive contributions in the collection.

## Conclusion

InfraTech suits engineers who read Chinese and want hands-on notebook code for AI infrastructure topics, particularly vLLM and SGLang internals, inference optimization, and distributed parallelism. It is not a course and not an installable library. The repository carries no license, which limits its use in commercial or redistributed contexts. Readers who need English-language material or a structured curriculum should look elsewhere.

## FAQ

### What topics does the InfraTech repository cover?

InfraTech covers inference optimization (chunked prefill, speculative decoding, KV cache, quantization, sampling), distributed parallelism (DP/TP/PP/SP/EP), attention mechanisms (RoPE, MLA), vLLM internals (scheduler, memory, CUDA graph), SGLang internals (RadixAttention, profiling, weight switching), and Multi-LoRA serving, all in annotated Python notebooks.

### Is InfraTech suitable for English-speaking engineers?

The explanatory text and README are written in Chinese. The Python code in the notebooks is readable without language knowledge, but the accompanying prose explanations and linked Zhihu articles are in Chinese. Engineers who do not read Chinese will lose most of the explanatory value.

### Can I use InfraTech notebooks in a commercial project?

The repository carries no license. Under default copyright rules, the code and content are owned by the author. Commercial use, redistribution, or adaptation requires explicit permission from CalvinXKY.

## Sources

- [CalvinXKY/InfraTech on GitHub](https://github.com/CalvinXKY/InfraTech)
- [Issues](https://github.com/CalvinXKY/InfraTech/issues)
- [README](https://github.com/CalvinXKY/InfraTech/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/calvinxky-infratech
