AI-fundamentals: A Structured GPU and LLM Infrastructure Learning Repository
AI 基础知识 - GPU 架构、CUDA 编程、大模型基础及AI Agent 相关知识。
At a glance
- What is it?
- AI-fundamentals is an open-source repository from ForceInjection that collects Chinese-language technical guides covering GPU architecture, CUDA programming, large language model theory, Kubernetes-based AI deployment, and distributed inference. It is organized as a navigable reference for AI engineers, system architects, and GPU developers who need depth across the full AI infrastructure stack.
- Who is it for?
- AI-fundamentals is worth bookmarking for AI infrastructure engineers who read Chinese and need a single reference that covers GPU hardware, CUDA programming, Kubernetes-based deployment, and LLM inference in one place. Engineers who need English-language documentation will not find it here; the repository's primary language is Chinese with HTML as the rendering format.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A Reference for Engineers Who Work Across the Full AI Stack
Most documentation for AI infrastructure lives in isolated silos: NVIDIA's documentation covers CUDA, Kubernetes documentation covers scheduling, and LLM framework documentation covers model serving. They rarely explain how these layers interact. AI-fundamentals is structured as an integrated learning path that covers the complete technology stack from hardware up to agentic systems.
The README describes the target audience as AI engineers, system architects, GPU programming developers, large model application developers, and technical researchers. The technology stack listed in the README includes CUDA, GPU architecture, LLM, AI systems, distributed computing, containerized deployment, and performance optimization. The repository is not a course with exercises or assessments; it is a curated collection of technical guides in Markdown and HTML, organized into numbered topic directories. Clone the repository, then navigate to the directory for the topic you need.
Repository Layout: Ten Topic Directories
The repository is divided into numbered directories. The first covers hardware architecture and interconnect technologies: GPU and TPU design, PCIe and NVLink buses, GPUDirect P2P and RDMA, CXL interconnect, the NVIDIA GB300 NVL72 architecture, and an AI latency pyramid model. The second covers GPU and CUDA programming: container environment setup, CUDA thread blocks and grids, CUDA streams, SIMT versus tile-based programming models, and TileLang kernel development.
Directory three covers AI cluster operations including GPU monitoring with `nvidia-smi` and `nvtop`, InfiniBand network architecture and health checks, and NCCL distributed communication benchmarks. Directory four covers cloud-native AI: NVIDIA Container Toolkit, Kubernetes device plugins, Kueue and HAMi GPU scheduling, LWS distributed inference, and high-performance storage with JuiceFS and DeepSeek 3FS. Directory five covers model training and fine-tuning; six covers LLM theory; seven covers RAG and tool use; eight covers agentic systems; nine covers inference. Directories ten and eleven cover AI courses and AI-native application patterns.
The Hardware and CUDA Content in Detail
The hardware section goes deeper than most introductory resources. The README lists specific articles on GPUDirect P2P between GPUs on the same node, GPUDirect RDMA for cross-node storage access, NVLink-C2C chip-to-chip interconnect, and the NVIDIA GB300 NVL72 rack-scale architecture. Each article is a separate Markdown file in the `01_hardware_architecture/` subdirectory with its own path listed in the README's table of contents.
The CUDA section covers the programming model from environment setup through kernel optimization. The README lists guides on GPU parallel computing basics, CUDA core concepts including thread blocks and grids, CUDA stream concurrency, and a comparison of SIMT versus tile-based programming models. A guide on TileLang covers its syntax, operator development, and performance tuning. A `nvbandwidth` best practices guide covers memory bandwidth and PCIe transfer measurement. The README also links to an external resource, CUDA-Learn-Notes, which it describes as covering over 200 Tensor Core and CUDA Core optimization kernel examples including HGEMM and Flash Attention implementations via MMA and CuTe.
Kubernetes AI Infrastructure Coverage
Directory four is the most operationally detailed section. It covers the NVIDIA Container Toolkit's underlying mechanism for giving containers GPU access, Kubernetes device plugin source analysis, and integration between Kueue job queuing and HAMi GPU sharing. HAMi provides fine-grained GPU virtualization at the container level; the README links to a usage guide, a Prometheus metrics reference, and a comparison with KAI Scheduler.
The distributed inference section covers LWS (Leader Worker Set), which the README describes as a Kubernetes-native abstraction for distributed LLM training and inference scheduling, and `llm-d`, described as an LLM inference architecture built on Kubernetes. The storage subsection covers JuiceFS (with a focus on its data-metadata separation architecture), DeepSeek 3FS design notes, and NVIDIA's ICMS architecture for KV Cache storage in inference scenarios. These topics are not typically covered in a single resource and represent the most specialized content in the repository.
The GPU resource management subsection includes four parts: basic theory, virtualization technologies (hardware-level, kernel-mode, and user-mode), resource management and optimization covering CUDA streams and MPS scheduling, and a practical deployment guide. Additional dedicated articles cover HAMi resource isolation and Flex AI configuration for production environments.
What This Repository Does Not Cover
AI-fundamentals does not provide runnable code examples, exercises, or lab environments. It is a documentation collection, not an interactive course. Engineers looking to practice CUDA programming will need to find separate exercise repositories; the README links to CUDA-Learn-Notes as an external resource with kernel examples, but the AI-fundamentals content itself is explanatory rather than hands-on.
The repository is also written primarily in Chinese. Engineers who read English but not Chinese will not benefit from the prose content. Package names, command-line tools, API names, and configuration keys appear in their original form, but the explanations are in Chinese. Official NVIDIA documentation and the Kubernetes documentation cover many of the same topics in English with additional depth and vendor support, at the cost of not connecting the layers together the way this repository does.
The repository also does not cover model fine-tuning workflows in the same depth as the hardware and infrastructure sections. Directory five is listed as covering model training and fine-tuning, but the README's table of contents for that section is not as detailed as directories one through four, suggesting it is less developed at present.
Maintenance and License
The last push to the AI-fundamentals repository was on 2026-09-27. The repository is not archived. It has two releases: v3.0 published on 2026-01-08 and v2.0 published on 2025-08-29. The release cadence suggests major content reorganizations happen roughly every few months. The Apache 2.0 license permits unrestricted use, modification, and redistribution of the content.
The top-level directory includes a `CLAUDE.md` and an `AGENTS.md` file, indicating the repository documents conventions for AI coding assistants working on its content. The `.github/` directory suggests CI or issue template configuration is in place. Because the primary language is HTML (as listed in the repository metadata), some guides may be rendered as HTML rather than plain Markdown, which affects how they display when browsed on GitHub versus accessed through the project's website at forceinjection.github.io.
There is also a `.pre-commit-config.yaml` file, suggesting the repository enforces formatting or linting checks on contributions. The `scripts/` directory and `img/` directory suggest build or illustration assets are used in the guides. The `99_misc/` directory at the end of the numbered sequence suggests miscellaneous reference material that does not fit the main topic categories.
Editorial conclusion
AI-fundamentals is worth bookmarking for AI infrastructure engineers who read Chinese and need a single reference that covers GPU hardware, CUDA programming, Kubernetes-based deployment, and LLM inference in one place. Engineers who need English-language documentation will not find it here; the repository's primary language is Chinese with HTML as the rendering format. The last push was on 2026-09-27, and the Apache 2.0 license permits unrestricted use and redistribution. The repository does not replace official product documentation from NVIDIA, Kubernetes, or any of the tools it covers, but it connects those tools in ways that official documentation does not.
Frequently asked questions
What topics does the AI-fundamentals repository cover?
The README lists ten major topic areas: hardware architecture and interconnects, GPU and CUDA programming, AI cluster operations, cloud-native AI infrastructure on Kubernetes, model training and fine-tuning, LLM theory, RAG and tool use, agentic systems, inference systems, and AI-native application patterns. Each area is a numbered directory in the repository root.
What language is the AI-fundamentals content written in?
The content is written primarily in Chinese. The repository description in the README is in Chinese, and the article text throughout the directories is in Chinese. Commands, package names, API names, and configuration keys appear in their original form.
Does AI-fundamentals require any specific tools to use?
No installation is required to read the guides. Clone or download the repository and navigate to the relevant directory. Some guides are HTML files that render best in a browser rather than the GitHub web interface. The repository does not include runnable code that requires a specific runtime environment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/forceinjection-ai-fundamentals)