Lab
A library of journeys. Each one is a book you read a page a day, written as interactive chapters instead of video courses. One is open.
JOURNEY · 30 DAYS1 / 30 CHAPTERS PUBLISHED
GPU Programming: Zero to Reading Production Kernels
A daily journey from “what is a warp” to reading production kernel code. One topic a day, fifteen to twenty minutes each, in plain language with real numbers from an RTX 4070. It ends at being able to open a production kernel and reason about it. Tensor cores, multi-GPU collectives, and PTX-level tuning are a deliberate Part 2.
start reading
TABLE OF CONTENTS
Foundations & mental model
DAYS 01–06- 01Why GPUs are shaped this way
- 02Anatomy of a GPU
- 03The thread hierarchy
- 04Warps and SIMT
- 05Warp divergence
- 06Occupancy
Memory
DAYS 07–12- 07Global memory & coalescing
- 08Shared memory & bank conflicts
- 09Registers & spilling
- 10Caches & bandwidth limits
- 11The roofline model
- 12Case study: why naive matmul is slow
Writing CUDA
DAYS 13–18- 13Your first kernel
- 14Launch configuration in practice
- 15Tiling with shared memory
- 16Parallel reductions
- 17Streams & concurrency
- 18Reading an Nsight trace
Triton & modern kernel authoring
DAYS 19–23- 19Why Triton exists
- 20Triton's programming model
- 21Writing a fused softmax in Triton
- 22Autotuning
- 23Kernel fusion
Reading real-world kernels
DAYS 24–27- 24Anatomy of FlashAttention
- 25Quantized kernels
- 26KV-cache kernels
- 27Reading production kernel code
Capstone
DAYS 28–30- 28Profile and optimize
- 29Capstone: a fused kernel from scratch
- 30Where to go next