Lab

A library of journeys. Each one is a book you read a page a day, written as interactive chapters instead of video courses. One is open.

JOURNEY · 30 DAYS1 / 30 CHAPTERS PUBLISHED

GPU Programming: Zero to Reading Production Kernels

A daily journey from “what is a warp” to reading production kernel code. One topic a day, fifteen to twenty minutes each, in plain language with real numbers from an RTX 4070. It ends at being able to open a production kernel and reason about it. Tensor cores, multi-GPU collectives, and PTX-level tuning are a deliberate Part 2.

start reading

TABLE OF CONTENTS

Foundations & mental model

DAYS 0106

Memory

DAYS 0712
  • 07Global memory & coalescing
  • 08Shared memory & bank conflicts
  • 09Registers & spilling
  • 10Caches & bandwidth limits
  • 11The roofline model
  • 12Case study: why naive matmul is slow

Writing CUDA

DAYS 1318
  • 13Your first kernel
  • 14Launch configuration in practice
  • 15Tiling with shared memory
  • 16Parallel reductions
  • 17Streams & concurrency
  • 18Reading an Nsight trace

Triton & modern kernel authoring

DAYS 1923
  • 19Why Triton exists
  • 20Triton's programming model
  • 21Writing a fused softmax in Triton
  • 22Autotuning
  • 23Kernel fusion

Reading real-world kernels

DAYS 2427
  • 24Anatomy of FlashAttention
  • 25Quantized kernels
  • 26KV-cache kernels
  • 27Reading production kernel code

Capstone

DAYS 2830
  • 28Profile and optimize
  • 29Capstone: a fused kernel from scratch
  • 30Where to go next
get in touch