LAB / GPU PROGRAMMING

DAY 06 / 30 · FOUNDATIONS & MENTAL MODEL · ~12 MIN READ

Occupancy

Registers and shared memory per block decide how many warps actually fit on an SM at once.

01 · CALLING BACK TO DAY 01

Latency hiding is a supply problem

Day 01 ended on a claim: the GPU hides memory latency by switching to other ready threads the instant one stalls. That only works if there are other ready threads — the chip needs warps queued on every SM so the scheduler never runs dry. Occupancy is the measure of that supply: how many of an SM's warp slots you've actually filled, out of 48.

What decides it is not luck. Three budgets are spent when your block lands on an SM: threads (each block brings threads per block of them), registers (threads × registers per thread, carved from the SM's 64K register file), and shared memory (reserved per block from the SM's pool). Whichever budget runs out first caps how many blocks fit, and that caps your warps.

02 · THE CALCULATOR

Shape a block, watch the warps

Move the three sliders and watch which constraint takes over. The occupancy bar is the supply of latency-hiding you're giving the scheduler.

OCCUPANCY · SIMPLIFIED ADA-CLASS SM (RTX 4070)

BLOCK BUDGETS · WHAT LIMITS YOU

threads / SM1536 / 1536
registers / SM49152 / 65536
block slots / SM6 / 24

BLOCK SIZE IS THE LIMITER · 6 BLOCKS RESIDENT

OCCUPANCY · RESIDENT WARPS

100%48 / 48 warps

plenty of warps in flight. the scheduler always has somewhere to hide latency.

simplified Ada-class model · real chips add carve-outs between L1 and shared, and compiler registers get spilled before they run out

03 · THE TRAP

High occupancy is not the goal

It's tempting to read the bar as a score. Resist that. Occupancy is potential — how much latency the scheduler could hide. Whether that potential converts depends on what the warps do: a kernel doing pure arithmetic benefits from every extra warp, while a kernel already saturating memory gains nothing past a point.

The subtler trap is chasing occupancy with waste. Shrinking your registers per thread to cram more warps in forces the compiler to spill registers to local memory (day 09's villain), and suddenly every "extra" warp is stalling on memory traffic you created. The right question isn't "how high can the bar go" but "does the scheduler ever run dry" — and you answer it with a profiler, on day 18.

CHEATSHEET · DAY 06
  • 01Occupancy = resident warps / 48. It measures the scheduler's supply of latency-hiding. Day 01's argument, quantified.
  • 02Three budgets cap it. Threads per SM, the 64K register file, and the shared memory pool. The tightest one wins.
  • 03Registers are the usual villain. Every register per thread multiplies across the whole block. Spilling to buy warps backfires.
  • 04100% is not the finish line. Occupancy is potential. Profiling (day 18) tells you whether the potential was ever the problem.

Next: day 07 leaves the SM and follows a warp's memory request out to DRAM. Why access pattern beats access count, and the single trick that keeps you on the fast path.

get in touch