Courses · in production
Courses are coming.
Built like the posts: measured, runnable, no hand-waving. Each course ships with working code, real benchmarks, and exercises against actual inference stacks, not slideware.
KV Cache: The Complete Mental Model
Memory math, paging, prefix reuse, eviction, and what it all costs. By the end you can reason about any serving stack from first principles.
LLM Inference Performance Engineering
The deep end: vLLM internals, writing and profiling GPU kernels, quantization trade-offs, and the discipline of measuring serving systems honestly.
Be first in line
Course-opening announcements, early-bird pricing, and a vote on what gets built next. Nothing else.
Double opt-in. Unsubscribe anytime. No spam, ever.