llm inference · measured

Deep, measured engineering content on how LLM inference actually works.

vLLM internals. KV cache mechanics. GPU kernels. Profiling. Quantization. Written for senior engineers who run models in production. Every claim benchmarked, every benchmark runnable.

// first entries are in the lab. the list hears about them first

First course in production: KV Cache: The Complete Mental Model · get notified →