llm inference · measured
Deep, measured engineering content on how LLM inference actually works.
vLLM internals. KV cache mechanics. GPU kernels. Profiling. Quantization. Written for senior engineers who run models in production. Every claim benchmarked, every benchmark runnable.
Latest from the lab
Read it before it hits the front page
One deep post at a time. No fluff, no reposts of the docs. Just measurements you can rerun and mental models that survive contact with production.
Double opt-in. Unsubscribe anytime. No spam, ever.