Attention Variants and KV-Cache Compression
How MHA, GQA, MLA, cache quantization, and token-selection methods trade memory for implementation and quality risk.
How MHA, GQA, MLA, cache quantization, and token-selection methods trade memory for implementation and quality risk.
Work through probability correction and timing examples, compare chains, trees, and self-speculation, and connect KV-state correctness to real speed gains.