Attention Variants and KV-Cache Compression
How MHA, GQA, MLA, cache quantization, and token-selection methods trade memory for implementation and quality risk.
How MHA, GQA, MLA, cache quantization, and token-selection methods trade memory for implementation and quality risk.