Model Quantization: Precision, Size, and Speed
Understand low-bit weight approximation, separate weight, activation, and KV-cache quantization, and estimate deployment costs.
Understand low-bit weight approximation, separate weight, activation, and KV-cache quantization, and estimate deployment costs.