Deep Dives
TurboQuant: Google's 6x KV Cache Compression Hits 3-Bit With Zero Loss
Google's TurboQuant compresses KV cache 6x at 3 bits with zero loss, speeding up attention.
Reporting, analysis and practical guides connected to llm-inference.
Google's TurboQuant compresses KV cache 6x at 3 bits with zero loss, speeding up attention.