Google's TurboQuant cuts AI working memory by 6x, but it won't fix the global RAM shortage
… Published on Google Research , the tech is described as a way to shrink AI's working memory, known as the " KV cache ", by using a form of vector quantization. …