
TurboQuant: How a Simple Spin Saves Gigabytes of GPU Memory
Start With a Restaurant Before we talk about AI, let me tell you about a busy...
Apr 8, 20266 min read0 reactions0 comments
Tag archive

Start With a Restaurant Before we talk about AI, let me tell you about a busy...

Two weeks after Google published their TurboQuant paper at ICLR 2026, five independent implementations exist -- including one running a 104B model on a MacBook. What the community built, what works today, and what it means.
The pervasive adoption of large language models (LLMs) and other deep neural networks has ushered in...
Preamble: What vector search is all about If you've spent any time near an LLM in the...

Hi everyone. My name is Jianyang Gao. I am currently a postdoctoral researcher at ETH Zurich, and I...