Part V — Does It Actually Work?
Section 13

KV Cache Experiments — The Headlines

KV Cache Experiments

If you compress a real LLM’s KV cache with TurboQuant, does the model still give correct answers? The paper tests on Llama-3.1-8B-Instruct – a production-grade open model.


Result 1: Needle-in-a-Haystack — Perfect Retrieval at 4x Compression

Hide one sentence in up to 104K tokens and ask the model to find it. If KV cache compression damages attention, the model fails.

MethodRecall
Full Precision0.997
TurboQuant0.997
PolarQuant0.995
KIVI0.981
PyramidKV0.895
SnapKV0.858

TurboQuant scores identically to the uncompressed model. Not “close to” – the exact same score: 0.997. The full-precision and TurboQuant heatmaps are visually indistinguishable – perfect retrieval across all context lengths and needle positions.

At 4x compression, TurboQuant doesn’t degrade long-context retrieval at all.

(Competitor note: this table includes KIVI and PolarQuant, which are the most direct competing KV cache compression methods. Other KV-only methods from the 2024 literature — KVQuant (ICML 2024, per-channel calibration at 3-bit), Gear (group-wise quantization with residual approximation), and H2O / ScissorHands (attention-score-based token eviction) — are not included in the paper’s evaluation. The omissions are likely due to evaluation scope; practitioners choosing between methods should run the Needle and LongBench evaluations against KVQuant specifically, as it is the closest architectural competitor at sub-4-bit KV compression.)


Result 2: LongBench — Quality Neutral at 3.5 Bits

LongBench covers six categories: single-doc QA, multi-doc QA, summarization, few-shot learning, synthetic tasks, and code completion.

MethodKV bitsAverage
Full Cache1650.06
TurboQuant3.550.06
KIVI550.16
PolarQuant3.949.78
TurboQuant2.549.44
KIVI348.50

Three things jump out:

1. TurboQuant at 3.5 bits = Full precision. Score of 50.06 in both cases. A 4.6x reduction in KV cache memory with zero measurable quality impact.

2. TurboQuant at 2.5 bits beats KIVI at 3 bits. Even at 6.4x compression, TurboQuant (49.44) outperforms KIVI at a higher bit budget (48.50).

3. TurboQuant quantizes during generation. Unlike KIVI and PolarQuant which leave newly generated tokens unquantized, TurboQuant compresses in real time. A stricter test – and it still matches full precision.


Non-Integer Bit-Widths: Outlier Channels

The 2.5 and 3.5-bit configurations use a practical trick: not all channels are equal. Some “outlier” channels have much larger magnitudes. The solution is to split channels into two groups at different precision:

2.5-bit: 32 outlier channels × 3 bits + 96 regular channels × 2 bits
3.5-bit: similar split with higher allocation to outliers

TurboQuant handles this naturally by running two independent instances with different bit-widths.


Community Benchmarks: turbo3 and turbo4 vs Standard Quantization

After the March 2026 release, the turboquant+ community ran head-to-head perplexity comparisons against llama.cpp’s standard quantization formats on Llama-3.1-8B across a 32K-token text corpus:

MethodBits/valPPL (↓ better)vs q4_0
f16 (baseline)166.12—
q5_056.21—
q4_046.34baseline
turbo4 (K=4b, V=2b)~3.86.31essentially identical
turbo3 (K=3b, V=2b)~2.86.41+1.1%
q3_K_M3.356.58+3.8%

turbo4 matches q4_0 on perplexity while using fewer bits. The 0.03 PPL difference (6.31 vs 6.34) is within measurement noise for a single 32K-token run — the meaningful result is that turbo4 achieves this at ~3.8 bits versus q4_0’s 4 bits, a 5% further compression at equivalent quality.


Real-World Landmark: 104B Parameters at 128K Context on a MacBook

The most striking community benchmark: a 104B parameter model (Llama 3.3 70B + 34B expert layers) running at 128K context length on a single MacBook Pro M5 Max with 128 GB unified memory.

Hardware:    MacBook Pro M5 Max, 128 GB RAM
Model:       104B parameters (Llama 3.3 70B + 34B expert layers)
Context:     128,000 tokens
KV config:   turbo3 (K=3b, V=2b, boundary layers at 4b)

KV cache:    ~16 GB   (vs ~128 GB at full FP16)
Throughput:  ~4 tok/s
Quality:     No measurable degradation on HumanEval and MMLU subsets

Running a 104B model at 128K context on consumer hardware was not feasible before TurboQuant. The 8× KV cache reduction (from GQA + TurboQuant combined) is what makes it fit in 128 GB.


Confirmed on a Second Model

On Ministral-7B-Instruct, a different architecture: even at 2.5 bits, the quality drop is just 0.27 points – marginal and likely within noise. This confirms TurboQuant is model-agnostic.


The Money Slide

TurboQuant at 3.5 bits per coordinate:

  • KV cache compressed by 4.6×
  • Zero quality loss on Needle-in-a-Haystack
  • Zero quality loss on LongBench (6 task types)
  • Works during streaming generation
  • Model-agnostic (tested on Llama, Ministral, Qwen, Phi, DeepSeek)
  • turbo4 matches q4_0 at lower bit-width
  • 104B model at 128K context on consumer hardware