Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%
Sentiment Mix
Geography
Expert Signals
arnav__1
author • 1 mention
Hacker News
source • 1 mention
AI-Generated Claims
Generated from linked receipts; click sources for full context.
Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%.
Supported by 1 story
We built OpenLake because KV caches are outgrowing GPU memory.A single 256K token conversation on Gemma 4 31B produces approximately 43GB of KV state, more than half the memory of an 80GB H100.
Supported by 1 story
The problem becomes even harder across a cluster: a prefix cached on one GPU host is unavailable when the next request lands on a different GPU, forcing the new GPU to repeat work the fleet has already completed.Once the KV cache is offloaded, network bandwidth becomes a major constraint on read latency.
Supported by 1 story
In our tests, this achieved:- 1.72× lossless KV compression.
Supported by 1 story
- Approximately 600GB/s decompression throughput on an H100 - 80GB/s of...
Supported by 1 story
Related Events
Show HN: HART OS – an open-source AI OS built so frontier AI needs no datacenter
Open Source • 7/27/2026
Show HN: Distill and serve small models with frontier quality for half the cost
Uncategorized • 7/27/2026
Show HN: Managing on-premise servers without Kubernetes
Uncategorized • 7/26/2026
Show HN: Reproducibility Benchmark a Risk Quantitative Model
Research • 7/26/2026
AMD to invest up to $5bn in Anthropic in massive AI server deal - Memeburn
LLMs • 7/27/2026