Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
LRU is harder to beat than the KV-cache papers suggest (news.ycombinator.com)
2.
Benchmarking Pocket-Scale Inference (news.ycombinator.com)
3.
Nvidia finds that simple linear math can replace costly AI model handoffs (venturebeat.com)
4.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025) (news.ycombinator.com)
5.
GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz (news.ycombinator.com)
6.
Can I Buy Your KV Cache? (news.ycombinator.com)
7.
KVarN: Native vLLM backend for KV-cache quantization by Huawei (news.ycombinator.com)
8.
KVarN: Native vLLM KV-cache quantization back end by Huawei (news.ycombinator.com)
9.
Users say Gemini starts forgetting long before it’s supposed to (androidauthority.com)
10.
Autoregressive next token prediction and KV Cache in transformers (news.ycombinator.com)
11.
KV Cache Is Becoming the Memory Hierarchy of Inference (news.ycombinator.com)
12.
High-Fidelity KV Cache Summarization Using Entropy and Low-Rank Reconstruction (news.ycombinator.com)
13.
From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem (news.ycombinator.com)
14.
Nvidia says it can shrink LLM memory 20x without changing model weights (venturebeat.com)
Today's top topics: anthropic fast company apple openai innovation by design iphone google android ai safety samsung
View all today's topics →