Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
Cache-to-Cache: Direct Semantic Communication Between LLMs (2025) (news.ycombinator.com)
2.
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression (news.ycombinator.com)
3.
Foundation Model Engineering: From Theory to Production (news.ycombinator.com)
4.
LRU is harder to beat than the KV-cache papers suggest (news.ycombinator.com)
5.
Benchmarking Pocket-Scale Inference (news.ycombinator.com)
6.
Nvidia finds that simple linear math can replace costly AI model handoffs (venturebeat.com)
7.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025) (news.ycombinator.com)
8.
Predictive Speculative KV Replication for Bursty LLM Inference (news.ycombinator.com)
9.
GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz (news.ycombinator.com)
10.
Can I Buy Your KV Cache? (news.ycombinator.com)
11.
KVarN: Native vLLM backend for KV-cache quantization by Huawei (news.ycombinator.com)
12.
KVarN: Native vLLM KV-cache quantization back end by Huawei (news.ycombinator.com)
13.
Users say Gemini starts forgetting long before it’s supposed to (androidauthority.com)
14.
SEM-Guided Low-kV FIB Finishing for Leading-Edge Semiconductor Failure Analysis (spectrum.ieee.org)
15.
KV Sharing, MHC, and Compressed Attention (news.ycombinator.com)
16.
Autoregressive next token prediction and KV Cache in transformers (news.ycombinator.com)
17.
KV Cache Is Becoming the Memory Hierarchy of Inference (news.ycombinator.com)
18.
High-Fidelity KV Cache Summarization Using Entropy and Low-Rank Reconstruction (news.ycombinator.com)
19.
From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem (news.ycombinator.com)
20.
Google's TurboQuant reduces AI LLM cache memory capacity requirements by at least six times — up to 8x performance boost on Nvidia H100 GPUs, compresses KV caches to 3 bits with no accuracy loss (tomshardware.com)
21.
Nvidia says it can shrink LLM memory 20x without changing model weights (venturebeat.com)
22.
Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0 (news.ycombinator.com)
23.
Lossless LLM 3x Throughput Increase by LMCache (news.ycombinator.com)
24.
Cloudflare: Outage not caused by security incident, data is safe (bleepingcomputer.com)
Today's top topics: ai safety openai gemini google dario amodei artificial intelligence claude tilly norwood apple watch series 12 ai force
View all today's topics →