1.
2.
Show HN: Morph Reflexes – Multi-head classifiers for agent traces
(news.ycombinator.com)
3.
Micro-Agent: Beat Frontier Models with Collaboration Inside Model API
(news.ycombinator.com)
4.
AMD Strix Halo RDMA Cluster Setup Guide
(news.ycombinator.com)
5.
Two Qwen3 models on one DGX Spark: the residency math
(news.ycombinator.com)
6.
7.
KVarN: Native vLLM KV-cache quantization back end by Huawei
(news.ycombinator.com)
8.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
(news.ycombinator.com)
9.
Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
(news.ycombinator.com)
10.
Boosting multimodal inference performance by >10% with a single Python dict
(news.ycombinator.com)
11.
Advanced Quantization Algorithm for LLMs
(news.ycombinator.com)
12.
13.
DeepSeek OCR
(news.ycombinator.com)
14.
Voxtral-Mini-3B-2507 – Open source speech understanding model
(news.ycombinator.com)
15.
Mistralai/Voxtral-Mini-3B-2507 · Hugging Face
(news.ycombinator.com)
16.
VLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
(news.ycombinator.com)
17.
Life of an inference request (vLLM V1): How LLMs are served efficiently at scale
(news.ycombinator.com)
18.
Lossless LLM 3x Throughput Increase by LMCache
(news.ycombinator.com)