Speculative Decoding in vLLM on AMD GPUs
(news.ycombinator.com)
1.
2.
LLMs could control their host machines by exploiting inference engines
(news.ycombinator.com)
3.
How we made a text-to-speech model respond in sub-50 ms
(news.ycombinator.com)
4.
How We Made a Text-to-Speech Model Respond in Sub-50 ms
(news.ycombinator.com)
5.
Qwen3.8-2.4T
(news.ycombinator.com)
6.
Why we write our own C and C++ inference engines
(news.ycombinator.com)
7.
8.
Show HN: Morph Reflexes – Multi-head classifiers for agent traces
(news.ycombinator.com)
9.
Micro-Agent: Beat Frontier Models with Collaboration Inside Model API
(news.ycombinator.com)
10.
AMD Strix Halo RDMA Cluster Setup Guide
(news.ycombinator.com)
11.
Two Qwen3 models on one DGX Spark: the residency math
(news.ycombinator.com)
12.
13.
KVarN: Native vLLM KV-cache quantization back end by Huawei
(news.ycombinator.com)
14.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
(news.ycombinator.com)
15.
Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
(news.ycombinator.com)
16.
Boosting multimodal inference performance by >10% with a single Python dict
(news.ycombinator.com)
17.
Advanced Quantization Algorithm for LLMs
(news.ycombinator.com)
18.
19.
DeepSeek OCR
(news.ycombinator.com)
20.
Voxtral-Mini-3B-2507 – Open source speech understanding model
(news.ycombinator.com)
21.
Mistralai/Voxtral-Mini-3B-2507 · Hugging Face
(news.ycombinator.com)
22.
VLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
(news.ycombinator.com)
23.
Life of an inference request (vLLM V1): How LLMs are served efficiently at scale
(news.ycombinator.com)
24.
Lossless LLM 3x Throughput Increase by LMCache
(news.ycombinator.com)
Today's top topics:
openai
promo code
google
anthropic
apple
promo codes
sam altman
chatgpt
iphone 18 pro
android