Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
Kimi K3's full weights are here, but they're 'open' with a caveat: What enterprises should know (venturebeat.com)
2.
Show HN: Morph Reflexes – Multi-head classifiers for agent traces (news.ycombinator.com)
3.
Micro-Agent: Beat Frontier Models with Collaboration Inside Model API (news.ycombinator.com)
4.
AMD Strix Halo RDMA Cluster Setup Guide (news.ycombinator.com)
5.
Two Qwen3 models on one DGX Spark: the residency math (news.ycombinator.com)
6.
Kimi K2.7-Code cuts thinking tokens 30% — but practitioners say the benchmarks don't check out (venturebeat.com)
7.
KVarN: Native vLLM KV-cache quantization back end by Huawei (news.ycombinator.com)
8.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA (news.ycombinator.com)
9.
Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team (news.ycombinator.com)
10.
Boosting multimodal inference performance by >10% with a single Python dict (news.ycombinator.com)
11.
Advanced Quantization Algorithm for LLMs (news.ycombinator.com)
12.
The team behind continuous batching says your idle GPUs should be running inference, not sitting dark (venturebeat.com)
13.
DeepSeek OCR (news.ycombinator.com)
14.
Voxtral-Mini-3B-2507 – Open source speech understanding model (news.ycombinator.com)
15.
Mistralai/Voxtral-Mini-3B-2507 · Hugging Face (news.ycombinator.com)
16.
VLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (news.ycombinator.com)
17.
Life of an inference request (vLLM V1): How LLMs are served efficiently at scale (news.ycombinator.com)
18.
Lossless LLM 3x Throughput Increase by LMCache (news.ycombinator.com)
Today's top topics: openai coldcard
View all today's topics →