Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri (news.ycombinator.com)
2.
Rotary GPU: Exploring Local Execution for Large MoE Models Under Limited VRAM (news.ycombinator.com)
3.
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request (news.ycombinator.com)
Today's top topics: apple openai google nvidia anthropic meta android authority amazon samsung microsoft
View all today's topics →