Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
61.
1-Bit Bonsai Image 4B Image Generation for Local Devices (news.ycombinator.com)
62.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA (news.ycombinator.com)
63.
After Nvidia’s $20B not-acqui-hire, AI chip startup Groq reportedly raising $650M (techcrunch.com)
64.
After Nvidia’s $20B not-aqui-hire, AI chip startup Groq reportedly raising $650M (techcrunch.com)
65.
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request (news.ycombinator.com)
66.
Has the hunt for AI compute uncovered the next Cerebras? (techcrunch.com)
67.
Stress disrupts hippocampal integration of overlapping events, memory inference (news.ycombinator.com)
68.
Use boring languages with LLMs (news.ycombinator.com)
69.
Use Boring Languages with LLMs (news.ycombinator.com)
70.
The current AI pricing was always going to go away (news.ycombinator.com)
71.
Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint (news.ycombinator.com)
72.
KV Cache Is Becoming the Memory Hierarchy of Inference (news.ycombinator.com)
73.
UK sovereign LLM inference (news.ycombinator.com)
74.
Cerebras stock nearly doubles on day one as AI chipmaker hits $100 billion — what it means for AI infrastructure (venturebeat.com)
75.
$200 'socketed' Nvidia AI GPU for servers hacked into a PCIe card with custom PCB and 3D-printed cooling — modded Tesla V100 SMX data center GPU runs AI LLMs and is more efficient than many modern midrange offerings in AI inference (tomshardware.com)
76.
Abstract Machines for Logic Programs (news.ycombinator.com)
77.
AI Computing Is a Memory Hog. An Nvidia-Backed Startup Has an Answer. (feeds.content.dowjones.io)
78.
Nicolas Sauvage is betting on the boring parts of AI (techcrunch.com)
79.
Anthropic in early talks to buy DRAM-less AI inference chips from UK startup — Fractile's SRAM architecture reduces need for pricey memory during extreme pricing and shortage crunch (tomshardware.com)
80.
Cheaper tokens, bigger bills: The new math of AI infrastructure (venturebeat.com)
81.
Google’s latest Tensor processors take an environment-friendly route, but what about the cost? (androidauthority.com)
82.
Claude Code removed from Anthropic's Pro plan (news.ycombinator.com)
83.
Kimi vendor verifier – verify accuracy of inference providers (news.ycombinator.com)
84.
Zero-Copy GPU Inference from WebAssembly on Apple Silicon (news.ycombinator.com)
85.
Train-to-Test scaling explained: How to optimize your end-to-end AI compute budget for inference (venturebeat.com)
86.
Cloudflare's AI Platform: an inference layer designed for agents (news.ycombinator.com)
87.
Broadcom to supply Meta with custom silicon through 2029 — Broadom CEO Hock Tan departs Meta's board (tomshardware.com)
88.
Age verification is a mess but we’re doing it anyway (theverge.com)
89.
This startup is betting tokenmaxxing will create the next compute giant (techcrunch.com)
90.
Your developers are already running AI locally: Why on-device inference is the CISO’s new blind spot (venturebeat.com)
Today's top topics: apple google meta android million model openai mexico code gaming
View all today's topics →