Laya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second)
(news.ycombinator.com)
1.
2.
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
(news.ycombinator.com)
3.
4.
OpenJev
(news.ycombinator.com)
5.
How GLM built its own inference infrastructure
(news.ycombinator.com)
6.
GLM Built Its Own Inference Infrastructure
(news.ycombinator.com)
7.
Breaking the 1.58-bit Barrier for Ternary LLMs
(news.ycombinator.com)
8.
The Inference Hardware Revolution of 2026
(news.ycombinator.com)
9.
The AI Inference Revolution Is Here
(spectrum.ieee.org)
10.
We do modern frequentist statistics: Using fake-data simulation
(news.ycombinator.com)
11.
D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
(news.ycombinator.com)
12.
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
(news.ycombinator.com)
13.
Mercury 2.5
(news.ycombinator.com)
14.
15.
17.
Gimlet's Series B
(news.ycombinator.com)
18.
GPT-6 Astra on OpenRouter
(news.ycombinator.com)
19.
Architecting memory and storage in the AI era
(technologyreview.com)
20.
The efficient frontier of LLM inference
(news.ycombinator.com)
21.
Benchmarking Pocket-Scale Inference
(news.ycombinator.com)
22.
GLM-5.3-Flash will likely handle 45% of your AI workloads
(venturebeat.com)
23.
OpenAI Jalapeño: Better than Nvidia Blackwell
(news.ycombinator.com)
24.
25.
LLMs could control their host machines by exploiting inference engines
(news.ycombinator.com)
26.
OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
(news.ycombinator.com)
27.
Agent Is Not the Model
(news.ycombinator.com)
28.
I built a low-latency AI companion that plays Skyrim with me
(news.ycombinator.com)
29.
Why your local LLM feels dumber than it is
(news.ycombinator.com)
30.
DFlash 2: Keep Drafting Parallel
(news.ycombinator.com)
Today's top topics:
openai
scott bessent
apple
gemini
meta
virtual wall
border patrol
anthropic
pixel 11
trump-xi summit