Smaller, faster, safer: running Kimi and GLM at scale
(news.ycombinator.com)
1.
2.
AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
3.
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
(news.ycombinator.com)
4.
5.
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
(news.ycombinator.com)
6.
Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
(news.ycombinator.com)
9.
Kimi K3-256k
(news.ycombinator.com)
10.
11.
Running Kimi K3 on a M1 Max
(news.ycombinator.com)
12.
Running Kimi K3 on a M1 Mac
(news.ycombinator.com)
13.
What to know about Moonshot AI and its new open-weight model Kimi K3
(feeds.feedburner.com)
14.
What to know about Moonshot AI and its new open weight model Kimi K3
(feeds.feedburner.com)
15.
Kimi K3 Architecture Overview and Notes
(news.ycombinator.com)
16.
Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
(news.ycombinator.com)
17.
Using an open model feels surprisingly good
(news.ycombinator.com)
18.
19.
20.
Kimi K3 Now Available via Telnyx Inference API
(news.ycombinator.com)
21.
22.
Kimi-K3 Technical Report [pdf]
(news.ycombinator.com)
23.
Why China is giving away its best AI models
(theverge.com)
24.
25.
Kimi-K3 on HuggingFace
(news.ycombinator.com)
26.
Kimi-K3 Releases on HuggingFace 7/27
(news.ycombinator.com)
27.
Making sense of the panic over Chinese AI
(techcrunch.com)
28.
Kimi K3 built a Windows XP in browser
(news.ycombinator.com)
29.
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
(news.ycombinator.com)