Who’s afraid of the big, bad GPU?
(theverge.com)
1.
2.
1-Bit LLM in the Browser
(news.ycombinator.com)
3.
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
(news.ycombinator.com)
4.
Matrix Orthogonalization Improves Memory in Recurrent Models
(news.ycombinator.com)
5.
Can Clothes Make You Invisible to Facial Recognition?
(darkreading.com)
6.
The Lab Mistake That Might Revolutionize Computing
(spectrum.ieee.org)
7.
Puzzling Success of Overparameterization: Lottery Tickets or Escape Dimensions?
(news.ycombinator.com)
8.
11.
"They're made out of weights"
(news.ycombinator.com)
12.
13.
SaySynth: A Brief History of Speaking Machines
(news.ycombinator.com)
14.
It Takes Two Neurons to Ride a Bicycle
(news.ycombinator.com)
15.
A sleep-like consolidation mechanism for LLMs
(news.ycombinator.com)
16.
Language Models Need Sleep
(news.ycombinator.com)
17.
AI with Model-Based Design: Virtual Sensor Modeling
(spectrum.ieee.org)
18.
Paul Schrader Says His AI Girlfriend Dumped Him
(futurism.com)
19.
Self-Distillation Enables Continual Learning [pdf]
(news.ycombinator.com)
20.
δ-mem: Efficient Online Memory for Large Language Models
(news.ycombinator.com)
21.
Δ-Mem: Efficient Online Memory for Large Language Models
(news.ycombinator.com)
22.
Learning the Integral of a Diffusion Model
(news.ycombinator.com)
23.
A Theory of Deep Learning
(news.ycombinator.com)
24.
A Simpler Parametrization for Modern Optimizers
(news.ycombinator.com)
25.
Why are neural networks and cryptographic ciphers so similar? (2025)
(news.ycombinator.com)
26.
27.
Softmax, can you derive the Jacobian? And should you care?
(news.ycombinator.com)
28.
There Will Be a Scientific Theory of Deep Learning
(news.ycombinator.com)
29.
Different Language Models Learn Similar Number Representations
(news.ycombinator.com)
30.
Approximating Hyperbolic Tangent
(news.ycombinator.com)