δ-mem: Efficient Online Memory for Large Language Models
(news.ycombinator.com)
31.
32.
Δ-Mem: Efficient Online Memory for Large Language Models
(news.ycombinator.com)
33.
Learning the Integral of a Diffusion Model
(news.ycombinator.com)
34.
A Theory of Deep Learning
(news.ycombinator.com)
35.
A Simpler Parametrization for Modern Optimizers
(news.ycombinator.com)
36.
Why are neural networks and cryptographic ciphers so similar? (2025)
(news.ycombinator.com)
37.
38.
Softmax, can you derive the Jacobian? And should you care?
(news.ycombinator.com)
39.
There Will Be a Scientific Theory of Deep Learning
(news.ycombinator.com)
40.
Different Language Models Learn Similar Number Representations
(news.ycombinator.com)
41.
Approximating Hyperbolic Tangent
(news.ycombinator.com)
42.
43.
Types and Neural Networks
(news.ycombinator.com)
44.
4-bit floating point FP4
(news.ycombinator.com)
45.
46.
Epicycles All the Way Down (2025)
(news.ycombinator.com)
47.
Epicycles All the Way Down
(news.ycombinator.com)
48.
Study: Back-to-basics approach can match or outperform AI in language analysis
(news.ycombinator.com)
49.
Cooperative Vectors Introduction
(news.ycombinator.com)
50.
The Training Example Lie Bracket
(news.ycombinator.com)
52.
There is No Spoon. A software engineers primer for demystified ML
(news.ycombinator.com)
53.
54.
Intuitions for Tranformer Circuits
(news.ycombinator.com)
55.
AI (2014)
(news.ycombinator.com)
56.
Evolution
(feeds.nature.com)
57.
Integrated photonic neural network with on-chip backpropagation training
(feeds.nature.com)
58.
Datasets for Reconstructing Visual Perception from Brain Data
(news.ycombinator.com)
59.
Ten years of deploying to production
(news.ycombinator.com)
60.
Microgpt explained interactively
(news.ycombinator.com)