AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
1.
3.
Inflect-Micro-v2: complete voice in 9.36M parameters
(news.ycombinator.com)
4.
5.
7.
8.
9.
10.
Nvidia DGX Spark as a daily driver
(news.ycombinator.com)
11.
Alternative(s) to run CUDA on non-Nvidia hardware
(news.ycombinator.com)
12.
13.
Reduce GVisor Cold Starts with GPU Snapshotting
(news.ycombinator.com)
14.
Zluda 6 release (run unmodified CUDA applications on non-Nvidia GPUs)
(news.ycombinator.com)
15.
What happens when you run a CUDA kernel?
(news.ycombinator.com)
16.
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(news.ycombinator.com)
17.
I need your clothes, your boots, and your motorcycle
(news.ycombinator.com)
18.
Show HN: cuTile Rust: Safe, data-race-free GPU kernels in Rust
(news.ycombinator.com)
19.
What about OpenCL and CUDA C++ alternatives?
(news.ycombinator.com)
20.
Tiny hackable CUDA language model implementation
(news.ycombinator.com)
21.
Use your Nvidia GPU's VRAM as swap space on Linux
(news.ycombinator.com)
22.
23.
AI Agent Guidelines for CS336 at Stanford
(news.ycombinator.com)
24.
25.
I put a datacenter GPU in my gaming PC
(news.ycombinator.com)
26.
I Put a Datacenter GPU in My Gaming PC for £200
(news.ycombinator.com)
27.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
(news.ycombinator.com)
28.
Show HN: KVBoost – chunk-level KV cache reuse for HuggingFace, 5–48x faster TTFT
(news.ycombinator.com)
29.
Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint
(news.ycombinator.com)
30.
CUDA Books
(news.ycombinator.com)