Show HN: TurboGPT trains a 22KB byte-level GPT in 13 seconds via CUDA C++
A developer released TurboGPT, an open-source (MIT-licensed) CUDA C++ implementation for training tiny byte-level GPT models. The project's Windows build instructions and CLI usage are documented, and a sample run on a dataset called hn1g reportedly reached 2.52435 bits-per-byte after 1.5 billion training tokens.
GoKawiil's interpretation of the reporting above, not reported fact.
The project's emphasis on training a very small model quickly in a low-level language suggests an interest in efficient, minimal-dependency machine learning tooling rather than large-scale production use. Its use of raw CUDA C++ instead of frameworks like PyTorch could appeal to developers wanting fine-grained control over GPU training performance, though the practical utility of a 22KB model remains limited to experimentation or educational purposes.
- TurboGPT is a minimal, MIT-licensed CUDA C++ tool for training byte-level GPT models.
- A sample model trained on 1.5 billion tokens achieved 2.52435 BPB in about 13 seconds.
- Checkpoints and logs are stored in TensorBoard-compatible formats for resuming and analysis.
NVIDIA GeForce RTX 4070 GPU — Since turboGPT relies on CUDA compute capability for training speed, a modern NVIDIA GPU like the RTX 4070 gives you the horsepower to run and experiment with these tiny transformer training jobs quickly. It's a solid choice for hobbyist ML tinkering, CUDA development, and local model training experiments like this one.
See NVIDIA GeForce RTX 4070 GPU on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.Source: github.com, 2026-09-29
Published there as: “Show HN: TurboGPT: train 22KiB transformer in 13s”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.