Skip to content
Tech News
← Back to articles

Rust-built PSSA model outpaces transformer in speed and loss on WikiText-103 test

read original more articles
GoKawiil Brief

A developer built PSSA, a non-transformer language model coded from scratch in Rust without any ML framework, using a recurrent state-space layer, an episodic memory bank, and weights that update while the model runs. Tested against a matched-parameter standard transformer on identical data, PSSA reached a training loss of 3.98 versus the transformer's 4.43 over 12.7 million tokens, and generated text about twelve times faster on the same CPU. On a held-out 198,939-token slice neither model trained on, PSSA also scored lower loss throughout every checkpoint.

Why It Matters

GoKawiil's interpretation of the reporting above, not reported fact.

The comparison suggests that for smaller-scale language modeling tasks, alternatives to the transformer architecture built around linear-cost recurrent state and explicit memory lookup could offer meaningful efficiency gains over the standard quadratic-cost attention approach. Because the tests were run at matched parameter counts and on a single corpus, it remains unclear whether these advantages would persist at the much larger scales and diverse datasets typical of production language models. The independent, framework-free implementation could also make the architecture easier to inspect and modify than mainstream deep learning stacks.

Key Takeaways

Source: github.com, 2026-09-30

Published there as: “PSSA: A non-transformer language model written from scratch in Rust”

Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.