Skip to content
Tech News
← Back to articles

Strata lets gamers run Qwen3.8-Flash-Next 125B model on RTX 4090-class PCs

read original get NVIDIA GeForce RTX 4090 Graphics Card → more articles
GoKawiil Brief

A free, open-source tool called Strata enables the 125-billion-parameter Qwen3.8-Flash-Next AI model to run locally on consumer gaming PCs with 12GB+ VRAM, rather than requiring server hardware. Testing on an RTX 5070 and RX 9070 XT showed response speeds ranging from roughly 44 to 94 tokens per second depending on model compression level, with all processing staying on the local machine.

Why It Matters

GoKawiil's interpretation of the reporting above, not reported fact.

Running a model this large locally, without cloud servers, could make advanced AI capabilities like coding assistance and image understanding more accessible and private for individual developers and hobbyists. The reported speeds—several exceeding typical human reading pace—suggest consumer GPUs are becoming viable for workloads previously reserved for data-center hardware, though real-world performance will vary by card and use case.

Key Takeaways
Worth a Look

NVIDIA GeForce RTX 4090 Graphics Card — This article is all about running a massive 125B parameter AI model locally, and the RTX 4090's large VRAM and horsepower make it a top choice for exactly that kind of on-device inference. It's the hardware enthusiasts reach for when they want to push local LLMs to their fastest token speeds without relying on the cloud.

See NVIDIA GeForce RTX 4090 Graphics Card on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Source: github.com, 2026-10-04

Published there as: “Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s”

Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.