Strata lets gamers run Qwen3.8-Flash-Next 125B model on RTX 4090-class PCs
A free, open-source tool called Strata enables the 125-billion-parameter Qwen3.8-Flash-Next AI model to run locally on consumer gaming PCs with 12GB+ VRAM, rather than requiring server hardware. Testing on an RTX 5070 and RX 9070 XT showed response speeds ranging from roughly 44 to 94 tokens per second depending on model compression level, with all processing staying on the local machine.
GoKawiil's interpretation of the reporting above, not reported fact.
Running a model this large locally, without cloud servers, could make advanced AI capabilities like coding assistance and image understanding more accessible and private for individual developers and hobbyists. The reported speeds—several exceeding typical human reading pace—suggest consumer GPUs are becoming viable for workloads previously reserved for data-center hardware, though real-world performance will vary by card and use case.
- Strata runs the 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB VRAM, supporting both NVIDIA and AMD cards.
- Benchmarks on an RTX 5070 and RX 9070 XT show response speeds between roughly 44 and 94 tokens per second across different compression settings.
- The tool is free and open source, and processes everything locally without sending data to external servers.
NVIDIA GeForce RTX 4090 Graphics Card — This article is all about running a massive 125B parameter AI model locally, and the RTX 4090's large VRAM and horsepower make it a top choice for exactly that kind of on-device inference. It's the hardware enthusiasts reach for when they want to push local LLMs to their fastest token speeds without relying on the cloud.
See NVIDIA GeForce RTX 4090 Graphics Card on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.Source: github.com, 2026-10-04
Published there as: “Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.