Strata lets gamers run Qwen3.8-Flash-Next 125B model on RTX 4090-class PCs
A free, open-source tool called Strata enables the 125-billion-parameter Qwen3.8-Flash-Next AI model to run locally on consumer gaming PCs with 12GB+ VRAM, rather than requiring server hardware. Testing on an RTX 5070 and RX 9070 XT showed response speeds ranging from roughly 44 to 94 tokens per second depending on model compression level, with all processing staying on the local machine.