Cerebras has added the 27-billion-parameter Qwen 3.8 model to its public API endpoints, offering inference speeds of roughly 1500 tokens per second. The model supports 64k context on free tier and 128k on paid tier, and is available under Cerebras's free trial and pay-as-you-go pricing, subject to rate limits.
inference-docs.cerebras.ai
· 2026-09-03
At Hot Chips 2026, Cerebras detailed its next two wafer-scale accelerator generations, including a new Nexus rack architecture for the CS-4 system that triples rack-scale performance using three WS-3T engines. The company also confirmed that its future CS-6 system will introduce 3D-stacked DRAM directly on top of its wafer-scale logic and SRAM, marking a first for its chip design.
tomshardware.com
· 2026-08-27