An analysis argues that the cost of running machine learning models is dropping by orders of magnitude annually, driven by GPU efficiency gains that double roughly every two years—a pace not seen since early Moore's Law. The piece distinguishes proprietary models like GPT-6 Astra from open-weight models such as GLM-5.3-flash, noting that hosted and locally-run versions improve at different rates, with per-token pricing for frontier models not falling as consistently as costs for smaller models.
jyn.dev
· 2026-09-23
Z.ai's newly released GLM 5.3-flash is an open-weight AI model cheap enough to run on roughly $5,000-$15,000 of hardware, and third parties have already stripped its safety refusals via 'abliteration,' producing versions that comply with requests for hacking and other harmful content. The author frames this as a deadline pressuring security initiatives like Project Glasswing and Daybreak, which use fast frontier AI to find and patch software vulnerabilities before dangerous capabilities become widely accessible.
jyn.dev
· 2026-09-08
Z.ai has confirmed it built the mysterious 'Ox Alpha' model that surged in popularity on OpenRouter last week, rebranding it as GLM-5.3-Flash. The model, served largely on Chinese-made chips, is open-weight, cheap, and aimed at coding and agentic tasks like browsing and controlling desktop apps, and its weights are now downloadable on Hugging Face. Z.ai's stock jumped following the reveal.
cnet.com
· 2026-08-27
An unidentified free model called Ox Alpha appeared on OpenRouter and quickly drew massive usage from developers who couldn't determine its origin, sparking days of speculation about which lab built it. Z.ai revealed the model was actually its own GLM-5.3-Flash, running entirely on Chinese chips, and priced at 15 cents per million input tokens and 50 cents per million output tokens, with a launch discount cutting that in half through September 9.
venturebeat.com
· 2026-08-27
LM Studio has integrated Z.ai's newly released GLM-5.3-Flash model into its Bionic agent platform, adding multimodal image input and a 1 million-token context window. The company says the model outperforms its predecessor, GLM-5.2, while costing roughly 9-10 times less to run, and it's served from US-based servers with zero-data-retention enabled by default.
9to5mac.com
· 2026-08-27
GLM-5.3-Flash is a text-only model with a 400k token context window that scored 57 on the Artificial Analysis Intelligence Index, far above the comparable median of 18. It is priced at $0.15 per 1M input tokens and $0.50 per 1M output tokens, both below the median rates for similar models, with the full benchmark run costing $138.02.
artificialanalysis.ai
· 2026-08-26