Tech News
← Home  ·  All topics

Inference Providers

2 GoKawiil briefs on this topic

Team's month-long GLM 5.3 Flash coding trial derailed by costs, outages

A development team set out to run an entire month of coding work on the GLM 5.3 Flash model, keeping within a $68 budget for the first two weeks. The second half of the month saw usage spike to 1 billion tokens across other models, driven by an expensive vibe-coded Wagtail MCP server prototype that alone cost $150 and 5kWh, plus infrastructure capacity issues that forced switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash.

Z.ai unmasks mystery model Ox Alpha as open-weight GLM-5.3-Flash

An unidentified free model called Ox Alpha appeared on OpenRouter and quickly drew massive usage from developers who couldn't determine its origin, sparking days of speculation about which lab built it. Z.ai revealed the model was actually its own GLM-5.3-Flash, running entirely on Chinese chips, and priced at 15 cents per million input tokens and 50 cents per million output tokens, with a launch discount cutting that in half through September 9.