Skip to content
Tech News
← Back to articles

DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

read original get O'Reilly "Build a Large Language Model (From Scratch)" by Sebastian Raschka → more articles
Why This Matters

DeepSeek says its cheaper V4.1 Flash model beats its own flagship V4 Pro on performance, cost, speed and task completion, and it will automatically reroute Pro traffic to Flash at Flash pricing until V4.1 Pro ships. That collapses the traditional premium/budget tier split and puts further downward pressure on frontier-model API pricing, with off-peak rates as low as $0.003 per unit for cached input and $0.6 for output.

Key Takeaways
Worth a Look

O'Reilly "Build a Large Language Model (From Scratch)" by Sebastian Raschka — If headlines about V4.1 Flash beating V4 Pro on cost and speed make you curious what's actually under the hood, this book walks you through building an LLM step by step in code. It's a great way to understand tokens, inference costs and caching — the very things those input-cache-hit prices are talking about.

See O'Reilly "Build a Large Language Model (From Scratch)" by Sebastian Raschka on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

DSeek plans to officially release the V4.1 Flash model around September 10, 2026 (Beijing Time). After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time. In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!

We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.