DeepSeek released V4.1 Flash, a model initially mistaken for a minor update but revealed via its technical report to be a significant architectural overhaul, effectively a V5-class release. It achieves near 420 tokens/second throughput while compressing KV cache by 4x through techniques including cross-layer compression, sparse attention indexing optimizations, and FP4 precision, alongside a YOCO-inspired prefill design that only activates 8B parameters during prefill versus 16B during decode across its 40 layers.
zartbot.github.io
· 2026-09-17
DeepSeek V4.1 Flash achieved code execution on all 11 vulnerable systems in an AI hacking benchmark while leaving four patched systems untouched, at a total cost of just $4.65 for accepted runs. A manual review found the model discovered five novel attack paths beyond the six expected solutions, including a faster exploit against Grafana that bypassed the intended vulnerability entirely.
enclave.ai
· 2026-09-16
DeepSeek has released V4.1-Flash on its API, replacing the earlier V4-Flash and V4-Flash-Vision-Exp models. The new model adds native multimodal capabilities and can be accessed by setting the model parameter to deepseek-flash.
twitter.com
· 2026-09-10
DeepSeek confirmed it will officially release its V4.1 Flash model on September 10, 2026 (Beijing Time), stating internal testing shows it outperforms V4 Pro on performance, cost, speed, and task completion time. Until V4.1 Pro arrives, all Pro-model requests will be automatically redirected to V4.1 Flash and charged at the cheaper Flash pricing tier.
news.ycombinator.com
· 2026-09-09