DeepSeek confirmed it will officially release its V4.1 Flash model on September 10, 2026 (Beijing Time), stating internal testing shows it outperforms V4 Pro on performance, cost, speed, and task completion time. Until V4.1 Pro arrives, all Pro-model requests will be automatically redirected to V4.1 Flash and charged at the cheaper Flash pricing tier.
news.ycombinator.com
· 2026-09-09
The NSA, CISA and FBI issued a joint advisory alleging that Chinese AI companies including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI extracted billions of tokens from US frontier models such as Claude, GPT, Gemini and Grok since 2024. The agencies say this data was used to train competing systems like DeepSeek's R1 and Moonshot's Kimi K2/K3, describing the practice as 'distillation activities at an industrial scale.'
engadget.com
· 2026-09-09
The developer behind the iMessage AI assistant Olly, which has processed over 18 million messages with roughly a third routed through OpenRouter's open-source models, published a technical breakdown of problems encountered at scale. The core issue: OpenRouter can send identical model requests to any of about 20 different hosting providers, each running the same weights but with different precision, optimizations and tool parsers, producing measurably different benchmark results. For DeepSeek V4 Flash, GPQA Diamond and TAU-Bench Airline scores varied by several points across providers on the same day.
mmoustafa.com
· 2026-09-09
The NSA, CISA, and FBI issued a joint advisory alleging that Chinese AI companies, including DeepSeek and Moonshot AI, have run large-scale campaigns since 2024 to extract capabilities from leading US models like ChatGPT, Claude, Gemini, and Grok. The agencies say these firms used millions of queries routed through multiple accounts, APIs, proxies, and cloud providers to generate synthetic training data and build competing models such as DeepSeek's R1/R3 and Moonshot's Kimi-K2/K3.
slashdot.org
· 2026-09-08
A research team reports developing a pretraining method that matches DeepSeek V4 Pro Base's performance using roughly 50 times less compute, at an estimated cost of about $0.5 million on GB200 chips. Scaling the same recipe up 10x further, to roughly $4 million, reportedly surpassed all publicly available open base models on perplexity benchmarks.
magic.dev
· 2026-09-08
A newcomer called StartLux, founded by veteran programmer Chen Danyan and formerly known as Yuandian Xinghui, placed second in China's CAICT MCP specialized benchmark with its StartLux-V1.0-27B-Preview model. The 27-billion-parameter model trailed DeepSeek-V4-Pro, a 1.6-trillion-parameter system, by just 1.3 percentage points, despite running locally on consumer-grade PCs rather than in the cloud. Chen has argued that local models will eventually outcompete cloud-based ones on efficiency and market share.
chinaonchina.com
· 2026-09-03
Writer Jordan Andersen expands on a post by Disesdi Shoshana Cox questioning why, despite AI supposedly boosting developer productivity tenfold, no wave of new billion-dollar startups comparable to Airbnb, Stripe, or Dropbox has emerged. Instead, the biggest winners of the generative AI boom are AI companies themselves—OpenAI, Anthropic, and High-Flyer—while broader effects include cluttered inboxes, spammy social content, and degraded search results.
hermit-tech.com
· 2026-09-01
Jerry James benchmarked an 8-GPU NVIDIA RTX PRO 6000 Blackwell system paired with an AMD EPYC 9555 CPU, offering 768GB of combined GDDR7 memory. Rather than splitting huge models like Llama-3.1 405B or DeepSeek-R1 671B across all eight PCIe cards, which introduces heavy latency, the team found the setup performs best running independent single-GPU model instances in parallel.
gpupartner.com
· 2026-08-31
The vLLM project released version 0.28.0, combining 584 commits from 270 contributors. The update centers on deep performance work for Kimi-K3, including new decode context parallelism, fused kernels, memory-saving expert sharding, and ROCm support, alongside DeepSeek V4 improvements such as end-to-end sparse MLA and AMD Quark NVFP4 support.
github.com
· 2026-08-29
DeepSeek, the AI lab founded by Liang Wenfeng, has relied on his hedge fund High-Flyer Quant for early funding and computing power, but is now courting outside investors as its needs expand. Meanwhile, High-Flyer's affiliates have taken sizable pre-IPO stakes in Chinese hard-tech firms like memory chipmaker CXMT and robotics company Unitree, even as volatility in AI and chip stocks has hit the fund's returns.
cnbc.com
· 2026-08-28
Nvidia announced hardware optimizations for popular Chinese open-source AI models, including DeepSeek's V4 Flash and Alibaba's Qwen 3.8, as part of a broader push to support open-weight systems. Alongside the announcement, an SEC filing tied to its quarterly earnings warned that potential White House restrictions on Chinese-origin AI could hurt its business by limiting its ability to support apps built on those models.
cnbc.com
· 2026-08-27
Chinese AI startup DeepSeek, based in Hangzhou, is said to be raising fresh capital that would value the company at roughly $74 billion. The funds are intended to support research and development along with an expansion of its computing infrastructure.
wsj.com
· 2026-08-27