Tech News
← Home  ·  All topics

Deepseek

36 GoKawiil briefs on this topic

DeepSeek to launch V4.1 Flash model, undercutting V4 Pro on cost and speed

DeepSeek confirmed it will officially release its V4.1 Flash model on September 10, 2026 (Beijing Time), stating internal testing shows it outperforms V4 Pro on performance, cost, speed, and task completion time. Until V4.1 Pro arrives, all Pro-model requests will be automatically redirected to V4.1 Flash and charged at the cheaper Flash pricing tier.

US agencies accuse DeepSeek, Alibaba and other Chinese AI firms of mass model distillation

The NSA, CISA and FBI issued a joint advisory alleging that Chinese AI companies including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI extracted billions of tokens from US frontier models such as Claude, GPT, Gemini and Grok since 2024. The agencies say this data was used to train competing systems like DeepSeek's R1 and Moonshot's Kimi K2/K3, describing the practice as 'distillation activities at an industrial scale.'

Olly's builder details how OpenRouter provider swaps cause hidden model quality drift

The developer behind the iMessage AI assistant Olly, which has processed over 18 million messages with roughly a third routed through OpenRouter's open-source models, published a technical breakdown of problems encountered at scale. The core issue: OpenRouter can send identical model requests to any of about 20 different hosting providers, each running the same weights but with different precision, optimizations and tool parsers, producing measurably different benchmark results. For DeepSeek V4 Flash, GPQA Diamond and TAU-Bench Airline scores varied by several points across providers on the same day.

US agencies say Chinese AI firms systematically distilled ChatGPT, Claude, Gemini models

The NSA, CISA, and FBI issued a joint advisory alleging that Chinese AI companies, including DeepSeek and Moonshot AI, have run large-scale campaigns since 2024 to extract capabilities from leading US models like ChatGPT, Claude, Gemini, and Grok. The agencies say these firms used millions of queries routed through multiple accounts, APIs, proxies, and cloud providers to generate synthetic training data and build competing models such as DeepSeek's R1/R3 and Moonshot's Kimi-K2/K3.

Startup claims new pretraining recipe cuts compute 10x vs open-weight models

A research team reports developing a pretraining method that matches DeepSeek V4 Pro Base's performance using roughly 50 times less compute, at an estimated cost of about $0.5 million on GB200 chips. Scaling the same recipe up 10x further, to roughly $4 million, reportedly surpassed all publicly available open base models on perplexity benchmarks.

StartLux's 27B local model nearly matches DeepSeek-V4-Pro in CAICT MCP test

A newcomer called StartLux, founded by veteran programmer Chen Danyan and formerly known as Yuandian Xinghui, placed second in China's CAICT MCP specialized benchmark with its StartLux-V1.0-27B-Preview model. The 27-billion-parameter model trailed DeepSeek-V4-Pro, a 1.6-trillion-parameter system, by just 1.3 percentage points, despite running locally on consumer-grade PCs rather than in the cloud. Chen has argued that local models will eventually outcompete cloud-based ones on efficiency and market share.

AI coding tools speed up typing but not startup creation, critic argues

Writer Jordan Andersen expands on a post by Disesdi Shoshana Cox questioning why, despite AI supposedly boosting developer productivity tenfold, no wave of new billion-dollar startups comparable to Airbnb, Stripe, or Dropbox has emerged. Instead, the biggest winners of the generative AI boom are AI companies themselves—OpenAI, Anthropic, and High-Flyer—while broader effects include cluttered inboxes, spammy social content, and degraded search results.

Testing shows 8x RTX PRO 6000 rigs excel at parallel serving, not giant model sharding

Jerry James benchmarked an 8-GPU NVIDIA RTX PRO 6000 Blackwell system paired with an AMD EPYC 9555 CPU, offering 768GB of combined GDDR7 memory. Rather than splitting huge models like Llama-3.1 405B or DeepSeek-R1 671B across all eight PCIe cards, which introduces heavy latency, the team found the setup performs best running independent single-GPU model instances in parallel.

vLLM 0.28.0 ships major performance upgrades for Kimi-K3 and DeepSeek V4

The vLLM project released version 0.28.0, combining 584 commits from 270 contributors. The update centers on deep performance work for Kimi-K3, including new decode context parallelism, fused kernels, memory-saving expert sharding, and ROCm support, alongside DeepSeek V4 improvements such as end-to-end sparse MLA and AMD Quark NVFP4 support.

DeepSeek seeks outside funding as High-Flyer's IPO bets grow riskier

DeepSeek, the AI lab founded by Liang Wenfeng, has relied on his hedge fund High-Flyer Quant for early funding and computing power, but is now courting outside investors as its needs expand. Meanwhile, High-Flyer's affiliates have taken sizable pre-IPO stakes in Chinese hard-tech firms like memory chipmaker CXMT and robotics company Unitree, even as volatility in AI and chip stocks has hit the fund's returns.

Nvidia expands optimization support for Chinese open AI models, flags US regulatory risk

Nvidia announced hardware optimizations for popular Chinese open-source AI models, including DeepSeek's V4 Flash and Alibaba's Qwen 3.8, as part of a broader push to support open-weight systems. Alongside the announcement, an SEC filing tied to its quarterly earnings warned that potential White House restrictions on Chinese-origin AI could hurt its business by limiting its ability to support apps built on those models.

DeepSeek Reportedly Targeting $74 Billion Valuation in New Funding Round

Chinese AI startup DeepSeek, based in Hangzhou, is said to be raising fresh capital that would value the company at roughly $74 billion. The funds are intended to support research and development along with an expansion of its computing infrastructure.