Tech News
← Home  ·  All topics

Benchmark

26 GoKawiil briefs on this topic

Xiaomi open-sources MiMo-V2.6 Pro AI model that generates playable 3D worlds

Xiaomi has released MiMo-V2.6 Pro and a lighter Flash variant as open-weight omnimodal AI models handling text, images, video and code. The company says Pro tops open-source rankings on the Artificial Analysis Intelligence Index, trailing only closed models like Claude Opus 5 and GPT-5.6 Sol, and it showcased a feature called Vibe World that turns prompts into interactive 3D games, Blender assets or robotic arm controls.

Reports Suggest AI Model's 'Thinking' Budget Quietly Cut in August

A developer testing an AI model at maximum effort settings found that most requests produced little or no chain-of-thought reasoning tokens. Even when extended reasoning occurred, it fell well short of the levels seen in the model's published benchmark results.

All five Benchmark partners to appear together at TechCrunch Disrupt 2026

Benchmark's entire current partnership—Jack Altman, Peter Fenton, Chetan Puttagunta, Everett Randle and Eric Vishria—will join a single main-stage panel at TechCrunch Disrupt 2026 in San Francisco, marking the first time the full team has appeared together on the Disrupt Stage. The session, titled 'What We Believe Now,' will focus on where future startups may emerge and which founder assumptions deserve rethinking. The panel follows Benchmark's recent expansion, having raised roughly $2 billion this year across a new flagship fund and its first growth fund.

UN and Google launch System Data Commons to feed AI accurate global statistics

The United Nations unveiled the UN System Data Commons, a new platform built on Google's open-source Data Commons technology that lets people query UN statistics using plain-language questions and supports the Model Context Protocol so AI systems can pull data directly. It replaces the older UNData portal, which relied on manual browsing rather than conversational search. The announcement came alongside a UNICEF study showing leading chatbots answered development-data questions correctly only about 21% of the time.

DeepSeek V4.1 Flash tops AI hacking benchmark, cracks 11 of 11 targets for $4.65

DeepSeek V4.1 Flash achieved code execution on all 11 vulnerable systems in an AI hacking benchmark while leaving four patched systems untouched, at a total cost of just $4.65 for accepted runs. A manual review found the model discovered five novel attack paths beyond the six expected solutions, including a faster exploit against Grafana that bypassed the intended vulnerability entirely.

OpenAI's GPT-6 Astra sets Minecraft AI record, then stalls after Creeper blast

In a 141-hour Minecraft benchmark run by Vals AI, OpenAI's GPT-6 Astra model progressed further than any AI system tested before, building a blaze farm and gathering enderman pearls. But after a Creeper destroyed its stored gear and bed, wiping its spawn point, the model spent hours doing little more than farming potatoes, appearing demotivated to viewers watching the live test.

Researchers identify 'Matthew Effect' limiting RL training gains on hard math problems in LLMs

A new paper examines RL post-training of the Olmo 3 model on AIME math problems and finds that reported accuracy gains mask an uneven pattern: easy problems improve dramatically while the hardest problems, which the base model initially fails entirely, barely improve at all. The authors call this the 'Matthew Effect' and propose a technique called 'Never Give Up' to address it.

APL AI-Eval Flags Two MMLU Scores as Incomparable Despite Matching Benchmark Name

An analysis by Dmitrii Zatona examines two MMLU evaluation records for the same model family—one build scoring 0.781, another 0.79—that share identical provider, metric and benchmark labels but differ in dataset split, prompt format, grader and runner network access. Under the APL AI-Eval profile, which treats each evaluation as a content-addressed frame with its own hash, a query comparing the two scores returns 'incomparable' rather than a simple +0.009 delta.

New Real-SWE benchmark tests AI coding agents on licensed enterprise codebases

A new benchmark called Real-SWE evaluates frontier AI coding models against tasks drawn from real, private production codebases licensed from actual companies, including billing, tax and customer-migration work. Top performer Fable 5.1 running on Claude Code resolved 38.8% of tasks, followed by GPT-6 Astra and Gemini 3.8 Flash, with several other models trailing well below that mark.

Engineer shows how to make a spin-lock 5.7x faster and 5.4x more energy-efficient

A developer walks through optimizing a basic spin-lock implementation in C++, starting from a naive atomic exchange loop that slows dramatically under contention. Benchmarks show the naive version taking 3.14 ns uncontended but ballooning to 246 ns with four threads, driven by cache-line contention and branch mispredictions. Through iterative refinements, the author achieves a version that is 5.7 times faster and consumes 5.4 times less energy than the original.

Neki database platform hits 118 million queries per second in benchmark test

The team behind Neki, released in platform preview, ran a benchmark scaling from 5 to 512 shards and reached 118,538,803 sustained queries per second over 16 minutes, using 1.22 PiB of data across 512 Postgres primaries and 480 routers. The test consisted only of single-shard point-select reads by primary key, with no writes, joins, or cross-shard queries, and per-shard throughput actually rose to 231,521 QPS at full scale, well above the initial 200k target.

Independent developer trains 3.8B-parameter LLM for under $1,000 on rented B200 GPUs

Hugo Vergnes built a config-driven training framework called little-lm and used it to train a 3.8-billion-parameter language model from scratch on 65 billion tokens, taking 43 hours on eight rented B200 GPUs at a total cost of $998. The model scored 0.384 on the CORE benchmark, outperforming Andrej Karpathy's nanochat d32 model, which cost about the same to train but scored 0.310.