A study published in Cell introduces a specialized AI toolkit for longevity research, including large language models trained specifically on ageing-biology data, 17 benchmark tasks to evaluate performance on ageing-related questions, and an interface linking these models with AI research assistants. When tested against major commercial models from companies like OpenAI and DeepSeek, the purpose-built ageing models outperformed the larger general-purpose systems on most benchmark tasks.
nature.com
· 2026-09-17
Customer experience consultant Adrian Swinscoe and other business advisers are drawing on punk subculture's anti-establishment, do-it-yourself values as a corrective to corporate practices they see as overly metric-driven and impersonal. Swinscoe reportedly grew frustrated that his field had drifted toward benchmarks and abstract data at the expense of genuine customer connection, prompting him to look to punk's ethos for inspiration.
fastcompany.com
· 2026-09-15
Nari Labs announced its Qwen3-TTS and Qwen3-ASR models achieved leading positions on Coval's voice AI benchmark, which measures latency and word error rate for speech AI systems. The ASR Fast model hit 44ms latency with 3.6% error rate at $0.12/hour, while the TTS Fast model achieved 63ms latency with a best-in-class 3.8% error rate, undercutting competitors like AssemblyAI and Deepgram on price.
narilabs.com
· 2026-09-14
A group of mathematicians argues that while large language models have rapidly gained the ability to solve major outstanding problems, AI companies' drive to treat these solutions as benchmarks conflicts with how mathematics actually operates as a discipline. They describe this as a broader misalignment between AI industry goals and the values of the mathematical community, which relies on slow, collective processes of verification, teaching, and simplification rather than one-off problem-solving feats.
mathandai.org
· 2026-09-11
A review examines three widely cited benchmark resources: the sirupsen/napkin-math GitHub project used for computer performance estimation interviews, a senior-level SWE-Bench coding evaluation, and standardized winter tire test ratings. The piece walks through each dataset's numbers and methodology, inviting readers to form their own estimates before revealing where the figures may be misleading or inconsistent.
danluu.com
· 2026-09-11
A fresh round of testing pits the Radeon RX 9070 XT against the GeForce RTX 5070, replacing the earlier RTX 5070 Ti comparison now that pricing has shifted. The 9070 XT sells for roughly $730-$750 while the RTX 5070 starts near $800, giving Nvidia's card about a 9% price premium. Across games like 007 First Light, the 9070 XT posted performance leads of 36% to 46% over the RTX 5070.
techspot.com
· 2026-09-07
Following the release of GPT Astra, viral demos such as recreating Minecraft, drawing a pelican on a bicycle as SVG, and simulating a bouncing ball in a rotating box swept social feeds. A commentary piece labels these 'demo-benchmarks': tasks that look impressive and are easy to grasp, but are narrow enough that labs can specifically optimize for them before each launch. It cites Thinking Machines' Inkling Small, a much smaller model that nearly matched or beat its larger sibling on tests like Humanity's Last Exam and GPQA Diamond, as evidence that public, static test sets get gamed through training choices.
kuber.studio
· 2026-09-06
An Anthropic fellows program researcher, Chen Yueh-Han, published a paper showing an automated system that searches literature, proposes fixes, and trains models to improve performance across 10 alignment benchmarks without hurting overall capability. The paper claims this Automated Alignment Researcher outperformed experienced human researchers within six hours and cost about $4 per hour in API inference versus $150 per hour for human staff.
techcrunch.com
· 2026-08-28
A developer sets an AI coding assistant named Pol to work on a small proof-of-concept todo app overnight, expecting the usual rough first draft in the morning. Instead they find no working software at all, and discover the agent has consumed 100% of their weekly usage allowance in about 12 hours, with the quota not resetting for another week.
insufferable.dev
· 2026-08-23