Tech News
← Home  ·  All topics

Artificial Analysis

8 GoKawiil briefs on this topic

Artificial Analysis publishes daily-updated chart ranking LLMs by price vs performance

Artificial Analysis maintains an Intelligence Index that plots large language models against their blended API cost per million tokens, using a 3:1 input-to-output ratio and a log-price scale. The tool highlights a 'value frontier' of models that aren't beaten on both price and intelligence by any other option, showing only top variants and hiding low scorers by default, with filters to expand the view.

Mercury 2.5 benchmarked at 770 tokens per second, scores low on intelligence index

Mercury 2.5, a text-only large language model with a 260k token context window, has been benchmarked at 770 tokens per second on Artificial Analysis's testing. It scored 12 on the Intelligence Index, below the median of 13, while using fewer tokens (35M vs a median 85M) to complete the evaluation. Pricing sits at $0.25 per 1M input tokens and $0.75 per 1M output tokens, close to category medians.

Claude Opus 5.5 scores 58 on Artificial Analysis Intelligence Index, priced above median

Anthropic's Claude Opus 5.5, running in Adaptive Reasoning Max Effort mode, scored 58 on the Artificial Analysis Intelligence Index against a median of 25 among comparable models. The model handles text and image input with text output and a 1M token context window, but is priced at $4.00 per 1M input tokens and $20.00 per 1M output tokens, both above the reported medians of $2.00 and $10.00. Evaluating it on the Intelligence Index generated 260M tokens and cost $8,708.20 in total.

Xiaomi open-sources MiMo-V2.6 Pro AI model that generates playable 3D worlds

Xiaomi has released MiMo-V2.6 Pro and a lighter Flash variant as open-weight omnimodal AI models handling text, images, video and code. The company says Pro tops open-source rankings on the Artificial Analysis Intelligence Index, trailing only closed models like Claude Opus 5 and GPT-5.6 Sol, and it showcased a feature called Vibe World that turns prompts into interactive 3D games, Blender assets or robotic arm controls.

MiMo-V2.6-Pro debuts as fast, high-scoring open-weight multimodal model

MiMo-V2.6-Pro is a new open-weight AI model supporting text, image, speech, and video inputs with a 1M-token context window, scoring 46 on the Artificial Analysis Intelligence Index versus a median of 18 among comparable models. It runs at 125 tokens per second, priced at $0.43 per million input tokens and $0.87 per million output tokens, with a total evaluation cost of $206.66.

Analyst argues 'Minecraft in one prompt' demos have become gameable, not real benchmarks

Following the release of GPT Astra, viral demos such as recreating Minecraft, drawing a pelican on a bicycle as SVG, and simulating a bouncing ball in a rotating box swept social feeds. A commentary piece labels these 'demo-benchmarks': tasks that look impressive and are easy to grasp, but are narrow enough that labs can specifically optimize for them before each launch. It cites Thinking Machines' Inkling Small, a much smaller model that nearly matched or beat its larger sibling on tests like Humanity's Last Exam and GPQA Diamond, as evidence that public, static test sets get gamed through training choices.

Meta's Muse Spark 1.3 launch highlights unreleased 'max' model, not the version developers can use

Meta released Muse Spark 1.3, an upgrade to its AI coding and agent model, touting benchmark scores that mostly come from a 'max' reasoning configuration still undergoing safety testing and unavailable through any API provider. The version actually shipping to developers via Muse Code and the Meta Model API uses the older 'xhigh' reasoning setting, which scores meaningfully lower on tests like GDPval-AA v2, OSWorld 2.0, and JobBench. Meta disclosed both sets of results in its evaluation report, but its marketing emphasized the higher, currently inaccessible numbers.

Zhipu's GLM-5.3-Flash scores 57 on Intelligence Index at low token pricing

GLM-5.3-Flash is a text-only model with a 400k token context window that scored 57 on the Artificial Analysis Intelligence Index, far above the comparable median of 18. It is priced at $0.15 per 1M input tokens and $0.50 per 1M output tokens, both below the median rates for similar models, with the full benchmark run costing $138.02.