Skip to content
Tech News
← Back to articles

Recreating Minecraft Is Not a Benchmark

read original get Minecraft: Java & Bedrock Edition Deluxe Collection → more articles
Why This Matters

The piece argues that viral AI launch demos — one-prompt Minecraft clones, SVG pelicans on bicycles — have become 'demo-benchmarks' that labs can deliberately optimize for, making them marketing artifacts rather than measures of capability. The same overfitting risk applies to public static evals, where smaller models can outscore larger ones on leaderboards while feeling worse in practice. For buyers and developers, it's a warning that launch-day wow moments say little about real-world model quality.

Key Takeaways
Worth a Look

Minecraft: Java & Bedrock Edition Deluxe Collection — If all this talk of AI "recreating Minecraft in one prompt" makes you want the real thing, grab the actual game and go build something yourself. It's the endlessly creative sandbox every benchmark demo is trying to imitate, and no model overfitting required.

See Minecraft: Java & Bedrock Edition Deluxe Collection on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

GPT Astra released a couple of days ago and, inevitably, within the hour my entire feed was the same five things: recreating Minecraft in one prompt, painting themselves in MS Paint, the pelican riding a bicycle as an SVG, a ball bouncing in a rotating box with believable gravity, and an SVG game controller.

On paper these look like harder, more visual problems for a model to solve, there’s a reason they’re as big as they are. I’ve started calling them demo-benchmarks, visual and understandable enough for everyone to get but finite enough for the next model to be “perfect” on.

That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.

It’s not really their fault either, honestly I’d say it’s dumb if they didn’t - nothing sells a launch like a pelican or a 3D game controller the timeline can’t stop quoting.

A fixed, famous target and eight weeks of runway is a solved pelican, these tests never change and anything that never changes can be overfit. Every launch cycle proves it again.

A test you can perfect on a schedule measures preparation instead of capability, to me that’s anti the very definition of a benchmark, it should be a hard test, something very hard to perfect.

The same dynamic runs through the open evals, smaller models that feel dumber in practice still outscore better ones on sites like Artificial Analysis. This isn’t hypothetical, Thinking Machines’ Inkling Small scored within a point of its flagship sibling on the Artificial Analysis Intelligence Index with less than a third of the parameters and beat it on Humanity’s Last Exam, GPQA Diamond and SciCode. Public, static, famous test sets leak into training data and fine-tuning choices.

A launch is a first impression and first impressions are marketing, that’s why you’re always bound to be shocked - the shock was scheduled.

So what’s the alternative? Honestly, I’m not sure

because if you think about it, the obvious fix somewhat already exists. LiveBench rotates its questions, ARC-AGI keeps a private set, Humanity’s Last Exam holds part of itself back. Tests where the tested party doesn’t know what’s being tested: you can’t teach to a test that hasn’t been written yet.

... continue reading