Skip to content
Tech News
← Back to articles

GPT-6 Astra on robot arms

read original get Melissa & Doug Wooden Shape Sorting Puzzle Board → more articles
Why This Matters

A controlled head-to-head benchmark shows OpenAI's GPT-6 Astra hitting 19/20 on a simple pick-and-place task versus Fable 5.1's 8/20, while running faster and at roughly half the cost per run. But on a precision insertion task both models stall at the same final step (2/20 each), suggesting general model gains don't automatically translate into fine manipulation. For anyone betting on general-purpose models as robot controllers, that gap is the key data point.

Key Takeaways
Worth a Look

Melissa & Doug Wooden Shape Sorting Puzzle Board — The peg-and-groove insertion task that stumps these models is exactly what a classic knobbed shape-sorting board tests: align the piece, orient it, and drop it home. It's a fun, tactile benchmark to try yourself (or with a kid) while you watch robot arms stall on the final millimeter. Sturdy wooden pieces with grab knobs make it easy to handle.

See Melissa & Doug Wooden Shape Sorting Puzzle Board on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

GPT‑6 Astra on robotic manipulation

A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra control of the same YAM arms under the same Inspect Robots agent policy, on the same two tasks:

“Pick up the red block from the table and place it inside the bowl.”

“Pick up the round blue puzzle piece by the knob at its center and place it into the matching circular groove in the board.”

On the bowl task Astra placed the block in 19 of 20 trials, against Fable 5.1's 8 of 20 and Fable 5 in 1 of 20, in 2.5 minutes per trial to Fable 5.1's 6.8, at an estimated $0.94 per run to $2.12.

The puzzle task is a different story: Astra completed the insertion 2 times in 20 against Fable 5.1's 2 in 20. It reaches the groove and stalls at the same final step Fable does, at $1.36 per run to $2.18.

Block into bowl: the best completed run of each model (highest stage, then shortest), each played in its own time at the same speed‑up. Timers show real elapsed time with thinking pauses removed.

Astra completes the bowl task far more often, at about half the cost per run

2026-09-04T18:50:02.475437 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/ $1 $2 $3 0 20 40 60 80 100 completion rate (%) Fable 5 Fable 5.1 GPT-6 Astra 2.4× higher completion rate 2.3× cheaper Block into bowl $1 $2 $3 Fable 5 Fable 5.1 GPT-6 Astra same completion rate 1.6× cheaper Puzzle piece into groove estimated cost per run (USD, list price)

Large dots are condition means; faint dots are individual trials (100 if completed, 0 otherwise) at their own cost.

... continue reading