Skip to content
Tech News
← Back to articles

Why does Opus 5 feel worse to work with?

read original more articles
Why This Matters

This article highlights the challenges faced when developing advanced AI models like Opus 5, which, despite outperforming previous versions in benchmarks, feels less user-friendly due to its less collaborative behavior. It underscores the tension between optimizing for benchmark performance and creating AI that aligns with real-world needs for clarification and adaptability, a crucial consideration for the future of AI development and deployment in the tech industry.

Key Takeaways

In my opinion and that of the colleagues I've spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable.

I'm not claiming a step backwards in capabilities – it is a more capable model than Opus 4.7 and Opus 4.8 and even rivals Fable in benchmarks, yet these other models feel better to work with. I believe this is because they:

stop and ask questions if my intent was unclear,

don't make assumptions without checking,

and don't reinterpret or update my plans without asking.

Because of this, they don't require the careful babysitting that Opus 5 does.

I suspect this is the result of two compounding forces at Anthropic, and in current frontier labs in general.

First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI.

Second, the pressure to score highly on benchmarks. Although it's an open secret that many benchmark tasks are ill-defined, unfair, hackable, or otherwise broken, a good benchmark task is self-contained. It can be solved. It doesn't require hints, reading the task creator's mind, or outside information to pass.

That doesn't mean a good task can only have one correct answer, just that it should score all unambiguously correct answers equally.

... continue reading