Four lanes, one game: same serve, same rules, same question, asked of a different model in each. One answer moves the ball one step, so every ball travels at its own model's answer time and nothing is skipped or sped up. Jev is a typed-decision model rather than a chat model, and Pong is the Atari game chat models handle worst, so a loop that won't wait is the fair place to see what that's worth.
Show HN: Jev vs. GPT-5.6 and Claude Haiku at Pong
Why This Matters
This piece is a lighthearted but telling benchmark of AI response latency and decision-making, pitting a specialized 'typed-decision' model against general-purpose chat models like GPT-5.6 and Claude Haiku in a real-time Pong game. It highlights a growing industry conversation about whether chat-oriented LLMs are well-suited for low-latency, action-based tasks, and whether purpose-built models can outperform them where speed and precision matter.
Key Takeaways
- The test compares four AI models in a real-time Pong game where each model's move speed depends on its own actual response time, with no artificial syncing.
- Jev is described as a typed-decision model, contrasting with chat-based models like GPT-5.6 and Claude Haiku, suggesting different architectures suit different real-time tasks.
- Pong is highlighted as a particularly hard task for chat models, making it a revealing stress test for latency and decision quality in interactive AI applications.
Get alerts for these topics