New benchmark PacBench tests AI models on one-shot Pac-Man coding task
A Show HN project called PacBench evaluates how well different AI models and their supporting harnesses can recreate the classic game Pac-Man from a single prompt: 'Create a Pac-Man game in a single html page.' The benchmark measures how closely each model's generated code approximates a functioning Pac-Man game.
GoKawiil's interpretation of the reporting above, not reported fact.
This kind of benchmark could offer a simple, visual way to compare coding capabilities across large language models, since results are immediately playable and easy to judge by eye. It suggests a growing trend of creative, task-specific benchmarks that go beyond standard coding tests to gauge real-world usability of AI-generated software.
- PacBench is a new benchmark for testing AI coding ability via a single-prompt Pac-Man recreation task.
- It evaluates both the model and the surrounding harness used to generate the game.
- The project was shared as a Show HN post, indicating it's an independent, community-driven experiment.
Source: jonclegg.github.io, 2026-09-28
Published there as: “Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.