Jevman benchmark pits AI language models against classic Pac-Man ghosts
Jevman ran 100 games per AI model against scripted Pac-Man ghosts, with each game ending after three lost lives or a 5-minute cap. At every junction the model was asked to weigh the maze, pellets and ghost positions and output movement probabilities, with a 2-second deadline before a backup rule took over. Results were ranked by mean score with a 95% margin of error, and models within that margin were treated as tied.
GoKawiil's interpretation of the reporting above, not reported fact.
The project suggests game-based benchmarks can probe how well language models handle real-time spatial reasoning and decision-making under time pressure, not just text tasks. The fact that no game lasted more than 2 minutes 24 seconds hints current models still struggle with the kind of continuous, reactive strategy humans find easy. Benchmarks like this could offer a simple, visual way to compare model reasoning speed and consistency across versions.
- Each AI model played 100 Pac-Man games against standard scripted ghosts.
- Models had 2 seconds to decide moves before a backup rule intervened.
- Scores were ranked using mean performance with statistical margins to determine ties.
Source: opper.ai, 2026-10-08
Published there as: “Show HN: Jevman – AI decision models play Pac-Man”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.