18 comments

  • jmathai 38 minutes ago ago

    This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving.

    Remember when people considered you a genius for prompting with "You are a skilled writer....".

    • wredcoll 37 minutes ago ago

      That moment in time was physically painful.

  • strataspace 5 hours ago ago

    I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.

    The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.

    • shoobiedoo 42 minutes ago ago

      My hope is that this drives a new generation of hyper creative content from those who don't give up. I mean, of course it does pacman well. Pacman and its clones have been done to the death over decades. But what if we ask AI to write finnegan's wake 2?

    • acomjean 2 hours ago ago

      I wonder if we need better programming abstractions/ languages that can make programming easier for people.

      It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.

      They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.

  • continuational 41 minutes ago ago
    • thefourthchime 32 minutes ago ago

      I need to add that one! It's the first one I've seen to have almost a drunk Pac-Man. It lags a little bit.

  • _matthew_ an hour ago ago

    I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.

  • dang 2 hours ago ago
  • hoistway an hour ago ago

    Always assumed Pac-Man was an easy solve for modern AI. Guess those ghost patterns are trickier than they look for one-shot learning.

    • thefourthchime 4 minutes ago ago

      The game mechanics are shockingly hard for them to get right. You also reveal some of their personality traits. The GPT‑6 and Astra models tend to include a lot of cringe copy as well as their own personal graphics style.

    • gedy 35 minutes ago ago

      I did a variant of Pac Man using Claude and pointed it at a thorough blog post about the patterns and that greatly helped.

  • nedo_var an hour ago ago

    Curious if models grasp the ghost patterns or just react. Pac-Man's more complex than it seems for one-shot learning.

  • blindflag an hour ago ago

    I'm curious; can you explain why you picked Pac-Man, in particular?

    • thefourthchime 40 minutes ago ago

      It was something I randomly tested about a year ago, and no models handled it well, so whenever a new model came out, I tried it.

      It has the advantage of taking good screenshots and letting me know if it’s decent within five seconds of playing.

  • Computer0 7 hours ago ago

    Opus 5-5 seemed like a perfect clone, with others displaying flaws in an initial look. Astra notably created a bunch of surrounding ugly crap to look at.

    • thefourthchime 30 minutes ago ago

      Yes, Opus 5.5 is the first one that I don't have any notes on. It's just as impressive on other tasks I've given it over the last week. A clear step change.

    • vunderba 7 hours ago ago

      I was just coming here to say this. I looked at all of them, and Opus 5.5 at high effort seems significantly better than every other clone.

      Buffered controls, different pathfinding AI for each ghost, level transitions, etc.

      Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.