Ideally yes but that would not work with the model latency. But it is already partly covered when a ghost enters its corridor it is asked again. Here's how the decisions work in detail, also below a copy https://github.com/opper-ai/jevman-benchmark/blob/main/READM...
How decisions work
When a character commits to a corridor its next junction is known, so the game asks jev about it straight away (one System One request per frame, one choice question per character, options = legal directions described with computed facts: distances to pellets, power pellets, fruit and ghosts, whether the nearest ghost is coming closer, and whether a ghost can reach the end of the corridor before Pac-Man). Up to three requests are in flight at once. If the character reaches the junction before the answer, it waits there (its panel card says "thinking…"). After 2 s, or on an error, it uses a greedy rule.
Pac-Man can also get a second question mid-corridor: when a dangerous ghost is in the corridor ahead, or can reach the junction at its end before he does, jev is asked whether to keep going or turn back right now (pacman_escape). Pac-Man keeps moving while it is open, and each situation is asked once.
Ideally yes but that would not work with the model latency. But it is already partly covered when a ghost enters pacmans corridor he is asked again. Here's how the decisions work in detail https://github.com/opper-ai/jevman-benchmark/blob/main/READM...
Quick update, we doubled the models tested based on community contributions, and a new model took the #1 spot. also well over 1000 games have been played against decision model ghosts.
https://opper.ai/jevman-benchmark/#leaderboard
This is neat. At least I'm still better than the models at pac-man.
Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.
Great to hear, if it's about inference credits we actually built a thing called an AI wallet where users can just easily pay for their own inference on apps. if that's interesting lmk, id be happy to support this https://opper.ai/ai-wallet
Running cloud API inference for Pac-Man is peak modern software engineering. A 2 MB local ONNX model on a CPU core beats network round-trip jitter every time.
Shouldn't it be running at every tick of the game rather than just the junctions? Pac-Man should be able to change direction at any point.
Ideally yes but that would not work with the model latency. But it is already partly covered when a ghost enters its corridor it is asked again. Here's how the decisions work in detail, also below a copy https://github.com/opper-ai/jevman-benchmark/blob/main/READM...
How decisions work
When a character commits to a corridor its next junction is known, so the game asks jev about it straight away (one System One request per frame, one choice question per character, options = legal directions described with computed facts: distances to pellets, power pellets, fruit and ghosts, whether the nearest ghost is coming closer, and whether a ghost can reach the end of the corridor before Pac-Man). Up to three requests are in flight at once. If the character reaches the junction before the answer, it waits there (its panel card says "thinking…"). After 2 s, or on an error, it uses a greedy rule.
Pac-Man can also get a second question mid-corridor: when a dangerous ghost is in the corridor ahead, or can reach the junction at its end before he does, jev is asked whether to keep going or turn back right now (pacman_escape). Pac-Man keeps moving while it is open, and each situation is asked once.
Ideally yes but that would not work with the model latency. But it is already partly covered when a ghost enters pacmans corridor he is asked again. Here's how the decisions work in detail https://github.com/opper-ai/jevman-benchmark/blob/main/READM...
It may not matter if the ghosts' next move is predictable over that maximum distance.
That requires much more thinking (because of calculation) though doesn't it? Which decision models are not optimized for.
Speculating, but it seems like needing to be able to change direction at any moment would be more calculating, not less.
Quick update, we doubled the models tested based on community contributions, and a new model took the #1 spot. also well over 1000 games have been played against decision model ghosts. https://opper.ai/jevman-benchmark/#leaderboard
i guess a model trained on this game will end up taking the #1 at some point
This is neat. At least I'm still better than the models at pac-man.
Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.
Great to hear, if it's about inference credits we actually built a thing called an AI wallet where users can just easily pay for their own inference on apps. if that's interesting lmk, id be happy to support this https://opper.ai/ai-wallet
yeah we had a lot of success with this also with another project, ai roundtable: https://opper.ai/ai-roundtable/history
Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?
I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy
I’ll try training one for this benchmark, seems like fun
Not yet but the repo is open source and set up so that you can benchmark your own model and we can add it to the leaderboard, would be cool to see if Jeffy can play: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...
Jev is useful to run locally overnight: it can classify the results while you sleep, ready for you to review in the morning.
Yes, try with toxic/toxichat and CFPB
Running cloud API inference for Pac-Man is peak modern software engineering. A 2 MB local ONNX model on a CPU core beats network round-trip jitter every time.
Super cool!
I wish the controls were a little easier on mobile.
Thank you and yea agreed fixing it now, try again in an hour from now :)
it would be nice to have a regular LLM for reference
Paying 2 cents per game to avoid playing it.
What even is this reality.
haha thats one way to look at it :)
very cool :)