Rendered at 05:37:30 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
yambam 23 hours ago [-]
This reminds me of a couple decades ago when some friends and I set ourselves a challenge to, individually, each create as much of a Pac-Man clone as possible in 10 hours. None of us had any experience in games programming or graphics coding. It was great fun and we all learned a lot.
No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.
I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.
OJFord 22 hours ago [-]
I think you figure out different things for yourself; the next generation is going to have 'programming' appeal to them for different reasons. More creatives. More engineers, perhaps, as the practice becomes more about the bigger picture.
HPsquared 19 hours ago [-]
It'll be more like architecture than engineering, I think. Sketch out the building plan, context (physical and social) and desired usage.
OJFord 9 hours ago [-]
I think we're saying broadly the same thing. (Old school) Software engineering is fairly unlike most engineering disciplines. We've been both the engineer and the technician (and the latter's taken a lot of time, especially as at lower levels) but now Claude (or equivalent) is capable of being the technician, at least.
zerr 18 hours ago [-]
Articulating about architecture (as opposed to engineering (as opposed to programming)) is a part of coping I believe. Yes, LLMs do architectures as well. I think the vibe-coding is the more precise term, also considering that less and less human reviews are being done.
drcxd 22 hours ago [-]
Interesting, recently I am working on my own clone of Pac-Man. LLM implementations lose lots of details. They are not 1:1 replication of the original game. For example, the behavior of the ghost is not the same as the original. If I have not implemented the game myself, I can not tell the differences. What LLM produced look like the original game, but they are not.
CamperBob2 13 hours ago [-]
But it takes -- what, two or three sentences? -- to explain how the ghosts should move. The interesting (and important) thing is not that the model gets it wrong at first, it's how easy it is to correct it.
One-shotting something like Pac-Man doesn't prove much. At the end of the day, one-shot fidelity is going to scale more or less linearly with model size/world knowledge. Why wouldn't it?
strataspace 1 days ago [-]
I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.
The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.
acomjean 1 days ago [-]
I wonder if we need better programming abstractions/ languages that can make programming easier for people.
It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.
They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.
Joel_Mckay 23 hours ago [-]
>I wonder if we need better programming abstractions/ languages that can make programming easier for people.
That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.
Successful ecosystems are just pedantically lame enough to keep silly folks from going YOLO, but empowering enough to allow people to still have fun.
>which model has the best training code that was Pac-Man.
You mean which model is more cautious about copyright bleed-through of the $9Tn in FOSS and user code they misappropriated though isomorphic plagiarism. =3
boxed 23 hours ago [-]
> That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.
Nodejs wasn't easier for people. JS is a horrible language where simple stuff like comparisons, array access, and member access are broken.
fuzzythinker 12 hours ago [-]
Hope this gives a speck of encouragement. Just want to call out how I love the clarity in Three.js Resources, in readability, design, and answering why. Eg. https://threejsresources.com/tsl
shoobiedoo 1 days ago [-]
My hope is that this drives a new generation of hyper creative content from those who don't give up. I mean, of course it does pacman well. Pacman and its clones have been done to the death over decades. But what if we ask AI to write finnegan's wake 2?
Vakaiser 23 hours ago [-]
I’m quite optimistic about how AI tooling will raise the floor for creative work.
I personally have been working on a Three.js project with Opus 5 and 5.5 that I never would have continued with had I needed to dive into documentation by hand.
Seeing immediate results is incredibly motivating.
8n4vidtmkvmk 23 hours ago [-]
I'm finding new things to demotivate me. It's good and quick at doing everything I ask but sometimes as soon as it comes together I realize I didn't really want this or it's just not fun.
But maybe iterating quicker and finding that out sooner is still good. I don't know.
rico735 9 hours ago [-]
I suspect any attempt to market will be where demotivation kicks in. When they entry level is so low, far fewer people are going to bother giving anything a click.
Whilst it is the Astra aesthetic, the expectation of any three.js game will be that it is all surface and no detail, whether someone has put the effort in or not.
lukan 22 hours ago [-]
"The ThreeJS dude posted ab how demotivated he was to continue his work"
This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving.
Remember when people considered you a genius for prompting with "You are a skilled writer....".
wredcoll 1 days ago [-]
That moment in time was physically painful.
_matthew_ 1 days ago [-]
I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.
Curious if models grasp the ghost patterns or just react. Pac-Man's more complex than it seems for one-shot learning.
xnx 19 hours ago [-]
This would be a much better benchmark if it contained some twist which was not in the training set (e.g. pelicans on bicycles were not common before simonw).
Make a pacman game where pacman can always eat ghosts, but ghosts drop pellets.
darepublic 17 hours ago [-]
The sad thing is that sol takes under 5 min here but takes 10 minutes to make focused targetted change to crud app
weitendorf 23 hours ago [-]
Gonna be rude and say I don't think this is an interesting or useful benchmark tbh.
Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.
Everything is going to go from 0->1 on this benchmark in short order because of that.
swingboy 20 hours ago [-]
Interesting the `high` effort on the Claude models. Did you find that going above that is unnecessary?
hoistway 1 days ago [-]
Always assumed Pac-Man was an easy solve for modern AI. Guess those ghost patterns are trickier than they look for one-shot learning.
thefourthchime 24 hours ago [-]
The game mechanics are shockingly hard for them to get right. You also reveal some of their personality traits. The GPT‑6 and Astra models tend to include a lot of cringe copy as well as their own personal graphics style.
dansquizsoft 19 hours ago [-]
Yeah I noticed that (the extra copy) as soon as I started playing each of the GPT-6 ones... Its gross!!!
RugnirViking 22 hours ago [-]
oh god, yeah. Thats difficult to look at
gedy 1 days ago [-]
I did a variant of Pac Man using Claude and pointed it at a thorough blog post about the patterns and that greatly helped.
mrblinky 23 hours ago [-]
Pretty underwhelming. Surely they’ve ingested some existing code on the web and regurgitating that.
jonatron 22 hours ago [-]
Try coming up with something that doesn't exist yet on the web, one of the good models and a harness will be able to do it. Yes it's hard to believe but you can't deny the current reality.
darepublic 17 hours ago [-]
This would be a more telling experiment perhaps
blindflag 1 days ago [-]
I'm curious; can you explain why you picked Pac-Man, in particular?
thefourthchime 1 days ago [-]
It was something I randomly tested about a year ago, and no models handled it well, so whenever a new model came out, I tried it.
It has the advantage of taking good screenshots and letting me know if it’s decent within five seconds of playing.
Computer0 1 days ago [-]
Opus 5-5 seemed like a perfect clone, with others displaying flaws in an initial look. Astra notably created a bunch of surrounding ugly crap to look at.
thefourthchime 1 days ago [-]
Yes, Opus 5.5 is the first one that I don't have any notes on. It's just as impressive on other tasks I've given it over the last week. A clear step change.
vunderba 1 days ago [-]
I was just coming here to say this. I looked at all of them, and Opus 5.5 at high effort seems significantly better than every other clone.
Buffered controls, different pathfinding AI for each ghost, level transitions, etc.
Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.
pranavmore 22 hours ago [-]
Tried opus 5.5 for the same?
Stitch4223 23 hours ago [-]
Someday we’re gonna miss the times where models create jank.
These results get better and better, but the half-baked pac mans games and pelicans are just funny.
You cannot prompt that level of brokenness. Same for those slop-posters you see everywhere.
Tanjreeve 21 hours ago [-]
I thought this was about playing a pacman game. I feel like we already established that LLMs can create simple games and "local" apps at this point
mikojan 22 hours ago [-]
Clearly the high-tech plagiarism machine will successfully plagiarise one of the most plagiarised games.
gwt4life 21 hours ago [-]
Yeah its not really a good benchmark. Its like having it implement a C compiler after it was trained on the GCC.
tamimio 23 hours ago [-]
It would be great to do the models tests but with different harnesses too to see how they impact on the same models and efforts.
No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.
I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.
One-shotting something like Pac-Man doesn't prove much. At the end of the day, one-shot fidelity is going to scale more or less linearly with model size/world knowledge. Why wouldn't it?
The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.
It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.
They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.
That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.
Successful ecosystems are just pedantically lame enough to keep silly folks from going YOLO, but empowering enough to allow people to still have fun.
>which model has the best training code that was Pac-Man.
You mean which model is more cautious about copyright bleed-through of the $9Tn in FOSS and user code they misappropriated though isomorphic plagiarism. =3
Nodejs wasn't easier for people. JS is a horrible language where simple stuff like comparisons, array access, and member access are broken.
I personally have been working on a Three.js project with Opus 5 and 5.5 that I never would have continued with had I needed to dive into documentation by hand.
Seeing immediate results is incredibly motivating.
Whilst it is the Astra aesthetic, the expectation of any three.js game will be that it is all surface and no detail, whether someone has put the effort in or not.
Mrdoob? If so, can you show a link?
https://x.com/dangreenheck/status/2102137910921240807
Remember when people considered you a genius for prompting with "You are a skilled writer....".
https://news.ycombinator.com/item?id=49882889
Make a pacman game where pacman can always eat ghosts, but ghosts drop pellets.
Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.
Everything is going to go from 0->1 on this benchmark in short order because of that.
It has the advantage of taking good screenshots and letting me know if it’s decent within five seconds of playing.
Buffered controls, different pathfinding AI for each ghost, level transitions, etc.
Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.
These results get better and better, but the half-baked pac mans games and pelicans are just funny.
You cannot prompt that level of brokenness. Same for those slop-posters you see everywhere.