Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

My goal post for "AI will definitely replace most SWEs" was to reproduce a particular 90s programming game one shot and then add multiplayer support with minimal prompting.

Opus 4.5 hit that point in November.



I tried this a while ago, haven’t tried again recently. The models were producing code that was clearly lifted from stuff in their training data, and what I ended up with was a fairly decent game in html and js after a bit of tidy up, though it felt like several code paradigms smooshed together rather than a coherent whole, but it mostly worked. Not something I’d want to maintain but it was impressive at the time.

They were able to one-shot famous games (like asteroid or pong), I suspect because they had been trained on multiple versions of that game. So like producing Harry Potter, with the right prompt it was able to produce a license stripped version of code it had seen. I tried another arcade game like frogger and it failed really badly and took a lot longer, never got it working.

The whole exercise left me feeling they have a long way to go, I don’t see how anyone could think they would replace SWE unless they didn’t look at the code produced, even now.


Out of curiosity - what harness did you use, and what model? And how are you prompting? In my mind prompting like:

“You’re going to make frogger in javascript. I want a complete clone of functionality for level 1, with amazing 80s era pixel art sprites. I’m super lazy, so you’re going to have to test everything, right from the start. Pick a test harness, write the tests, including tests for having amazing graphics, gameplay, input, UI, sounds, etc, and write a full workplan, then work through that workplan, in parallel where you can. The workplan should emphasize getting a stripped down version up immediately and have workstreams for all the major requirements after that. Add a final test that assesses how fun the game is by reviewing a real video of a test run. Loop on that final test until you can’t improve things any more.”

Should produce something playable with no further input. As you say, I’m not sure it would produce a codebase we’d want to look at or work on. But, I’d be surprised if this weren’t successful.


Sure give it a go, perhaps it will work better now with frontier models, I haven't tried it in a while (this was a year ago, things have improved since then). I'm not sure what tests for having amazing graphics, gameplay, input, UI, sounds, etc would look like, but it would be interesting to see the results!


okay hold my beer. both claude and codex running now.

EDIT: both agents took about 20 minutes. I used that exact prompt in a clean directory for each, and then said "deploy to netlify" - so a total of two prompts.

Codex: https://astounding-bavarois-27b5a2.netlify.app

Claude: http://strong-hotteok-91dfb0.netlify.app

Netlify is having trouble claiming the Claude project, so if you need a password it's "My-Drop-Site"

FYI, Claude rated itself 7.7/10 for fun, and Codex 98/100 during the fun test loop. As you'll see if you poke at them, Claude needs a physics bug fix round. But I think these both did about what I would have expected.


Nice, very retro (looking at the codex one)!

Claude one doesn't really work (collision detection was the problem I had before too), but fairly close.

Yes when I tried previously I had a few gameplay issues in frogger and I couldn't manage to one-shot this sort of thing at the time (a year ago), so last year definitely saw some good progress at this sort of thing. The asteroids game I was very happy with though, had a very cool retro feel and was wireframe only. Wasn't so keen on the code produced as it had a patchwork feel to it.


To your point, I didn't even look at the code.. :) Okay, I looked at the codex code. it's super reasonable -- separation of concerns, operating on a state model, it's not over designed. I did not hate it. I also noted that codex put in a CRT simulator loop which is a nice touch.

I think a year ago this would have taken a lot of back and forth and arguing; to me that's kind of the point of Simon's article -- a lot more just 'works' now.


Sorry I meant the code a year ago - it took a bit more hand-holding at that point and it was a mishmash of different things, but I feel it’s just slightly easier now - still similar. Haven’t looked into this one just had a quick play. Thanks for trying it out!

I think his article is for the last 6 months - my feeling is progress with LLMs has stalled recently and generated code still has problems with accuracy and coherence and subtle bugs, but everyone has a different experience.


I agree with that. Right now, you choose:

Subtle bugs in understanding the spec but strong arch and coding (codex)

Or

Subtle bugs in implementation but good understanding of the spec (claude).


Frogger is kind of too well known such that there is ample training data for building that specific game.

The game I was thinking of is relatively obscure -> Panel de Pon


Yes I was surprised at the time that it failed so badly at Frogger, I think from memory it was colission detection it just couldn't get right, plus the positioning of various game elements as it has quite a lot going on (the examples above still have some problems with these things). I thought there would be open source examples out there in js/html but perhaps not so much for frogger.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: