Mate its been a couple of days, I would think you'd be mad to expect perfection from an LLM reverse engineering job straight away.
As a proof of concept however it shows that we are at the stage where any malicious actor has a very low barrier to entry to cause large scale corporate espionage etc.
The endavour doesn't have much value at all except as an experiment how well translating C code to Rust via LLM works (but you don't need an LLM for that, C2Rust was already a thing).
Also, the C version of Quake compiled to WebAssembly and running in browsers is just as "safe".
What this exploits is the mental shortcut of "rust = good", which might be a good thing, as that was always wrong. But now it is being pushed to its breaking point so that that idea will eventually collapse.
Accelerationalism on a micro scale, basically.
AI keeps breaking things that were broken before like this constantly. It's the great cleanup of old bullshit. (Unfortunately through even more bullshit, but at least there is a silver lining)
Oh my god, and the author even supplied a "proof"[0] visual diff harness... that it replicates the original game pixel for pixel.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
If someone recreates Amazon's logistics and distribution systems they could try to compete with Amazon? But they'd also need the connections, distributors, transportation, etc. same with banking software, you need capital to be a bank not just software, and if they have the capital then the technology is working we intended making it easier to make new things and innovate, or at least just compete?
No, I am not talking about "taking over" companies and trying to emulate them and do business yourself. You just need to be able to break trust in the api calls and no one knows if a purchase order or transaction is legitimate.
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
Or they can just take your money and not send out anything.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
I mean some people (me included) have been begging society to pull that plug since over 10 years now.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live.
This was wrong even before LLMs.
The visual diff harness was probably how they got rid of a lot of visual bugs, just tell the LLM to keep going until the pixels match exactly as the verification criteria
tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options
I would never have imagined a phone web browser capable of this in 1996 between playing Quake and testing out the hot new JavaScript powered mouse rollover image effects in Netscape Navigator 2.0
This may sound funny but I feel games would lose a lot of fun if they were all written in rust and had classes of bugs just not available to them. For better or for worse quirks and bugs in games have shaped how people approach games, and also have given games charm for decades.
Fortunately for gamers, Rust doesn't do anything to stop physics engines from going haywire or preventing players from clipping out of bounds. A Mario 64 written in Rust still has parallel universes (well, assuming that you translated the out-of-bounds float-to-short cast as a modulo, which is actually UB in the C implementation).
‘Safe rust’ just being used to mean ‘no unsafe’ is kind of a shallow understanding of the language. You can write code without unsafe that is not really ideal at all eg abusing vector/slice indexing to create a kind of interior mutability that the borrow checker is blind to. Which as a C port I’m going to guess it probably ends up doing
Almost hard to believe that someone using an LLM to mindlessly migrate software from one language to another for no practical reason at all would have a shallow understanding of those languages...
Vibe coders need to stop trying to use the browser as a platform for complex games. It's never going to work out and will always lead to a slow, unplayable, piece of shit. If you want to make a game then just make it a desktop program... There's a reason why literally every modern game is built in this way.
These aren’t JavaScript apps, they compile to WASM which actually has really good performance. Probably not as good as native assembly but for an old game it’s easily more than enough.
I think 'standing on the shoulders of giants' is the phrase for something like this. It's the confluence of browser rendering, WASM, Rust, and LLMs. For me it's less a demo of what AI can do, and more a showcase of the human effort from the past few decades on the parts that needed to fall into place for an LLM to come in at the (relatively speaking) last second and claim a win. Sure, an LLM did the port from C to Rust, but think of all the things needed for it to all work. That's pretty damn amazing, and it wasn't done with AI.
I remember seeing the QuakeC line of code that halved self-damage from rockets. It enabled rocket jumping, but it also made the rocket launcher a much, much better close quarters deathmatch weapon than in Doom, so I thought it was a bit lame.
I don't see any need for this. Rust is brilliant but the Quake C++ code was already more or less bug-free.
Because I did and it's the playable allegory of the cave but in vibecoded rust.
As a proof of concept however it shows that we are at the stage where any malicious actor has a very low barrier to entry to cause large scale corporate espionage etc.
It's easy to create a set that looks as if it was a real city. It's infinitely more hard to create that real city.
But you're absolutely right that a set is all it needs for all sorts of (cyber) attacks.
From what I understand, the adobe guys software is not great.
Consequently it’s missing tons of features.
Also, the C version of Quake compiled to WebAssembly and running in browsers is just as "safe".
It would have been impressive before LLMs, but now? Who cares? Why is this here?
What this exploits is the mental shortcut of "rust = good", which might be a good thing, as that was always wrong. But now it is being pushed to its breaking point so that that idea will eventually collapse.
Accelerationalism on a micro scale, basically.
AI keeps breaking things that were broken before like this constantly. It's the great cleanup of old bullshit. (Unfortunately through even more bullshit, but at least there is a silver lining)
But "this" has been shapeshifting somewhat constantly. We're dynamically crossfading from one dysfunction being blown up to the next.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
[0] https://github.com/terrapapagalli1516/quake-srp/tree/main/or...
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
Currently a plague in some European countries.
It looks like the real site, and you pay twice, in the fake app, and later the police.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live. This was wrong even before LLMs.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
Obviously gameplay bugs are still possible in Rust, but many of them are not.
[0] https://github.com/terrapapagalli1516/quake-srp
[1] https://dreadsweeper.franzai.com/
I like these
I think 'standing on the shoulders of giants' is the phrase for something like this. It's the confluence of browser rendering, WASM, Rust, and LLMs. For me it's less a demo of what AI can do, and more a showcase of the human effort from the past few decades on the parts that needed to fall into place for an LLM to come in at the (relatively speaking) last second and claim a win. Sure, an LLM did the port from C to Rust, but think of all the things needed for it to all work. That's pretty damn amazing, and it wasn't done with AI.
I wholeheartedly bless this slop.
Found the repo in the Reddit post:
https://github.com/terrapapagalli1516/quake-srp
https://www.reddit.com/r/quake/comments/1x1ch14/quake_srp_sl...
Quake in Flash player from 2009 I think
Fucking missed it.
https://play.neverball.org/
.... /s, if it needed to be said
You can play Half Life right now, with one click.