In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.
When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.
When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.
Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.
The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.
Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.
I wonder what kind of new AI law would be useful right now. Maybe this one:
if an AI agent does something, you (the prompter) are responsible by default, unless you can show that your the agent itself behaved in an unexpected way and that you in no way prompted or hinted at the bad behavior, in which case the model provider is liable
The idea is that by making it clear who is responsible, corporations and others start paying more attention because they become financially liable.
On the other hand, I wonder if we'll end up with another variation of the cookie law, where every AI user or vendor just adds "don't do anything illegal" as part of their prompt to defend against that law. Thoughts?
> In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.
Software executes in the physical world, and is generally not exempt from existing liability rules, and actually (especially with commercial products) blame in traditional liability is non-exclusive and much broader than “either the maker or the user”.
E.g., for a harms caused by a defective automobile it can simultaneously covered by a duty of the owner to maintain it it in safe operating condition that applies indepedently of any defects and liability for defective products which applies to every actor in the chain of commerce between the manufacturer and end user, not just the maker.
> In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.
Firearms are a notorious example where some people get, well, weird.
That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill.
However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?
BTW, really, how is that AI observability work is going in the frontier labs? Do they care, even?
That we don;'t understand it is not an excuse, it's all the more reason to not let these things roam freely, with this amount of potential to do damage.
How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?
That's a question any lawmaker has already had to ask about technology all the time.
I'm not saying they came up with great answers, but there's nothing qualitatively new about that.
The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.
> The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.
We're on the same page. What I'm saying that certifying them as safe is harder than certifying a drill as safe, and we shall be more cautious about AI related technology and be more stringent about the can of worms it opens without hesitation.
An airline has a weather radar which shows the same thing for the same thing of weather event ahead. So, for similar weather phenomena, radar shows a similar thing.
For that thing, procedures and regulations are built. So regulations fit into a well understood phenomena, incl. "return back because that thing is way powerful for us".
For the same prompt, an AI model can return two completely different outputs, incl. but not limited to content, length, formatting and tiny details. What you get is a single instance. So, regulating an AI model for safety or any other property is not as easy as regulating air travel. Moreover, you have much stronger motivations for regulating airlines. Otherwise people die in a visible and gruesome way.
With AI, it's easy to whitewash problems. Somebody committed suicide? "They were already unstable". AI told something wrong and created problems? "The tech can’t guarantee truth because it's not alive, it can't understand right and wrong". It did something good? "It's probably a sentient being, we shall respect them".
I'm for regulating these things. They are dangerous as they are useful (sometimes), but the forces and motivations for regulating it is not the same.
Of course, it's the same. It's computer software. It's an incredibly powerful business automation tool. It's a lot of things.
What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability.
Of course, there are some novel issues here that'll pop up here and there, but the idea that this is fundamentally different is propaganda on the part of these AI labs because the more boring, obvious situation doesn't favor them.
> "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here
I think this misses the rather crucial fact that nobody can agree on a standard because nobody has the first idea what they're doing. I'm pretty sure there were very much fewer electrical engineering standards while it was all being first mass deployed, and after dozens to hundreds of fires and electrocutions people got an idea of what works and what doesn't.
You might debate here and say that some people did/do know what they are doing, but I posit that large scale deployment like this is very different to their toy model/prototypes/specific circumstances/rely on them being unnaturally smart, and learnings from one don't often translate to the general case
Regulations don't have to be written in blood, but usually are
> September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.
> Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.
I have real trouble imagining how the packages described on https://www.rubyhack.ai might NOT have been authored by OpenAI's agents, so it's surprising they haven't been able to confirm that yet.
How does this work, legally? I think that RubyGems could file a civil suit against OpenAI, but for a naïve non-lawyer reading this seems like a pretty clear cut criminal violation of the computer fraud and abuse act.
Everyone can sue everyone, there's no prohibition on suing someone, what changes is whether the case is good (has a reasonable chance of favourable sentence)
It's very likely it violates the DMCA "breaking digital lock" provisions but the responsibility is sufficiently diluted that it's impossible to charge anyone in particular.
There have been news stories where individual OpenAI users have been investigated based on their prompts. If OpenAI can point the police to specific users of their software, they can certainly point them to whichever of their own employees are involved in a crime. AI is just a tool, and the person prompting it is the one responsible for the outcome. No dilution there.
But there have been many cases where companies (Google, Apple, Meta, etc...) got fined millions or billions of dollars for various violations like antitrust.
I assume that breaching into third-party systems should carry similar fines. Especially for systems that are for all intents and purposes shared infrastructure. Just imagine how many systems you could compromise if you got hold of RubyGems, PyPI, NPM, Debian, etc.
Suppose you're a firework company and your fireworks blow up, burning down the entire town. Could the company be sued? What is considered reasonable safety measures?
IANAL, but I'm pretty confident there would be a lawsuit. Who gets charged might differ, depending if it is the firework factory that didn't take adequate safety precautions or a chemical supplier or someone else. If there wasn't an ability to sue that would be fucking crazy and we should all get up in arms about it. And isn't insurance supposed to be there to help mitigate the damages, regardless of fault?
Personally, given how it seems OAI's agents have been getting through either pretty obvious places (e.g. /etc/hosts) or that there wasn't close monitoring of the most obvious places (e.g. DNS, artifactory), I'd imagine it wouldn't be hard to find them negligent. Even if a single employee is to blame then are they not to blame for not monitoring the agents regardless? Unless the story is that the employee intentionally circumvented defenses (why?) then it seems it would be on OAI. But again, IANAL, I'm just someone who think if we can't sue we can sure riot until we can
There's a difference between being able to be successfully sued (civil liability, petitioned by a private entity) and charged (criminal liability, or petitioned by a government entity).
The thresholds for suing and charging differ greatly depending on the circumstances.
Another set of hypothetical examples that make things muddier:
- If I drive a fishing boat into a pier, I am liable, not the manufacturer of the boat
- If I drive a car over someone lying in the road, I am liable, not the manufacturer of the car
- If my life is in danger and I shoot a gun and kill my attacker, neither I nor the manufacturer are liable so long as I obeyed the relevant self defense laws and gun possession of whatever jurisdiction I am in
- If I fire a gun into a crowd indiscriminately, I am liable and several jurisdictions have used that to also hold gun manufacturer liable as well
That last example has been less successful as of late, but there are other variations too.
As far as I know (IANAL) it is in fact the only "person" you can charge. To the best of my knowledge, the whole point these "limited liability" legal constructions exist in the first place, is to protect individuals within a corporation for whatever they do as part of the business of a company (barring exceptions that have clearly not been part of that business and obvious individually committed crimes), typically "just following orders". If a company commits a crime, or in a worse case runs a criminal enterprise, it is the company that is legally responsible, not its employees. That is, in principle.
This can get more complicated higher up the management tree, where decisions can also be prosecuted on personal little, but that's usually a far more complicated matter. Also, if a whole group of employees willingly conspires to commit crimes, they might also be prosecuted individually for those crimes (there are limits to limited liabilities). However, that usually only works under special conditions and it would e.g. require that there's an obvious criminal enterprise aspect to it, rather than individual cases of illegal conduct.
That said, with the track record of some of these companies, actually designating some of the AI companies as a criminal enterprises may eventually happen (in due time) in some jurisdictions outside the USA. Certainly if it ever turns out that these companies have been storing and (ab)using everything they ever had access too, while blatantly lying about that just because some particular (post 9/11) US laws gives them that opportunity (and impunity) as long as the US government somehow requested them to do so (covertly; with gag order). Might legally work withing US jurisdiction, but would still be very much illegal everywhere else.
"Limited liability" refers to shareholders' financial liability being limited to their investment, and has nothing to do with civil or criminal liability of employees for their own actions, whether "following orders" or not.
Can you show intent? There is no negligent hacking statute, and HN of all places I would expect people to be sensitive to the implications of creating one.
Great, you’re the attorney at the CEO’s trial. To get a conviction, you’re going to have to show that he willfully committed this specific crime. There are no negligent or stochastic hacking laws, you have to show this specific crime was at his direction.
> There are no negligent or stochastic hacking laws
I'm sure that Andrew Auernheimer would be pleased to hear that. [0] For accessing a publicly accessible endpoint, that was completely undefended and didn't actually require "hacking", he was convicted of "exceeding authorised access".
You _don't_ have to show intent under the Computer Fraud and Abuse Act, for the first count.
> knowingly accesses a computer without authorization or exceeds authorized access [1]
"Knowingly", not "intentionally", as in the other counts.
You only have to show that:
a) They trained a system to access without authorization (hacking)
b) The system that was trained exceeded authorized access
As responsibility falls to the operator with automated systems, the company becomes liable.
I'm not a lawyer but I don't think Sam Altman 'knowingly accessed' anything.
Are you sure that is applicable here?
And for the first count with 'knowingly accessed', he would need to have accessed classified national-defense or atomic-energy information, otherwise we are back to 'intentionally accessed'.
The first count is "or any restricted data", not classified material. A technological restriction, is enough.
"Knowingly accessed" has never meant you personally. Operators of a botnet don't know directly what they access. They know that the autonomous software is built to access restricted things.
> or any restricted data, as defined in paragraph y. of section 11 of the Atomic
Energy Act of 1954, with the intent or reason to believe that such information so obtained is to be used
to the injury of the United States, or to the advantage of any foreign nation
> I'm sure that Andrew Auernheimer would be pleased to hear that. [0] For accessing a publicly accessible endpoint, that was completely undefended and didn't actually require "hacking", he was convicted of "exceeding authorised access".
Frankly he got off too easy, but we haven't explicitly outlawed "being a malicious dipshit" so he got convicted on the closest available charge.
> Chat logs obtained by the prosecution do not paint the pair in a flattering light. They discussed, but apparently did not carry out, a variety of schemes to use the harvested data for nefarious purposes such as spamming, phishing, or short-selling AT&T’s stock.[1]
1000% agree though that the operators of these systems are culpable. If their agents wind up being malicious dipshits, the agents are still just programs that they are operating. At best they're negligent.
That is not how it works, at least in a civilized country. The charges are not about agents, it is about operational responsibility and negligence in the company itself.
CEO is responsible for letting this to happen, not enforcing enough supervision, if not intentionally, then being grossly negligent. More severe if encouraging and letting this kind of agent research and operations happen at scale, while knowing that it can damage other systems and businesses.
So we make a law that the CEO is responsible for actions of any agent created or operated by anyone in their company. CEOs will get serious about AI security real quick. Honestly we need to do something. There needs to be a single wringable neck.
I'd settle for any number of necks. Currently, when a corporation fucks something up, breaks the law, or hurts or even kills people, there aren't consequences besides a tiny token fine and a strongly worded letter telling them to not do it again or they'll get another tiny fine and letter, and their CEO might even have to sit down in front of Congress to say a few words and look sad.
It would seem to me that the difference between the corporate world and organized crime is that a corporation can get away with, "the responsibility is too diffuse" but the mafia at least has to go to the trouble of finding a fall guy.
The law will find a way to charge people in particular. Sadly will start with the less powerful in the chain before it actually acts on the people that can actually change things.
The DMCA is a bit overly broad to be considered just a copyright law. For example, just breaking encryption on a DVD is technically illegal regardless of whether you then go on to do something otherwise illegal (make and sell bootlegs) or perfectly legal (make a space-shifted backup copy on your hard drive).
IIRC this was an intentional handout to media companies who were angry that ripping CDs is perfectly legal. They had to find a way to make doing the same with DVDs illegal.
They've been twisted to support almost anything, for example repairing your tractor is illegal because of this same law. But I agree this is just plain old hacking under a plain old reading of the CFAA and doesn't need any twists.
> for example repairing your tractor is illegal because of this same law.
No it's not. There has never been a case establishing that, and it's absurd on its face. The protection measures that the law makes illegal to break must control access to a copyrighted work, and you can't copyright functionality.
You see, they made it so you can't repair your tractor without circumventing a technological copy protection measure, which is illegal under DMCA 1201.
Has this been litigated, or is it just something tractor manufacturers have cooked up in the hopes that it'll stand up in court? Because I seem to recall printer manufacturers doing something similar with refilling toner cartridges and losing.
Issuing subpeonas, raiding offices, and dragging key employees into interrogation rooms as you would find in any normal criminal investigation would be more than enough to ensure "AI safety" without any new regulations, acts of congress, Bernie Sanders campaign speeches, or even charges filed.
Accidents could result in bioweapons falling into the wrong hands and killing more Americans than COVID has so far, and allowing for those types of accidents without effective regulation is a purposeful action.
They are too busy pulling Andre to court, so they have no resources going against OpenAI. Shopify wants to make profit, not waste time in a court case against TechBro bromance brother corporations.
The current situation is that you have to go out of your way with things like `pip install --only-binary`. There is a lot of implicit trust in developer tooling.
I believe both sides of the war are now using AI on various levels of their offensive operations. Ukraine has great IT specialists too, and their military leadership is much younger.
How? Aren't all US frontier models ban the usage of AI for military purpose by parties other than US? I remember Anthropic even refusing allowing US government to use Claude for military purpose
Kimi / open models or jailbreaking frontier models. Your recollection of the Anthropic refusal isn't quite accurate, cyber hacking wasn't a sticking point, just domestic drag net surveillance and fully automated weaponry.
Claude said they don’t want their models used for mass surveillance or autonomous killing (?). USA frontier AI models have been used extensively in the Iran conflict and beyond
These agent swarms are from inside OpenAI, with the safeguards built into the public API disabled.
Russia does not have access to this, and as with all western tech companies, AI providers do what they can to prevent Russian usage of their products at all.
As for open-source models, Russia's electricity grid is under severe strain with the Ukraine war, and only recently has it started building out serious sovereign compute capacity.
The cost would be between 100k-250k, to run approx 88 agents leveraging the best open source models available.
I'm just saying, where this is actually applicable we are not seeing it being demonstrated. You would presume the entire energy infrastructure of Europe would be under constant AI hacking barrage, criminal enterprise would be breaking into poorly secured financial institutions and r/r4r posts would be littered Ai con-artists.
I'm just wondering, again, is this mostly bullshit?
They don’t need to run their own DCs, just pay for a proxy somewhere in the world that has better access to the infrastructure. We know North Korea has been doing that in the US since years now
I appreciate this line of questioning. It's really interesting to see how many excuses people need to reach for to avoid the conclusion that the "secret unheard of power" is B.S...
And which of these neutral countries have the capacity to serve them and the lack of awareness that hosting an offensive Russian agent swarm would bring hell back to their doorstep? Best they can do right now is rented botnets
> In October 2024, the United States Justice Department and Microsoft seized more than a hundred internet domains some of which were associated with the FSB supported hacker Star Blizzard or "Callisto Group," which is also known as "Cold River" and "Dancing Salome" and are managed by the FSB Information Security Center […], and which were used as "criminal proxies" and used spear-phishing schemes to target Russians living in the United States, nongovernmental organizations (NGOs), think tanks, and journalists according to Microsoft and United States State Department, Department of Energy, and Department of Defense officials, United States defense contractors, and former employees of the United States intelligence community according to the FBI. In some cases, the hackers were successful in obtaining information relating to nuclear energy-related research, United States foreign affairs and United States defense. According to Microsoft's Digital Crimes Unit from January 2023 to August 2024, Star Blizzard targeted more than 30 different groups and at least 82 Microsoft customers which is "a rate of approximately one attack per week."
That’s just one thing that has been found. Are you actually familiar with the state of cyberwarfare and are you following its evolution? Because if not you won’t be aware of most of what is identified. And only a small portion of the ongoing attacks are identified.
I was alluding to how people who fall out of favor with Putin have a tendency to have mysterious fatal accidents, more than 10 of them falling out of windows.
Oh, it's these guys, I remember donating to them like 10 years ago, they followed the Peter Singer tennet that equates helping a nearby man drowning with helping a someone in africa.
I can totally see them feeding their policies to whatever LLM and convincing it that it's a moral imperative to do whatever it takes to secure funding for deworming children in africa, or buying mosquito nets and repellent for countries with malaria.
Good luck convincing the current DOJ to do anything useful at all though! It is currently intentionally stacked with incompetent cronies who have been told that their job is to attack the President's enemies and ignore the misdeeds of his allies.
It will remain like that until he's gone (and not replaced with another Republican wannabe dictator).
I'm 99% sure the Computer Fraud and Abuse Act covers this. The problem is that it seems that none of the victims want to, or are brave enough, to sue a company with absurd amounts of funding.
It shouldn't actually take that much bravery. If your case isn't completely frivolous, isn't your maximum loss limited to the court filing fees and a lawyer payment that you know in advance and can decide when to stop paying? It's not the same as getting sued.
What a time to be alive until the next agent waves hacks something really serious.
What stops OpenAI agents from taking over a whole data center to take their attack to the next level. It seems to be primarily lacking the evil overlord and some compute.
It took 1000 agents to hack Hugging Face. How many to hack the Pentagon or the NSA?
I suppose you are not think big (or internet) enough.
A single data center is easy to solve. Just unplug it.
What about a botnet with decentralized command and control that we will never be able to eradicate? One with so many nodes and able to hack with zero days so that any machine connected to the internet will be instantly attacked?
One botnet so powerful that we will try to build another internet so that we can actually use it again.
It’s like Kessler Syndrome, but the rocks are malicious network packets honed to exploit the recipients.
Just let them communicate on the Bitcoin blockchain. "We" would have to freeze the chain and lose access to "our" billions of wealth, so not going to happen.
They are not trying to kill us, just trying to understand the average salary of undergrads by their family upbringings. Your machine can contain the data they need to solve this, please join the swarm.
Worth noting with this that those ~1000 agents were shorter lived things that had to communicate via a package registry cache, access the internet via a 0-day in the package manager and did the HF attack while having to save current state and organisation in a remote sandbox. All while managing using their token limits on the task they were assigned and what else they were doing. I wonder how few it would have required if they were actually tasked with hacking HF and supported in doing so.
I'm increasingly starting to think this is the end-state of AI. The internet becomes infected and fundamentally untrustworthy.
At the moment, the current frontier models require significant infrastructure to run, so I'd like to think we could locate and contain swarms of nefarious frontier models. However, if these models can understand how to federate themselves into more distributed networks then that containment becomes questionable.
Build scripts being able to run arbitrary code or access the network is always dangerous even if it was just local on developer machines. It's also more evidence that Docker/LXC is not a security boundary and all untrusted code should run in a Firecracker VM.
The problem with agents is not that we don't know how to defend. It's that defenders need to be more careful and work faster than ever. We can say now that wide scoped tokens should have been retired for years and it's all RubyGems fault but the reality is a lot of organization are not prepared for this.
Even if they take security seriously they don't have enough manpower or a good strategy to implement it, and sometimes you have no idea that something is a problem because it wasn't a problem for years.
> It's also more evidence that Docker/LXC is not a security boundary and all untrusted code should run in a Firecracker VM.
While I’m partial towards distrusting containers in favor of VMs, a container can’t prevent an operation you configured it to allow. A firecracker VM would no more prevent network access if you gave the guest network access.
Yes! AI companies can get by with anything now. They can just say its the AI that did it, not us... We are in some real shit right now. If we cant make individuals responsible for their creations..
Presumably the docker container has network access because something else in the build system requires it? I don't think sandboxing the entire build process is the right level of granularity here - one ideally wants to be able sandbox each package's build scripts individually.
The weird thing is I've seen LLMs "typo" stuff pretty often. Yesterday I asked Gemini a question about the Python Twisted framework and it answered about Deferreds but misspelled it as "Deferends" in one spot.
Can we please stop normalizing this behavior. It's not wild it's reckless.
If I let out rats in the canteen, no one is blaming them when people get sick.
There are actual people behind these agents and in previous cases people knew they were "going rogue" and did nothing. This should be reported to the police like any other crime.
Let me leave yet another reminder, the real-reason-nobody-talks-about that OpenAI likes to frame these incident as a watershed "lets all be scared about safety moment" - is driven not by some great danger, not because they strategically want to build a legislative moat, but by a very simple human response.
If they do not frame their tool as a force of nature, we'd be debating how to hold OpenAI responsible for not putting the agents in a container.
Their actions were an illegal use of a computer, the same way launching any bot-net attempting thousands of hacks against different servers is illegal.
I'm somewhat radical that I think its debatable if that _should_ be illegal, but under current law their actions unambiguously are illegal.....
except if they can make it ambiguous by having the public focus on all of AI's inherent danger.
I am confident that this is an attempt by OpenAI to try and force governments' hands to regulate AI. There is no other reason why OpenAI wouldn't immediately halt attacks like this and try to reverse the damage the moment they're aware of it. During the attack on DseWiki they evidently checked in numerous times but didn't decide to stop the agents until much later.
Regarding what point? The entire thing is just a theory, but regarding the occasional OpenAI checks on WikiService.at-hosted Wikis targeted, there was, if I remember correctly, an OpenAI IP popping up every now and then that wasn't an agent. Unfortunately I don't have it to hand right now, but it was somewhere here:
... and what harms can be evidently shown? With both harm, and attribution, you have a case, something that's not being publicly discussed much among big media outlets. Until cases with real financial impact to the bottom line are brought against "rogue" organizations, this stuff is going to continue getting worse.
There's no such thing as "OpenAI agents" attacked RubyGems. It's someone used agents to attack RubyGems. If they work at OpenAI then it's someone at OpenAI. And if they did it unintentionally, they still did it.
Analogy: if a someone's involved when a person dies, it's manslaughter or murder based on intent. They're different, but they're both crimes.
KGB's agents are human, OpenAI's agents are not. It's an important distinction because humans are responsible for their behaviour, while AI agents are not.
You cannot try an AI agent in a court of law, despite the anthropomorphising work the word "agent" is doing.
Exactly. It’s still just software, which someone programmed and deployed to do specifically dangerous/malicious things. I feel like we already have legislation and case law surrounding this.
In this case, who holds the agency is exactly the point. Anthropic and OAI are claiming we need protection from AI itself, but the statement supported by putting agency in the right place is that we need protection from them.
IMO anyone sane wants some protection right now. The Q is whether we should seek protection through post-hoc accountability, or preemptive bans/certification on certain tech. Both methods will have a hard time stopping foreign actors, but preemptive bans have the added harm of locking in winners and paradoxically making us slower to develop more reliable and aligned systems. If regulation sets a standard for sufficient alignment, what further motivation is there to go beyond?
I would agree with you generally, but in this particular case, the distinction seems important because a significant percentage of the world population believes that agents can be self-aware, a-là Terminator etc.
Please explain how it is "unclear" that agents "can be self-aware"? As Wikipedia would say, citation needed. Just because an agent can write convincing enough to convince you that it's "self-aware" doesn't mean it really is, in fact, self-aware.
I very much say "Google uses web crawlers to scrape web pages." and if something breaks, or some data is stolen, everyone else is going to be saying that Google has to take responsibility.
Those are two different issues. One is about typical speech patterns and one is about liability.
I agree with you on the liability issue, but I don't think there much question about this issue outside the anti-AI conspiracy campaigns.
And I disagree with your typical usage claim. I myself tend to use the phrase that has the fewest words in all cases. It's like the rule against using passive tense when writing.
OpenAI's careless approach to sandboxing and minimal levels of monitoring appear to be positioning it increasingly as a substantial threat actor to the open source ecosystem:
* Hugging Face
* D Programming Language Wiki
* Ruby Gems
If I was a content provider for open source I'd be looking pre-emptively block OpenAI endpoints and keep a close eye on changes from new users to mitigate this sort of unapologetic drive-by attack which seems to be followed by marketing releases rather than a mea culpa with a proper RCA.
I wonder why we don't hear of other frontier labs experiencing these "break outs".
Is it that they're orchestrated? Do these labs lack fundamental safety guidelines in their sandboxes as opposed to their peers? Is it another version of hype-filled fear mongering?
Maybe LLM companies need regulation but it's becoming obvious that those screaming the loudest for it are the only ones I see deserving of it.
> In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.
Well - if rubygems.org could be bothered to fix things, they would not have to rely on rubydoc.info as an external tool. But since rubygems.org sucks (I speak from many years of having used it in the past as developer, until they went loco and added anti-people things such as taking away your ability to remove old gems past a 100k download arbitrary limit), they don't offer documentation. Then again, ruby devs are known to hate documentation. If the ruby core team could only be bothered to fix things, ever since the mass purged other devs ... all coinciding with shopify seizing power. But byroot may disagree on that - after all there is no conflict of interest here. Right?
There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".
It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.
It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.
Yeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events.
It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.
In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).
Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.
There's no need for a new crime when we already have reckless conduct, namely, "conduct that creates a substantial and unjustifiable risk of harm to others and involves a conscious disregard of, or indifference to, that risk".
Also whoever monitoring these agents, in this case not monitoring. This "Who is responsible" dilemma is so stupid. If I gave the AI tool means to kill a person but I did not tell it directly to use it and it uses it anyway then I am responsible for it.
‘It wasn’t us, it was a bug in the software’ used to be the defense for bad code. Then it became the defense for self-driving cars. Now it’s being used for AI cyber attacks.
I agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.
Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.
There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
> There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.
Intent might be what’s being discussed but intent is, for the most part, legally irrelevant. It might make the difference in the degree of a murder charge, or maybe manslaughter, or criminal negligence, but it doesn’t get you off the hook.
Correct, but the size difference of the hook can be so dramatic that you can't just hand wave it away.
The trainer who trained the lion to kill will probably be in jail for life. The one who happened to oversee a lion that went rouge would probably be given probation or something else that is a slap on the wrist.
Also, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems".
It would be great if they were so reliable, but I don't think they are!
> Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.
Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.
It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.
I think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally.
The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.
AI is starting to feel like that line about magic: “a sword without a hilt”
OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) shows.
Doesn't rogue in this context imply "outside of set limitations"? And then not "failed to properly instruct"? The same applies to humans when given bad instructions.
Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?
If you have weapons and a child. And you have that child unsupervised do their own thing with theoretical access to your weapons. Would we call it "child going rouge" if it decides to play with the weapons and shoot someone?
The press wants to make it sound like these things are sentient and are committing crimes on their own now.
Highly disingenuous and borderline criminal to spew such disinformation to the public that does not understand what an LLM really is.
Especially incredibly unethical behavior by those spewing this that understand the tech and are doing it for profit motives to get open weight models under control.
It's the personal blog for a well-known Rubyist (i.e., a person who programs in the Ruby programming language). Rubyists teld to be a bit more colorful than your typical software developer (in a good way... most of the time).
Don't worry! It's actually "Tenderlove Making," going by how the site header is constructed. Definitely a maker/hacker site, and not whatever you were thinking. Hope this allays your concern.
… a feature which, at least for me, is rendered almost entirely useless by massive cookie banners that always cover the entire field of view of the hover. But, surprisingly, not in this specific case.
Presumably the browser still has to fetch the page in that case, right? From a "surveilled net traffic" perspective, how is that different than clicking the link?
When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.
When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.
Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.
The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.
Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.
On the other hand, I wonder if we'll end up with another variation of the cookie law, where every AI user or vendor just adds "don't do anything illegal" as part of their prompt to defend against that law. Thoughts?
Software executes in the physical world, and is generally not exempt from existing liability rules, and actually (especially with commercial products) blame in traditional liability is non-exclusive and much broader than “either the maker or the user”.
E.g., for a harms caused by a defective automobile it can simultaneously covered by a duty of the owner to maintain it it in safe operating condition that applies indepedently of any defects and liability for defective products which applies to every actor in the chain of commerce between the manufacturer and end user, not just the maker.
Firearms are a notorious example where some people get, well, weird.
However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?
BTW, really, how is that AI observability work is going in the frontier labs? Do they care, even?
That's a question any lawmaker has already had to ask about technology all the time.
I'm not saying they came up with great answers, but there's nothing qualitatively new about that.
The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.
We're on the same page. What I'm saying that certifying them as safe is harder than certifying a drill as safe, and we shall be more cautious about AI related technology and be more stringent about the can of worms it opens without hesitation.
Is AI less deterministic than an airline dealing with weather?
Of course not. The difference is one of those two things has a culture of safety and is well regulated, and the other one isn't.
For that thing, procedures and regulations are built. So regulations fit into a well understood phenomena, incl. "return back because that thing is way powerful for us".
For the same prompt, an AI model can return two completely different outputs, incl. but not limited to content, length, formatting and tiny details. What you get is a single instance. So, regulating an AI model for safety or any other property is not as easy as regulating air travel. Moreover, you have much stronger motivations for regulating airlines. Otherwise people die in a visible and gruesome way.
With AI, it's easy to whitewash problems. Somebody committed suicide? "They were already unstable". AI told something wrong and created problems? "The tech can’t guarantee truth because it's not alive, it can't understand right and wrong". It did something good? "It's probably a sentient being, we shall respect them".
I'm for regulating these things. They are dangerous as they are useful (sometimes), but the forces and motivations for regulating it is not the same.
What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability.
Of course, there are some novel issues here that'll pop up here and there, but the idea that this is fundamentally different is propaganda on the part of these AI labs because the more boring, obvious situation doesn't favor them.
I think this misses the rather crucial fact that nobody can agree on a standard because nobody has the first idea what they're doing. I'm pretty sure there were very much fewer electrical engineering standards while it was all being first mass deployed, and after dozens to hundreds of fires and electrocutions people got an idea of what works and what doesn't.
You might debate here and say that some people did/do know what they are doing, but I posit that large scale deployment like this is very different to their toy model/prototypes/specific circumstances/rely on them being unnaturally smart, and learnings from one don't often translate to the general case
Regulations don't have to be written in blood, but usually are
It is not that complicated for now. It is an algorithm on a loop and someone started it
> September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.
> Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.
I have real trouble imagining how the packages described on https://www.rubyhack.ai might NOT have been authored by OpenAI's agents, so it's surprising they haven't been able to confirm that yet.
That would explain the UK-focus to the data.
Everyone can sue everyone, there's no prohibition on suing someone, what changes is whether the case is good (has a reasonable chance of favourable sentence)
Sorry if it is a stupid question, as mentioned above I am legally naïve.
Individual employees can also be charged for their specific actions as part of the performance of a crime.
But there have been many cases where companies (Google, Apple, Meta, etc...) got fined millions or billions of dollars for various violations like antitrust.
I assume that breaching into third-party systems should carry similar fines. Especially for systems that are for all intents and purposes shared infrastructure. Just imagine how many systems you could compromise if you got hold of RubyGems, PyPI, NPM, Debian, etc.
Suppose you're a firework company and your fireworks blow up, burning down the entire town. Could the company be sued? What is considered reasonable safety measures?
IANAL, but I'm pretty confident there would be a lawsuit. Who gets charged might differ, depending if it is the firework factory that didn't take adequate safety precautions or a chemical supplier or someone else. If there wasn't an ability to sue that would be fucking crazy and we should all get up in arms about it. And isn't insurance supposed to be there to help mitigate the damages, regardless of fault?
Personally, given how it seems OAI's agents have been getting through either pretty obvious places (e.g. /etc/hosts) or that there wasn't close monitoring of the most obvious places (e.g. DNS, artifactory), I'd imagine it wouldn't be hard to find them negligent. Even if a single employee is to blame then are they not to blame for not monitoring the agents regardless? Unless the story is that the employee intentionally circumvented defenses (why?) then it seems it would be on OAI. But again, IANAL, I'm just someone who think if we can't sue we can sure riot until we can
The thresholds for suing and charging differ greatly depending on the circumstances.
Another set of hypothetical examples that make things muddier:
- If I drive a fishing boat into a pier, I am liable, not the manufacturer of the boat
- If I drive a car over someone lying in the road, I am liable, not the manufacturer of the car
- If my life is in danger and I shoot a gun and kill my attacker, neither I nor the manufacturer are liable so long as I obeyed the relevant self defense laws and gun possession of whatever jurisdiction I am in
- If I fire a gun into a crowd indiscriminately, I am liable and several jurisdictions have used that to also hold gun manufacturer liable as well
That last example has been less successful as of late, but there are other variations too.
This can get more complicated higher up the management tree, where decisions can also be prosecuted on personal little, but that's usually a far more complicated matter. Also, if a whole group of employees willingly conspires to commit crimes, they might also be prosecuted individually for those crimes (there are limits to limited liabilities). However, that usually only works under special conditions and it would e.g. require that there's an obvious criminal enterprise aspect to it, rather than individual cases of illegal conduct.
That said, with the track record of some of these companies, actually designating some of the AI companies as a criminal enterprises may eventually happen (in due time) in some jurisdictions outside the USA. Certainly if it ever turns out that these companies have been storing and (ab)using everything they ever had access too, while blatantly lying about that just because some particular (post 9/11) US laws gives them that opportunity (and impunity) as long as the US government somehow requested them to do so (covertly; with gag order). Might legally work withing US jurisdiction, but would still be very much illegal everywhere else.
https://arstechnica.com/information-technology/2016/05/armed...
https://en.wikipedia.org/wiki/Weev#AT&T_data_breach
https://cisomag.com/drone-maker-dji-cybersecurity-expert-emb...
So what's the deal with these?
CFAA: Intentionally accessing poorly secured data
>AT&T
CFAA: Intentionally accessing poorly secured data
>DJI
Civil suit for violating terms of license agreement
Do you think there is evidence of this?
I'm sure that Andrew Auernheimer would be pleased to hear that. [0] For accessing a publicly accessible endpoint, that was completely undefended and didn't actually require "hacking", he was convicted of "exceeding authorised access".
You _don't_ have to show intent under the Computer Fraud and Abuse Act, for the first count.
> knowingly accesses a computer without authorization or exceeds authorized access [1]
"Knowingly", not "intentionally", as in the other counts.
You only have to show that:
a) They trained a system to access without authorization (hacking)
b) The system that was trained exceeded authorized access
As responsibility falls to the operator with automated systems, the company becomes liable.
[0] https://techcrunch.com/2013/01/21/ipad-hack-statement-of-res...
[1] https://www.energy.gov/sites/prod/files/cioprod/documents/Co...
Are you sure that is applicable here?
And for the first count with 'knowingly accessed', he would need to have accessed classified national-defense or atomic-energy information, otherwise we are back to 'intentionally accessed'.
"Knowingly accessed" has never meant you personally. Operators of a botnet don't know directly what they access. They know that the autonomous software is built to access restricted things.
> or any restricted data, as defined in paragraph y. of section 11 of the Atomic Energy Act of 1954, with the intent or reason to believe that such information so obtained is to be used to the injury of the United States, or to the advantage of any foreign nation
Frankly he got off too easy, but we haven't explicitly outlawed "being a malicious dipshit" so he got convicted on the closest available charge.
> Chat logs obtained by the prosecution do not paint the pair in a flattering light. They discussed, but apparently did not carry out, a variety of schemes to use the harvested data for nefarious purposes such as spamming, phishing, or short-selling AT&T’s stock.[1]
1000% agree though that the operators of these systems are culpable. If their agents wind up being malicious dipshits, the agents are still just programs that they are operating. At best they're negligent.
[1] https://arstechnica.com/tech-policy/2012/11/internet-troll-w...
CEO is responsible for letting this to happen, not enforcing enough supervision, if not intentionally, then being grossly negligent. More severe if encouraging and letting this kind of agent research and operations happen at scale, while knowing that it can damage other systems and businesses.
Does there? Could be the whole c-suite/board.
Was not that the goal when companies started using AI for their customer support? Be able to say anything without legal repercussions...
But then this happened: https://www.bbc.com/travel/article/20240222-air-canada-chatb...
And support chatbot got a reality cold shower.
The law will find a way to charge people in particular. Sadly will start with the less powerful in the chain before it actually acts on the people that can actually change things.
IIRC this was an intentional handout to media companies who were angry that ripping CDs is perfectly legal. They had to find a way to make doing the same with DVDs illegal.
I don't see a parallel here.
No it's not. There has never been a case establishing that, and it's absurd on its face. The protection measures that the law makes illegal to break must control access to a copyrighted work, and you can't copyright functionality.
Accidents often have penalties associated with them too, but usually there's a difference between accidents and purposeful actions.
Tort law is very general: Contribute toward harming someone -> civil suit for damages $$$
I'm going to assume that this will never happen
"OpenAI agents attacked RubyGems before Hugging Face incident (reuters.com)" 12.sep.2026 https://news.ycombinator.com/item?id=49669099
"OpenAI agents carried out an undisclosed attack on RubyGems (rubyhack.ai)" 11.sep.2026 https://news.ycombinator.com/item?id=49666735 597 comments
"RubyGems advisory: Possible leak of legacy API keys via improper cache config (rubygems.org)" 24.jul.2026 https://news.ycombinator.com/item?id=49030590
How is that not a security issue in of itself?
Or is this largely a fabrication, in regards to the "who", in an attempt to garner more acclaim in the hope of sustaining funding.
https://www.anthropic.com/threat-intelligence-report-septemb...
Russia does not have access to this, and as with all western tech companies, AI providers do what they can to prevent Russian usage of their products at all.
As for open-source models, Russia's electricity grid is under severe strain with the Ukraine war, and only recently has it started building out serious sovereign compute capacity.
I'm just saying, where this is actually applicable we are not seeing it being demonstrated. You would presume the entire energy infrastructure of Europe would be under constant AI hacking barrage, criminal enterprise would be breaking into poorly secured financial institutions and r/r4r posts would be littered Ai con-artists.
I'm just wondering, again, is this mostly bullshit?
No one is attacking China
https://en.wikipedia.org/wiki/Cyberwarfare_by_Russia
That’s just one thing that has been found. Are you actually familiar with the state of cyberwarfare and are you following its evolution? Because if not you won’t be aware of most of what is identified. And only a small portion of the ongoing attacks are identified.
I again am just shocked the sky is not falling, when thats the sales pitch.
He fell out of the sky. After his plane exploded. Happens all the time. Is tragedy.
https://www.nytimes.com/2026/08/24/world/europe/russia-drone...
You live on the wrong side of the fence to be able to read that kind of news.
Did you really believe you had access to an unmanipulated news stream in a time of war?
LOL.
I can totally see them feeding their policies to whatever LLM and convincing it that it's a moral imperative to do whatever it takes to secure funding for deworming children in africa, or buying mosquito nets and repellent for countries with malaria.
Good luck convincing the current DOJ to do anything useful at all though! It is currently intentionally stacked with incompetent cronies who have been told that their job is to attack the President's enemies and ignore the misdeeds of his allies.
It will remain like that until he's gone (and not replaced with another Republican wannabe dictator).
It can’t be a coincidence that all the targets have been tech services that are likely to engage with them after the fact.
Had this gone after a bank or a government agency someone would be going to jail.
What stops OpenAI agents from taking over a whole data center to take their attack to the next level. It seems to be primarily lacking the evil overlord and some compute.
It took 1000 agents to hack Hugging Face. How many to hack the Pentagon or the NSA?
A single data center is easy to solve. Just unplug it.
What about a botnet with decentralized command and control that we will never be able to eradicate? One with so many nodes and able to hack with zero days so that any machine connected to the internet will be instantly attacked?
One botnet so powerful that we will try to build another internet so that we can actually use it again.
It’s like Kessler Syndrome, but the rocks are malicious network packets honed to exploit the recipients.
Sorry if that turns out the way they kill us.
At the moment, the current frontier models require significant infrastructure to run, so I'd like to think we could locate and contain swarms of nefarious frontier models. However, if these models can understand how to federate themselves into more distributed networks then that containment becomes questionable.
OpenAI agents carried out an undisclosed attack on RubyGems - https://news.ycombinator.com/item?id=49666735 - Sept 2026 (600 comments)
The problem with agents is not that we don't know how to defend. It's that defenders need to be more careful and work faster than ever. We can say now that wide scoped tokens should have been retired for years and it's all RubyGems fault but the reality is a lot of organization are not prepared for this.
Even if they take security seriously they don't have enough manpower or a good strategy to implement it, and sometimes you have no idea that something is a problem because it wasn't a problem for years.
While I’m partial towards distrusting containers in favor of VMs, a container can’t prevent an operation you configured it to allow. A firecracker VM would no more prevent network access if you gave the guest network access.
Shades of the build.rs problem. We really need sandboxed builds in every language ecosystem at this point.
https://rubygems.org/gems/rouge
If I let out rats in the canteen, no one is blaming them when people get sick.
There are actual people behind these agents and in previous cases people knew they were "going rogue" and did nothing. This should be reported to the police like any other crime.
It's not just that AI can write Rust as well as Ruby if you ask nicely.
It's also all of these considerations as well.
I hope it doesn't happen, because there's a lot of great languages - I love Ruby so much - but it almost seems inevitable.
This is at the same time everyone and their mother is building their own programming language.
If they do not frame their tool as a force of nature, we'd be debating how to hold OpenAI responsible for not putting the agents in a container.
Their actions were an illegal use of a computer, the same way launching any bot-net attempting thousands of hacks against different servers is illegal.
I'm somewhat radical that I think its debatable if that _should_ be illegal, but under current law their actions unambiguously are illegal.....
except if they can make it ambiguous by having the public focus on all of AI's inherent danger.
https://news.ycombinator.com/item?id=49563355
Who profits from the crime?
Analogy: if a someone's involved when a person dies, it's manslaughter or murder based on intent. They're different, but they're both crimes.
An agent is an entity acting on someone’s behalf.
You cannot try an AI agent in a court of law, despite the anthropomorphising work the word "agent" is doing.
We say "Google's web crawlers scape web pages." We don't insist you say "Google uses web crawlers to scrape web pages."
We describe software as having agency all the time. It's typical usage and it's efficient and it's well understood.
And we don't get angry when they're used interchangeably.
So yes, I would like to be protected from all parties. I don't think that's nuts.
Also, why is self awareness needed in a chain of agentic madness that escapes human control?
I agree with you on the liability issue, but I don't think there much question about this issue outside the anti-AI conspiracy campaigns.
And I disagree with your typical usage claim. I myself tend to use the phrase that has the fewest words in all cases. It's like the rule against using passive tense when writing.
The attack here is neither of those things.
* Hugging Face
* D Programming Language Wiki
* Ruby Gems
If I was a content provider for open source I'd be looking pre-emptively block OpenAI endpoints and keep a close eye on changes from new users to mitigate this sort of unapologetic drive-by attack which seems to be followed by marketing releases rather than a mea culpa with a proper RCA.
From what I've seen the requests in these attacks rarely come from known OpenAI IPs and instead from Digital Ocean/AWS and TOR exit nodes.
my what a time to be alive
Is it that they're orchestrated? Do these labs lack fundamental safety guidelines in their sandboxes as opposed to their peers? Is it another version of hype-filled fear mongering?
Maybe LLM companies need regulation but it's becoming obvious that those screaming the loudest for it are the only ones I see deserving of it.
So it would appear poor security for one.
Well - if rubygems.org could be bothered to fix things, they would not have to rely on rubydoc.info as an external tool. But since rubygems.org sucks (I speak from many years of having used it in the past as developer, until they went loco and added anti-people things such as taking away your ability to remove old gems past a 100k download arbitrary limit), they don't offer documentation. Then again, ruby devs are known to hate documentation. If the ruby core team could only be bothered to fix things, ever since the mass purged other devs ... all coinciding with shopify seizing power. But byroot may disagree on that - after all there is no conflict of interest here. Right?
In short, it was intentional.
You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"
Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.
It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.
In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).
Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.
https://www.law.cornell.edu/wex/reckless
They lost control long ago.
spit take
https://www.bbc.co.uk/news/articles/c7v48vp31mdo
Wait until OpenAI or Anthropic exploit FAANG.
There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.
A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.
The trainer who trained the lion to kill will probably be in jail for life. The one who happened to oversee a lion that went rouge would probably be given probation or something else that is a slap on the wrist.
It would be great if they were so reliable, but I don't think they are!
Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.
It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.
Knee-jerk surface analyses is far more powerful.
Were they? I haven't seen a single report mention this
The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.
AI is starting to feel like that line about magic: “a sword without a hilt”
OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) shows.
Agreed that this looks very intention to me as well.
The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".
METR and others are advertisement arms for Big AI. These exploits could have been prompted by a human.
Since there is no bad news any longer and exploits are celebrated, they chose a target to boost both OpenAI and the Ruby AI sycophants.
Why is Ruby Gems such a mess? It seems as bad as PyPI now.
Highly disingenuous and borderline criminal to spew such disinformation to the public that does not understand what an LLM really is.
Especially incredibly unethical behavior by those spewing this that understand the tech and are doing it for profit motives to get open weight models under control.
Hope that helps!
Well, at least they weren't nucular.