The correct title: Researcher brakes one specific stubborn historic enigma message with good help from Astra.
Stubborn for a long time because the message used a completely different key from the rest of that day's traffic. Everyone assumed it shared the daily key. The original transcription had errors. The left rotor turned over at letter 72, which is rare and breaks standard crib attacks.
What is cool, if true, is that it was a 2 day collab between the Leffer and Astra. To me this shows the importance of human in the loop, was still all also showing how immensely power of llm tools. But I think it’s getting a bit silly how much anrticles ignores the driving force (the person) in breakthroughs like this.
I agree with this. I think the researchers who's harnessing the llm's power should be credited more than the model itself. We also need to understand the thought process and the prompts that are given to the model so we can learn and collab to ensure humanity's progress as much as the llm itself.
So the LLM would have done all of this on its own? Why is it ok to acknowledge the human was needed but it’s not a collaboration? Is there a defined percentage of ownership required to make the word collaboration valid?
"However, the most astonishing thing about this break is that the GPT–6 Astra did it entirely on its own. Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page."
I mean... I'm all for collaboration but I think this case is pretty clear, no?
Do we have any kind of transcript as to how the message was cracked, and whether this was cheaper or more expensive than simply Bombe-style trying all the combinations?
The SS-Totenkopf Division was advancing east during the opening weeks of Operation Barbarossa, the German invasion of the Soviet Union. 10 July 1941, the division had just fought its way through the Soviet border defenses around Sebezh. It had moved through Lithuania and Latvia, crossed the Dvina area, and advanced through Dagda toward a place German records called "Rosenow." The division moved out of the Rosenow area around 6 July, fought around Sebezh on 8-9 July, and then continued east/northeast toward Opochka and eventually Porkhov.
interesting, YouTube channel Veritasium just published a video on how Enigma was broken during WWII. at the very end they also give message that has yet to be decoded, although apparently they're different.
WRT the timing, Veritasium maybe looked at the last few weeks and decided there's a fast-closing window in which to report on any famous messages yet to be solved.
If the problem was solved by anyone before or if a similar problem has been solved, then LLMs seem to be able to solve them which is an astonishing piece of technology.
I'm personally not sure if it can come with original thinking and techniques to solve completely novel problems. For that, some imagination and thinking outside the box are required, and I doubt the current architecture can do any of this.
I don't know..I doubt for example it can come up with special relatively if it has knowledge up until 1905.
But I think that is what makes it so good at coding, because coding and building software in general has a lot of repeated problems in different context. Same things for human lives, many think their story or situation are unique, but reality is that the shape of human life has been repeated many many times.
I'd say novel math or scientific theories..let us say we send a robot to space, and we ask to build a colony. A lot of the challenges this robot will face will be novel, it could use inspirations of what humans did on earth, but it might get stuck when things don't work as expected and training data has nothing to build on..but then again we might teach it how to run experiments etc, which would result in data that it can use..but some of those experiments might require imagination or breakthrough in understanding..my guess is that it will get stuck there...
A child walking for the first time. Novelty is easiest agent-relative. A problem is novel for an agent if there is no prior experiences of techniques which work to solve it.
Well it's an old one at this point, but the story around the invention of the 1-time pad is pretty interesting. Long story short, a new engineer who didn't know the problem was considered "impossible" was tasked with sorting it out, and he did. I'm sure I left out a lot of details.
just because humans couldn't solve does not mean the necessary technique were not already discovered...we have agents that don't get tired and has access to all humanity knowledge, the building blocks could be there already..
this is not moving the goalposts, this is try to understand what this tech truly able and not able to do.
It's easier to evaluate certainly. Did it solve the problem? Yes/No
When you design a new thing it will have a dozen drawbacks and a dozen and one benefits. If people then have a bias that everything ai is bad, the signal won't be strong enough to convince.
I believe an LLM can solve pretty much any problem for which we can define a fast iterative loop and for which we have reliable tools to automatically verify the correctness of a result. That’s how they are able to solve some math problems, and how they are able to generate working code. Then it’s a question of how much you have to pay for the model to explore the space of solutions in a reasonable timeframe
In math a lot of the spectacular results have been made by finding a counterexample or worming their way toward a proof that is very well defined.
There's some debate over to what degree current generation AI can be creative at all, or whether it can only crawl around its latent space and explore within constraints. One might ask: were all the solutions to all the math problems AIs have solved already "there" latent in the training data and just hadn't been spotted by humans and put together?
But then... isn't everything latent in our training data if training data is "all observations made about the universe?"
But then... what even is creativity? That gets into philosophy and metaphysics. Creativity, like consciousness and sentience and self-awareness, is not a rigorously well defined concept. So to a degree we don't even know how to ask the question of whether these things are creative.
Yea, when you start digging into intelligence/learning/creativity/language you realize that you're talking about more information on these topics than a human could ever learn in their lifetime, and you learn just how much we don't know about ourselves and these processes.
Whenever I hear "AI can never" I know I can disregard them as an unserious person when it comes to anything around AI, learning, or philosophy.
OpenAI solving all these math/etc problems (reportedly 100 coming) is both impressive and insanely unimpressive. Unimpressive because to me it sorta signals that OpenAI has nothing better to be working on than obscure mathematical curios?
It seems like OpenAI's takeaway after Sora is to not stretch themselves thin and focus on what's important.
Based on what's publically available, they're focusing on hacking uncontesting orgs using misconfigured sandboxes and math puzzles.
Your statement is essentially unfalsifiable. We can't possibly discuss whatever Sam Altman is doing in his private office room, nor should we assume OpenAI is working on anything other than what has some public traces.
Attempting to solve & solving open math problems probably is a good benchmark for comparing models and gauging model progression. A lot of useful info is obtained like time needed to solve the problems, identifying when not to chase dead ends, thought processes & logic steps, etc.
Disagree. We are at the point where coming up with good evals for these models is extremely difficult. Solving unsolved math problems is a valid way of evaluating model progress and somewhat necessary to understand how far the current crop of models can go.
Even the Millennium Prize Problems have, in a way, become benchmarks for model companies to prove themselves. The smartest individuals among humans are becoming replaceable. Intelligence has become a product you can quantify and buy with electricity. That feels awful.
> Even the Millennium Prize Problems have, in a way, become benchmarks for model companies to prove themselves
Well, let them have these. They'll play around with open problems which generate media hype and then they might run out and move on to something else, because "AI came up with a problem and solved it in 3 days" won't have the same effects as "AI solved a problem in 3 days that humans couldn't solve in 100 years".
I think that's right. Expert experience used to be almost the most precious and valuable part of the computer field, but today that experience has been "distilled" into SOTA models.
Put it bluntly: the weavers who could be replaced by the spinning jenny were clearly doing repetitive labor. People writing code and maintaining project pipelines a few years ago relied heavily on experience, but in a sense that was also "repetitive labor." Replacing repetitive labor and freeing up productivity is of course progress.
But reform always has its victims. Like the textile workers who starved in the streets centuries ago, and me, kicked to death in the street by AI today...
All technology on the tech tree which requires intelligence to unlock will soon be available to humanity – mind control, population exterminating bioweapons, new ultra destructive kinetic weaponry, perhaps even a cure for cancer.
That won’t happen. But also, whatever benefits are unlocked will be owned mostly by a small group of individuals, definitely not available to humanity as a whole.
Solving obscure puzzle samples that approximately ~0 humans on Earth ever attempted to solve, mostly by pattern matching known solutions to similar puzzles, is not intelligence. DeepBlue has been outperforming the best humans at a specific puzzle-like task since the last century.
Do any of the people proclaiming this shit actually use these models? No matter how many headlines are coming out, every day I deal with reams of the most horrific code I've ever seen technically compile, with routine mistakes that any human would get fired for if they made.
But humans have been confusing pattern matching against known solutions for intelligence for a hundred years!
Seriously though, it ends up looking like that. To take a stupid example a couple of weeks ago I asked an agent to look at porting my hand written WebGL renderer (+ shaders etc) to WebGPU. It estimated a human would take 6-10 weeks, and I would agree. (Which is why I hadn't done it). 24 hours later it was deployed and live. This is classic tedious, difficult, low level if quasi mechanical work (rather like cracking an enigma message), and LLMs absolutely fly through it.
You do understand this is intentionally trained into recent models for marketing purposes? "Wow, it saved me months of work in a day! This is the most amazing technology ever!!!!"... is what it intends to evoke by underpromising and overdelivering. I routinely have it helpfully suggest it will take something like "three engineer-months" to do something I do by hand without any LLM assistance in a day. The estimates may be accurate if you have literally never touched a computer in your life before and are starting to learn from there.
In the games industry I was tech lead of teams of hundreds of devs and had to deal with their estimates of this sort on a daily basis. 6-10 weeks for a total renderer rewrite is on the low end.
Even those of us that are pro AI need to acknowledge this is the current reality.
The smartest humans now need to move to being less concerned about status games among humans and more with how to provide value to a mix of intelligent machines and humans. i.e. if you're starting an SaaS in 2026 you better be assuming half your revenue is going to come from machines acting by themselves.
"GPT–6 Astra mentions a private collection, but it is not clear what this is"!, my spidey senses makes me think it hacked something? Or am I misreading this?
Phew, never have I seen in my life the goalposts move so fast.
It seems like even yesterday that the threshold for impressing someone is that the machine would have to be good at pretending to be a person. Now the threshold is that they have to be able to invent special relativity.
I didn't know this - from wikipedia page on Enigma[1]:
> Despite the seeming difficulty in decrypting its messages, Enigma contained a number of design issues that left patterns in the cyphertext. Poland first cracked the machine as early as December 1932 and was able to read messages prior to and into the war. Poland's sharing of their achievements enabled the Allies to exploit Enigma-enciphered messages as a major source of intelligence.
Ok interesting, so why do people talk about Turing in this connection then?
> Turing devised techniques for speeding the breaking of German ciphers, including improvements to the pre-war Polish bomba method, an electromechanical machine that could find settings for the Enigma machine
Ok so Turing just improved an existing method. Without being an actual expert it's impossible to know how much credit he actually deserves.
Two more references: the Polish method was called "Bomba" [3] invented by Marian Rejewski [4]
Yea, it's kind of odd how "sour grapes" people can be when something stops being as special as they thought it was.
Take someone from a few hundred years ago and drop them into today, and if they don't go catatonic and die, then they'd tell you that we created magic. "Wow, you live in a world of magic and all you do is bitch about it".
>"You're flying! You're sitting in a chair, in the sky!"
Actuallllly these hypothetical people from a few hundred years ago hypothetically said you are an awful person for misrepresenting them, stop putting words in their mouth, thank you!
A person inventing something out of dumb luck is interesting.
A machine running a loop through an expensive LLM for an undisclosed amount of time, which cost an undisclosed amount of money, which was told to keep looping into a solution was found, for a problem that nobody was very concerned about... That just seems like PR, and it's not so interesting.
This dichotomy is counter-productive. The valuations floating around are insane, and the claim that some software can replace every single laborer is one step removed from fiction. At the same time, this stuff is clearly going to change how to world works in countless, deep ways. But the idea that some fancy autocomplete can replace humans ignores the reality of humanity and the fancy autocomplete.
It’s another dot com bubble, not a crypto bubble. Trillion dollar valuations burst once you leave lesswrong.
Edit: very fancy autocomplete. I know what these things are capable of. It’s still not “intelligence”, for X definition of intelligence. And it certainly benefits from having obscene amounts of compete thrown at it. It is awesomely impressive synthesis of data, yet it’s clearly still that.
i dont think humans are going to be replaced either, but calling the thing that solves millennium problems and is currently changing multiple industries entirely a "fancy autocomplete" makes it harder to take any point you are making seriously.
Stubborn for a long time because the message used a completely different key from the rest of that day's traffic. Everyone assumed it shared the daily key. The original transcription had errors. The left rotor turned over at letter 72, which is rare and breaks standard crib attacks.
What is cool, if true, is that it was a 2 day collab between the Leffer and Astra. To me this shows the importance of human in the loop, was still all also showing how immensely power of llm tools. But I think it’s getting a bit silly how much anrticles ignores the driving force (the person) in breakthroughs like this.
"However, the most astonishing thing about this break is that the GPT–6 Astra did it entirely on its own. Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page."
I mean... I'm all for collaboration but I think this case is pretty clear, no?
Edit: Found it from here: https://mvueh-enigma-solved.carterl.chatgpt.site/
(Is it possible that this is a misreported detail? It feels like a singular ROSENOW would be an equally effective crib)
https://www.youtube.com/watch?v=JsBZOcqZerk
And then I come to hackernews and well, not quite, but I'm sure that one will be done shortly too.
I'm personally not sure if it can come with original thinking and techniques to solve completely novel problems. For that, some imagination and thinking outside the box are required, and I doubt the current architecture can do any of this.
But I think that is what makes it so good at coding, because coding and building software in general has a lot of repeated problems in different context. Same things for human lives, many think their story or situation are unique, but reality is that the shape of human life has been repeated many many times.
I'd say novel math or scientific theories..let us say we send a robot to space, and we ask to build a colony. A lot of the challenges this robot will face will be novel, it could use inspirations of what humans did on earth, but it might get stuck when things don't work as expected and training data has nothing to build on..but then again we might teach it how to run experiments etc, which would result in data that it can use..but some of those experiments might require imagination or breakthrough in understanding..my guess is that it will get stuck there...
Um, I'm not sure if you've noticed, but we have bipedal robots that walk and run rather well now.
https://docs.nvidia.com/learning/physical-ai/index.html
https://www.nvidia.com/en-us/use-cases/robot-learning/
"Prove or disprove string theory in pure mathematics, reply in Caveman speech"
Reminds of what Einstein said, imagination is more important than knowledge..might be he deepest insight ever.
this is not moving the goalposts, this is try to understand what this tech truly able and not able to do.
When you design a new thing it will have a dozen drawbacks and a dozen and one benefits. If people then have a bias that everything ai is bad, the signal won't be strong enough to convince.
There's some debate over to what degree current generation AI can be creative at all, or whether it can only crawl around its latent space and explore within constraints. One might ask: were all the solutions to all the math problems AIs have solved already "there" latent in the training data and just hadn't been spotted by humans and put together?
But then... isn't everything latent in our training data if training data is "all observations made about the universe?"
But then... what even is creativity? That gets into philosophy and metaphysics. Creativity, like consciousness and sentience and self-awareness, is not a rigorously well defined concept. So to a degree we don't even know how to ask the question of whether these things are creative.
This gets interesting.
Whenever I hear "AI can never" I know I can disregard them as an unserious person when it comes to anything around AI, learning, or philosophy.
I think they should now focus on robotics, so it can do my dishes while I work on fun math games.
why would they compete with Nvidia on that?
Based on what's publically available, they're focusing on hacking uncontesting orgs using misconfigured sandboxes and math puzzles.
Your statement is essentially unfalsifiable. We can't possibly discuss whatever Sam Altman is doing in his private office room, nor should we assume OpenAI is working on anything other than what has some public traces.
I think the publicity is a nice to have. They need models like this for in-house use.
Technically, this isn’t OpenAI directly.
There are many things I want to do - that would require me to hire a team of 50 people.
I don’t want to do that nor can I afford to. Can OAI focus on enabling me to do this? I don’t care about this other stuff.
Just like many things in life - if it doesn’t show up in the economy it’s irrelevant.
Well, let them have these. They'll play around with open problems which generate media hype and then they might run out and move on to something else, because "AI came up with a problem and solved it in 3 days" won't have the same effects as "AI solved a problem in 3 days that humans couldn't solve in 100 years".
Only if you subscribe to the "humans are special" rhetoric, in which case I'm - maybe - sorry to say the feeling will only intensify.
I think the more accurate description of what's happening is that access to expertise is becoming commodified.
John Henry.
There's a reason we made folklore about when the machines came for the strength of men, and now 150 years later it comes for our minds.
There used to be days when women would make blankets, when men would make chairs, when children would make brooms...
But PROGRESS I tell you!
But reform always has its victims. Like the textile workers who starved in the streets centuries ago, and me, kicked to death in the street by AI today...
So not any different from right now.
>definitely not available to humanity as a whole.
[taps on forehead meme]
The whole of humanity can have it available, if there is a whole lot less humanity.
Way, way more concentration of wealth and power
Why?
> But also, whatever benefits are unlocked will be owned mostly by a small group of individuals, definitely not available to humanity as a whole.
It's unclear to me if this is the good scenario or bad scenario. This would be good in your view right? At least I hope you're right.
Do any of the people proclaiming this shit actually use these models? No matter how many headlines are coming out, every day I deal with reams of the most horrific code I've ever seen technically compile, with routine mistakes that any human would get fired for if they made.
Seriously though, it ends up looking like that. To take a stupid example a couple of weeks ago I asked an agent to look at porting my hand written WebGL renderer (+ shaders etc) to WebGPU. It estimated a human would take 6-10 weeks, and I would agree. (Which is why I hadn't done it). 24 hours later it was deployed and live. This is classic tedious, difficult, low level if quasi mechanical work (rather like cracking an enigma message), and LLMs absolutely fly through it.
You do understand this is intentionally trained into recent models for marketing purposes? "Wow, it saved me months of work in a day! This is the most amazing technology ever!!!!"... is what it intends to evoke by underpromising and overdelivering. I routinely have it helpfully suggest it will take something like "three engineer-months" to do something I do by hand without any LLM assistance in a day. The estimates may be accurate if you have literally never touched a computer in your life before and are starting to learn from there.
The smartest humans now need to move to being less concerned about status games among humans and more with how to provide value to a mix of intelligent machines and humans. i.e. if you're starting an SaaS in 2026 you better be assuming half your revenue is going to come from machines acting by themselves.
5.6 Sol is much more capable.
Being trained on a mountain of stolen material for guessing cribs helps. Up to now no group had that much funding to steal. Congratulations.
Then see if it can come up with E=mc^2
It seems like even yesterday that the threshold for impressing someone is that the machine would have to be good at pretending to be a person. Now the threshold is that they have to be able to invent special relativity.
> Despite the seeming difficulty in decrypting its messages, Enigma contained a number of design issues that left patterns in the cyphertext. Poland first cracked the machine as early as December 1932 and was able to read messages prior to and into the war. Poland's sharing of their achievements enabled the Allies to exploit Enigma-enciphered messages as a major source of intelligence.
Ok interesting, so why do people talk about Turing in this connection then?
> Turing devised techniques for speeding the breaking of German ciphers, including improvements to the pre-war Polish bomba method, an electromechanical machine that could find settings for the Enigma machine
Ok so Turing just improved an existing method. Without being an actual expert it's impossible to know how much credit he actually deserves.
Two more references: the Polish method was called "Bomba" [3] invented by Marian Rejewski [4]
[1] https://en.wikipedia.org/wiki/Enigma_machine
[2] https://en.wikipedia.org/wiki/Alan_Turing
[3] https://en.wikipedia.org/wiki/Bomba_(cryptography)
[4] https://en.wikipedia.org/wiki/Marian_Rejewski
GPT-6 Astra Solves a WWI German Radio Cipher
https://news.ycombinator.com/item?id=49763987
This for me, isn't interesting, it required no skill, no imagination, in fact it seemed like it happened by dumb luck.
So we have entered an age where an army of know-nothings direct models to old forgotten tasks so they can get 15 minutes of un-deserved attention?
[1] https://www.prinzai.com/p/gpt-6-astra-solves-a-wwi-german-ra...
I agree with you that these are not "trilling" discovery but they can be worth something anyway.
Like you, I do not like this news also because I think they will be used to just "push" the next two IPOs (Anthropic, OpenAI).
So like half of all useful human inventions are to you not interesting just because it happened by dumb luck?
Take someone from a few hundred years ago and drop them into today, and if they don't go catatonic and die, then they'd tell you that we created magic. "Wow, you live in a world of magic and all you do is bitch about it".
>"You're flying! You're sitting in a chair, in the sky!"
A machine running a loop through an expensive LLM for an undisclosed amount of time, which cost an undisclosed amount of money, which was told to keep looping into a solution was found, for a problem that nobody was very concerned about... That just seems like PR, and it's not so interesting.
They don't care about the advancements themselves, only that the advancements are some sort of cheating that shouldn't "count".
Humanity is profoundly unsettled by AI and is responding with avoidance and denial.
It’s another dot com bubble, not a crypto bubble. Trillion dollar valuations burst once you leave lesswrong.
Edit: very fancy autocomplete. I know what these things are capable of. It’s still not “intelligence”, for X definition of intelligence. And it certainly benefits from having obscene amounts of compete thrown at it. It is awesomely impressive synthesis of data, yet it’s clearly still that.
Complete this sentence: "You can solve world hunger by..."
Intelligence is required to finish sentences. It's not just a Markov chain.
The internet was pretty important afterall even with a bubble.
I'm pretty AGI-pilled, and I feel perfectly emotionally prepared for if AI stayed at its current capabilities and the S&P dropped 25%.