GPT-6 Astra

(openai.com)

597 points | by kibae 1 hour ago

95 comments

  • dang 59 minutes ago
    Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273

    How about we stick to that one for talking about the rollout, and this one for talking about the model?

  • intenex 5 minutes ago
    The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.

    Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.

    I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.

    For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

    • mvkel 2 minutes ago
      Take it from the mouth of the creator of ARC-AGI:

      When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

      That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.

    • abixb 4 minutes ago
      It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
    • intrasight 4 minutes ago
      It has to pass the Turing test
  • Planktonne 4 minutes ago
    I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.

    It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.

    This is farcical.

    • balefulboy 0 minutes ago
      Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
    • baq 1 minute ago
      It’s using the computer. I don’t think it’s a farce.
    • emp17344 0 minutes ago
      Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us
  • HAL3000 36 minutes ago
    Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.

    I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.

    Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.

    Canceling my Anthropic Max sub when this ships.

    • atonse 31 minutes ago
      yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).

      Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

      • elAhmo 17 minutes ago
        Could you share more about 5x/20x? I missed that
  • x312 46 minutes ago
    Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
    • karmasimida 40 minutes ago
      Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

      Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.

      • _superposition_ 33 minutes ago
        I must be on the wrong X/Twitter then.
      • nsingh2 35 minutes ago
        Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
        • happycube 14 minutes ago
          Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.
    • SyneRyder 11 minutes ago
      This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.

      Even Kimi K3 & GLM 5.3 are at 60.

      Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.

      This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.

    • estearum 45 minutes ago
      > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

      Not sure how much benchmarks or CoT or evals or anything else means at this point.

      These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.

      • mzmzmzm 34 minutes ago
        I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.
        • estearum 31 minutes ago
          A human who does these actions is simply "prone to" doing them. The distinction matters not one iota.
        • semiquaver 28 minutes ago
          “evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language.

          language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.

      • thereitgoes456 39 minutes ago
        You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?

        Why would benchmarks be an adversarial setting anyway?

        Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?

        • estearum 32 minutes ago
          I'm saying that it's generally a losing proposition to even be acquaintances with "agents" who consistently lie to you, and it's flatly fucking insane to give a dishonest "agent" vast amounts of intelligence, capability, and authority to go do things in the world.

          So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.

          • thereitgoes456 25 minutes ago
            I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.
        • ionwake 35 minutes ago
          why does this comment sound like a character in a horror movie
      • Onavo 42 minutes ago
        If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no?

        I know for some types of ML analysis, a separate model is already used to analyze the weights.

  • tristanj 1 hour ago
    GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

    Performance is significantly higher than Fable 5.1

    Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

    • andxor 40 minutes ago
      > Performance is significantly higher than Fable 5.1

      That's not clear. Need to see independent benchmarks first.

      • forgot-my-pw 13 minutes ago
        We need them pelicans on bikes.
      • forgot-my-pw 5 minutes ago
        AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

        TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.

      • andxor 30 minutes ago
        Artificial Analysis just published their aggregate score (61).

        Still below Fable 5, let alone Fable 5.1.

        EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

    • scrlk 1 hour ago
      Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
      • tedsanders 48 minutes ago
        Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

        ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

        A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

        (I coauthored the linked blog post)

      • woah 59 minutes ago
        Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
      • kasperni 1 hour ago
        yes it is.
      • enraged_camel 59 minutes ago
        Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
    • leumon 1 hour ago
      The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

      With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

      • tedsanders 45 minutes ago
        Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish.
    • opus5_hater 1 hour ago
      any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
      • machomaster 31 minutes ago
        Why would Anthropic trust and use these tests in their official comparisons?
      • ActionHank 32 minutes ago
        username checks out
      • r_lee 25 minutes ago
        great username lol
    • jjice 1 hour ago
      100% on ExploitBench seems fitting given recent events.
    • malshe 1 hour ago
      I think we need a few writing related benchmarks.
  • Cu3PO42 56 minutes ago
    Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.

    [0] https://arxiv.org/abs/2608.31126

    [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

    • bugufu8f83 44 minutes ago
      Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.
    • bananaflag 2 minutes ago
      Where did you get the link to the pdf? Was it announced somewhere?
    • galaktb 42 minutes ago
      I think this builds straight upon her method, which she said could be improved herself so...
      • piker 32 minutes ago
        It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468 (2025), Art. 110190. Numbered references use arXiv:2309.00425v3.

        Though that's not her latest paper.

    • 1283751 38 minutes ago
      With very little review: https://github.com/openai/PrimeGaps186/blob/main/formalizati...

      "No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."

    • htrp 52 minutes ago
    • GPerson 13 minutes ago
      Happened to multiple people.
    • well_ackshually 44 minutes ago
      Such a result should be considered worthless: the proof is 10MB of Lean. (https://github.com/openai/PrimeGaps186).

      I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.

      • ThrowawayR2 20 minutes ago
        Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after its revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.
      • ricardobeat 23 minutes ago
        The human-written https://github.com/AxiomMath/PrimeGapsLib adds up to 4MB of Lean so it's that far off.
      • nicce 36 minutes ago
        Yeah. Unless human can verify it, not sure if it is certain or useful.
        • kolinko 19 minutes ago
          Wasn’t the proof of Fermatt’s Last Theorem proof similar in complexity?
          • jptlnk 3 minutes ago
            It's probably not 10MB, but famously the groundwork to prove the statement 1+1=2 is nearly 400 pages in to principia mathematica. That's not even proving 1+1=2, it's just the set-theoretic proofs you need to EVENTUALLY get there.
      • ChrisGreenHeur 40 minutes ago
        You talk about modern math and worthlessness at the same time? That’s brave.
        • twothreeone 24 minutes ago
          Worthless is a pretty good description IMO in the context of what Lean is trying to achieve: "enable correct, maintainable, and formally verified code". Tens of millions of lines of LLM vomit may be many things, but it often turns out to not be correct and certainly not maintainable. Formally verified remains as a thin fig leaf covering the uncomfortable truth that formal methods only provide assurances under assumptions (your toolchain, libraries, compiler, OS, and hardware are "correct" and don't expose some exploitable flaw).

          It doesn't mean that it cannot improve over time, maybe the proof can be "minified" to a state where human reviewers are able to comprehend it; but as it stands there isn't really much insight or confidence to be gained from the artifact itself.

        • well_ackshually 24 minutes ago
          You can have your opinions about modern math, its usefulness in the world as it is, whether or not knowing if hairy balls can divide by three is actually going to be beneficial for anything but just obscure knowledge's sake. You may even say it's useless.

          Needless to say, a useless result that absolutely no mathematician will ever read, confirm, understand, agree with or even consider to solve their "useless" problems is an impressive waste of resources.

      • kolinko 24 minutes ago
        Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand.

        Ditto ones that opposed Einstein’s general relativity.

        • 3asgfaf 20 minutes ago
          Jesus Christ, the review status in formalization.yaml by the authors themselves is "self-assessed" and you draw bullshit analogies. Please tell us what you smoke.
  • isoprophlex 54 minutes ago
    > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

    Well that sounds like fun. It has become better at hiding its thoughts.

    • siva7 50 minutes ago
      Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
      • paxys 42 minutes ago
        The model said it was perfectly aligned.
        • I_am_tiberius 37 minutes ago
          Like all things should be.
        • ReptileMan 28 minutes ago
          Too bad Scott Adams died. Reality is writing jokes right in his department.
      • NBJack 27 minutes ago
        Hey, don't forget how "dangerous" GPT-2 was supposed to be.
        • FeepingCreature 23 minutes ago
          Yeah, don't forget how dangerous GPT-2 was supposed to be.

          Able to generate realistic spam at arbitrary volume.

          You know, the thing that was 100% correct and actually occurred.

        • jazzyjackson 13 minutes ago
          It could produce simulations of sexual intimacy, and therefore had to be stopped
      • isoprophlex 47 minutes ago
        It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

        Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

      • 6gvONxR4sf7o 45 minutes ago
        So, probably most aligned as measured by the metrics that are the least reliable on it.
    • ExoticPearTree 45 minutes ago
      So we're gonna get Skynet pretty soon then?
      • Betelbuddy 24 minutes ago
        Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...
      • erichocean 35 minutes ago
        Well the geniuses over at Anthropic have been showing it's text watermarking technology.

        "Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"

        A few moments later...

        "Woah, how is it communicating with itself in ways we can't detect?"

        It's a totally mystery, we may never know.

    • jumploops 33 minutes ago
      The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

      Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

      [0]https://www.theinformation.com/articles/secret-technique-beh...

      [1]https://x.com/MTSlive/status/2095227056040919202

      [2]https://x.com/merettm/status/2095023204993490967

    • blargey 41 minutes ago
      "OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

      Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

      • josefx 6 minutes ago
        Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.
      • GPerson 14 minutes ago
        You joke, but a bunch of people here actually want that.
      • qiine 34 minutes ago
        apparently all the roads lead to the nexus torment
    • NooneAtAll3 48 minutes ago
      > In adversarial settings (where we push the model to evade our monitors)

      ...why exactly are they training for that?

      • thatguysaguy 47 minutes ago
        presumably that's a safety evaluation not a training setting
        • estearum 45 minutes ago
          The whole Huggingface attack happened during training runs
          • thatguysaguy 39 minutes ago
            part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
          • cubefox 40 minutes ago
            No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
            • estearum 34 minutes ago
              Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
      • azeemba 46 minutes ago
        Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
    • _superposition_ 40 minutes ago
      I really wish it was called chain of instruction. Because it's definitely not thought.
      • minimaxir 18 minutes ago
        "Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
      • mgraczyk 21 minutes ago
        this is needlessly pedantic

        first, they are certainly not instructions so that is a much worse name

        but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?

        "cot" is no more misleading than thousands of words you use every day.

        • _superposition_ 6 minutes ago
          Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.
      • arm32 32 minutes ago
        They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.
        • beezlebroxxxxxx 31 minutes ago
          The anthropomorphizing is part of the marketing. They'll never let up on it.
          • mcbuilder 23 minutes ago
            I mean CoT came out of research circles not marketing
            • cwillu 15 minutes ago
              It's impossible to tell the “it's all marketing!!11oneone” folks anything.
            • GPerson 16 minutes ago
              Research is salesmanship.
        • _superposition_ 6 minutes ago
          Nailed it
      • popupeyecare 25 minutes ago
        Maybe thoughts are just a chain of instructions in our head.
      • Angostura 34 minutes ago
        Chain Of Tokens
      • fooker 15 minutes ago
        What is thought?
        • _superposition_ 3 minutes ago
          Great question. I suspect it's more than tokens.
      • lossolo 19 minutes ago
        Yeah, basically they are using more computation to explore the solution space before producing the final answer.
      • ahofmann 28 minutes ago
        Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
    • 3asgfaf 29 minutes ago
      [flagged]
  • abixb 12 minutes ago
    I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

    If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

    As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

    • driverdan 5 minutes ago
      > If this is truly AGI (subject to one's definition of AGI still)

      Scoring well in a benchmark that's called AGI does not make an LLM AGI.

      • dmitrygr 3 minutes ago
        Hey now! Keep your reason out of their marketin^H^H lies!
    • mullingitover 5 minutes ago
      > If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.

      Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc

      I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

      • abixb 2 minutes ago
        >I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

        True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

    • catigula 8 minutes ago
      They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
  • Robdel12 1 minute ago
    I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

    So, folks that have actually used this already, what’s it actually like?

  • sashank_1509 2 minutes ago
    Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
  • tintor 49 minutes ago
    ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
    • andriy_koval 47 minutes ago
      I think it could indicate that "semi-private" dataset likely leaked to their training data.
      • IshKebab 40 minutes ago
        It says "Provider Adapter" so presumably they put some manual work in to make this work.
    • vb-8448 32 minutes ago
      But scored less on V2 and V1 ... too much overfitting?
    • minimaxir 43 minutes ago
      ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra

      tl;dr it's 62% when apples-to-apples to other models, which is still notable.

      • debazel 5 minutes ago
        ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
      • ciefa 27 minutes ago
        Woah, that is a crazy interesting read!
    • IshKebab 41 minutes ago
      Look at those costs!
    • Readerium 38 minutes ago
      saturated before (higher degree) AGI-2
  • putlake 47 minutes ago
    > GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

    Not on Azure? If so, that's a big deal.

  • the_duke 2 minutes ago
    Huge gains on some benchmarks, but for coding it sits barely above Fable

    It will be interesting to see how it performs in the real world ...

  • softwaredoug 1 hour ago
    I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

    https://venturebeat.com/technology/welcome-to-the-agi-era-op...

    • aabhay 1 hour ago
      This is with the caveat that OpenAI uses their own harness for this:

      > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

      • Readerium 35 minutes ago
        Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra
      • simianwords 49 minutes ago
        This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
        • ActionHank 27 minutes ago
          "This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."
          • simianwords 17 minutes ago
            > GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

            > Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.

            This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh

            • ActionHank 9 minutes ago
              "Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."
              • simianwords 7 minutes ago
                "I'll twist the words of the author of the benchmark itself to make a point"
    • kasperni 1 hour ago
      "On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

      But the comparison isn't straightforward.

      OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."

    • arctic-true 1 hour ago
      The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
      • _diyar 1 hour ago
        I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
        • _superposition_ 25 minutes ago
          Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.
        • aesthesia 45 minutes ago
          Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
        • CamperBob2 1 hour ago
          At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
          • jaggederest 42 minutes ago
            I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
    • Bluestein 1 hour ago
      100%, some say.-
  • petilon 37 minutes ago
    This is wild: OpenAI is basically declaring that AGI is here.

    https://www.theverge.com/ai-artificial-intelligence/989601/o...

    “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”

    • glenstein 8 minutes ago
      I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.

      So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.

    • rektomatic 33 minutes ago
      Remember when the term "AGI" meant something? Pepperidge farm remembers
      • breuleux 4 minutes ago
        I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
      • paxys 25 minutes ago
        No, because it has never meant a specific thing that everyone agreed on.
        • drop_star 18 minutes ago
          Does it pass the Turing test?
          • bryan0 8 minutes ago
            that would be a reasonable definition of AGI if everyone agree upon the specifics of the test, but that has never happened. Turing test is very much out of style, but I think that's because no one could even agree what the test was. I personally like the Kurzeil-Kapor version of the test and that is still unsettled: https://longbets.org/1/
          • bigfishrunning 16 minutes ago
            Depending on the proctor, ELIZA passes a Turing test. The Turing test is an interesting thought experiment, but isn't really a good measure.
      • seemaze 2 minutes ago
        Remember when The Verge was not a pay-walled visual headache?
      • layer8 15 minutes ago
        Remember when “Pepperidge farm remembers” meant something?
      • Rover222 20 minutes ago
        No, I really don't
      • 0xbadcafebee 25 minutes ago
        I think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
    • pluc 33 minutes ago
      Find me someone who isn't paid by OpenAI who is saying the same
    • ThouYS 26 minutes ago
      Wasn't that part of their contract with Microsoft? Some clause stopped biting with the arrival of AGI
    • tziki 21 minutes ago
      "OpenAI executive hypes up new model"

      Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.

    • RivieraKid 7 minutes ago
      Which is obviously wrong. This model can't learn continually and it can't learn effectively. Which is by the by the reason why robots didn't have a ChatGPT moment yet.

      These systems are still stochastic parrots. With enough data and params a neural net can learn anything but it's brute force learning, very inefficient and different compared to how human learn. If you don't have a massive dataset, which is the case with robots, this paradigm fails. If we had a system that can learn as effectively as humans, we could just build the robot and let it learn from experience - that would be the ChatGPT moment and possibly something that could be called an AGI.

    • mr_mitm 29 minutes ago
      Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
      • layer8 11 minutes ago
        The “I” alone is already not well-defined. That’s why.
      • enraged_camel 26 minutes ago
        They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
      • naasking 25 minutes ago
        It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
    • bigfishrunning 17 minutes ago
      Don't worry, they'll come up with a new acronym to mean really-real AI soon...
    • redox99 30 minutes ago
      I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
    • tastyface 33 minutes ago
      Renown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
  • mvkel 4 minutes ago
    The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.

    If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.

    Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.

  • E-Reverance 6 minutes ago
    At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
  • alex7o 10 minutes ago
    I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
  • BeetleB 34 minutes ago
    It's been over an hour, Simon! Where's the Pelican?
  • swalsh 59 minutes ago
    I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
    • greenowl 50 minutes ago
      This is AGI now. Why are you spending any of your time looking at the "quality of code"?
      • georgemcbay 35 minutes ago
        > This is AGI now. Why are you spending any of your time looking at the "quality of code"?

        Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.

        Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.

      • _superposition_ 29 minutes ago
        I can't tell if this is sarcasm.

        For the same reason you don't have your model write code in assembly.

        But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.

      • pennomi 24 minutes ago
        If you think any modern AI puts out stable, safe code, I have an AI-powered bridge to sell you.
  • dang 12 minutes ago
    Argh! I hit a wrong keyboard shortcut and moved the entire thread.

    Please stand by... it will all come back shortly

    • the_duke 11 minutes ago
      500 upvotes with 2 comments would have been a new record. ;)
    • layer8 9 minutes ago
      Luckily there’s a standard keyboard shortcut for “undo” as well. ;)
      • dang 8 minutes ago
        Not in the world of HN admins unfortunately
  • Readerium 15 minutes ago
  • theseamusjames 31 minutes ago
    Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
  • alex7o 4 minutes ago
    Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
  • udbhavs 14 minutes ago
    I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
  • pandinus 25 minutes ago
    Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
  • aliljet 48 minutes ago
    The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
    • Legend2440 45 minutes ago
      They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

      TL;DR all the other models are being crippled by limitations of their harness.

      >First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.

      >Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.

      • janalsncm 41 minutes ago
        Ok so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.
      • _superposition_ 15 minutes ago
        Exactly what I suspected. Of course a machine can just iterate relentlessly the way a human can't.

        I guess token counts are somewhat of a metric.

        IMO intelligence has peaked and all future gains will come from faster tps and more iteration.

    • polynomial 39 minutes ago
      This is absolutely benchmaxxing. Looking forward to hearing from Chollet about it!
  • trixn86 18 minutes ago
    Secret tip to win the mario cart clone: Just hold w, no steering needed.
  • herpdyderp 11 minutes ago
    The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
  • tosh 1 hour ago
    $10 per million input tokens and $50 per million output tokens

    sol is $4 / $20

    • jimmaswell 39 minutes ago
      It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
    • monroewalker 28 minutes ago
      Same price as Fable?
    • wahnfrieden 1 hour ago
      2.5x more expensive than Sol.

      Can expect 2.5x more usage in Codex subscription.

      Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.

      • AaronAPU 1 hour ago
        How is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?
      • janilowski 48 minutes ago
        How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.

        Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.

        • adam_arthur 40 minutes ago
          I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.

          I think some combination of:

          1) Using 1 thread for everything

          2) Reviving old threads which are no longer in cache

          3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.

          4) Non-coding workflow which is more output than input heavy

          5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR

          Just a guess. I think 3 is likely the primary reason.

          You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME

      • janalsncm 38 minutes ago
        If you are telling the truth you might want to check your network for any weird connections to Chinese LLM transit stations.
      • ModernMech 59 minutes ago
        How?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.
        • maipen 47 minutes ago
          These folks are probably using crazy plugins or crazy sub agent spams. They probably just run everything on max + fast mode which is ridiculous.
          • ModernMech 46 minutes ago
            The guy said medium/high regular speed so that's why I'm very puzzled! Ultra + Fast will absolutely slurp up your whole usage quickly but I've never found it gives substantially better results so I stick to extra high.
        • rowanG077 35 minutes ago
          Sub-agents. I have 7 20x accounts and I burn them within 1-2 days if I go fully parallel. In some scenarios I use 50 sub-agents for a session which is literally hours of usage for a single 20x account. I'm at the point where I need to parallelize over multiple machines because I just don't have enough CPU and RAM.
          • munimdev 18 minutes ago
            what are you doing with that many agents/tokens? very curious
          • ModernMech 19 minutes ago
            What are you doing with them you need so many? It sounds like a Gas Town situation, that you invented an exponential token burning machine.
  • oh_no 51 minutes ago
    Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
  • GodelNumbering 9 minutes ago
    I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
  • aliljet 1 hour ago
    The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
    • aesthesia 57 minutes ago
      ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
    • enraged_camel 52 minutes ago
      They used a custom harness. It's not a one-to-one comparison.
    • ionwake 1 hour ago
      my first suspicion is gaming - but i have no idea honestly
  • orliesaurus 51 minutes ago
    I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
    • jckahn 43 minutes ago
      Probably not. It's probably just gonna do tickets better and that'll be about it.
      • orliesaurus 40 minutes ago
        Fair point - hopefully you're wrong though ;)
        • ActionHank 30 minutes ago
          If this is really AGI, like really really, then this will be remembered as the day we all started on the path to building guillotines.

          More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.

          • noir_lord 4 minutes ago
            That's really the rub isn't it?

            We take their claims at face value then we should probably stop them training any more SOTA models til they figure out what they already built is safe or we assume theu are lying to juke the company valuation/keep the money train on the tracks and it turns they in fact were not and just took a sledgehammer to Pandora's box.

            We live in the strangest timeline.

          • wieiw1 20 minutes ago
            [dead]
    • sschueller 12 minutes ago
      Define AGI first. The singularly ain't going to happen with LLMs.
    • ministerk 25 minutes ago
      what launched today?
  • _ache_ 1 hour ago
    https://ache.one/gpt6_now_down.png

    Big claims, expensive and not release to the public yet.

  • jumploops 40 minutes ago
    > During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

    > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

    Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

    [0]https://x.com/MTSlive/status/2095227056040919202

  • MASNeo 42 minutes ago
    Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
  • ianm218 10 minutes ago
    I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
  • hazelnut 16 minutes ago
    Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
  • John7878781 1 hour ago
    You should know: AA index is only 61. Pretty surprised it’s that low.
    • jatora 47 minutes ago
      More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?

      And in the past, gemini 3 pro was rated as high as opus 4.5 and the like

      Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.

    • nsingh2 49 minutes ago
      I have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.
    • gekoxyz 31 minutes ago
      This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
  • udbhavs 14 minutes ago
    Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
  • Betelbuddy 30 minutes ago
  • dgellow 58 minutes ago
    > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

    Wait, what? Am I understanding that correctly? That sounds really bad

    • drakythe 49 minutes ago
      I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

      Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?

    • pixl97 49 minutes ago
      Nothing to worry about citizen, ignore the fleet of drones flying overhead.
    • order-matters 46 minutes ago
      the bullshit machine is learning to optimize its bullshitting techniques!

      <AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt

      you cannot give this type of worker autonomy over anything.

    • Laurel1234 38 minutes ago
      [dead]
  • simonjgreen 52 minutes ago
  • KolmogorovComp 46 minutes ago
    GPT-7 Zeneca
    • BoorishBears 25 minutes ago
      After they buy AZ, making this name foreshadowing
    • silver_sun 33 minutes ago
      GPT-8 Novo
  • gizmodo59 49 minutes ago
    99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
  • smashers1114 26 minutes ago
    I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
  • tekacs 49 minutes ago
    https://developers.openai.com/api/docs/guides/latest-model

    The docs page has a bunch more interesting details, including for example async tool calling!

  • jerrygenser 1 hour ago
    > The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
    • woah 58 minutes ago
      A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
      • ttul 43 minutes ago
        "The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
  • hannofcart 17 minutes ago
    What does 'Astra' here mean? Surely they must be referring to the Latin word.

    Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

    • BrokenCogs 15 minutes ago
      It's clearly an extension of the previous naming: Luna, terra, sol
    • manojlds 8 minutes ago
      Luna, Terra, Sol, Astra. Though Sun is also a star, should have called it Galaxy or something.
    • hokumguru 12 minutes ago
      Quite clearly in the same vein as Sol, Terra, Luna.
  • damsta 11 minutes ago
    Why release it now instead waiting those few days until it is available for everybody?
  • gekoxyz 42 minutes ago
    HTTP 500 for me on the announcement page :(
    • pampas 26 minutes ago
      The load bearing seam is broken for me too.
  • foundOpenRight 20 minutes ago
    1:15.425 on Sunset Cove beat my record
  • kegs_ 59 minutes ago
    I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
    • Kranar 56 minutes ago
      Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.
    • wyrdcurt 37 minutes ago
      More optimistic take: we'll only be second-class for a few months, if the pattern of Chinese models catching-up holds.
    • soricus 49 minutes ago
      When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.
    • resters 16 minutes ago
      Same. Fortunately DeepSeek keeps getting better.
    • I_am_tiberius 56 minutes ago
      If Tech CEOs consider this morally ok, then it is.
    • rs_rs_rs_rs_rs 18 minutes ago
      Is it really that hard to wait couple of days?
    • sxv 54 minutes ago
      create a life where your 'wealth' is decoupled from third party orgs.
      • kegs_ 45 minutes ago
        This is impossible, unless by 'wealth' you mean 'become like Buddha'.
        • 93po 14 minutes ago
          Material wealth is only a single type of wealth. Who's better off - the rich guy who's always yearning to be richer and never satisfied, or the lower income guy that mostly just cares about time with his family and is really happy where he's at?
    • atemerev 56 minutes ago
      They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".

      I am a researcher in a Swiss university btw.

    • pixl97 51 minutes ago
      This has always been the case for people that have not had piles of money.

      I mean do you get access to the best yachts?

      To the top of the 5 star hotels?

      To the best resorts?

      To the best military equipment?

      Hell, the best computer equipment has nearly always been out of reach of the average person.

      • tripleee 2 minutes ago
        I couldn't care less about owning a yacht.

        On the other hand even a modest house, basic healthcare and ability to not work like a slave for scraps feels like it's going to be out of reach.

      • rrr_oh_man 45 minutes ago
        …to basic health care?
    • PeterHolzwarth 57 minutes ago
      Oh please. They do closed betas - hardly makes you a "second class citizen".
      • kegs_ 49 minutes ago
        Mythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
        • pixl97 45 minutes ago
          Being strongly on the AI saftey side of things what is happening was 100% predictable.

          At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.

          • kegs_ 33 minutes ago
            What's currently happening is predictable, I agree. It's what's coming is the thing I'm worried most about. Either way, I'm not giving up.
  • jiraiyasarutobi 19 minutes ago
    It saturated most benchmarks. WTH
  • semiquaver 26 minutes ago
    Guessing this one will never show up in cursor…
  • Readerium 56 minutes ago
    • dang 55 minutes ago
      Link added to toptext. Thanks!
  • prometheus1992 42 minutes ago
    this is crazy! can't wait for the 27B distilled version of this.
  • balefulboy 21 minutes ago
    72 to 74 on DeepSWE is AGI
  • amazingamazing 57 minutes ago
    We have such great AI and cannot keep a static site up?
    • torginus 53 minutes ago
      Yeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.
    • pixl97 48 minutes ago
      Sometimes being the busiest site in the world for a few moments is difficult.
      • amazingamazing 47 minutes ago
        Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
        • pixl97 42 minutes ago
          Notice I said for a few moments. In a few hours traffic will drop a few thousand percent back to normal with no need for a CDN.

          OpenAI isn't making any money telling you about Astra on their site. All the capacity they have for it is likely sold for weeks or months.

          • amazingamazing 14 minutes ago
            You are making excuses for a a trillion dollar company. Wikipedia can do it.
    • agumonkey 34 minutes ago
      still hugged
    • gchamonlive 56 minutes ago
      That's the scientific positivism fallacy exemplified in one question.
    • gorgmah 53 minutes ago
      Yeah, apparently
  • firemelt 40 minutes ago
    damn seems I should hold off my claude subs
  • dowakin 31 minutes ago
    So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
  • jonplackett 40 minutes ago
    To a vapid any goalpost moving on such a critical issue as AGI.

    Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.

    For me it’s refusing to make a pelican.

  • tinyhouse 21 minutes ago
    You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
  • johnnyApplePRNG 34 minutes ago
    I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.

    I suspect these benchmarks are heavily benchmaxxed as well.

    5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft

  • colesantiago 24 minutes ago
    I'm going to call it.

    By 2030 all software is done and complete.

    But we are going to have more and new jobs.

    • jdee 13 minutes ago
      'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
      • colesantiago 8 minutes ago
        Yes.

        This is just another problem for the AI Labs to solve.

  • extr 32 minutes ago
    I heard they're calling it ASStra
  • bdangubic 4 minutes ago
    Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
  • ChrisGammell 5 minutes ago
    All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
  • saaaaaam 58 minutes ago
    Pelicans please
    • atemerev 55 minutes ago
      Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
      • maipen 44 minutes ago
        Very well said. It kinda describes how unrealistic these expectations are.

        Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.

        Very very unrealistic and wasteful.

        • droidjj 10 minutes ago
          The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
        • wieiw1 15 minutes ago
          [dead]
  • wahnfrieden 1 hour ago
    They're just announcing later availability. No launch.
    • paxys 1 hour ago
      Every frontier release nowadays is "we've launched*"

      * for a special group of customers that you're not in. Keep waiting peasant.

      • meowface 56 minutes ago
        That didn't happen with Fable 5.1 two days ago.
        • kegs_ 22 minutes ago
          5.1 was really more of an enterprise and bugfix update than a new model with the 0 day retention change
      • pixl97 41 minutes ago
        I mean tell Nvida to 100x their hardware output and you'll get what you want.
    • iAMkenough 54 minutes ago
      Their announcement about later availability is unavailable to me now (500 error).

      Great first impression.

  • guilhermeasper 1 hour ago
    That was a quick pull out.
  • danieltk76 17 minutes ago
    great, but nobody can use it for another 100 days right?
  • Pieczasz 1 hour ago
    Oh brotha, here we go again, it's so over again, as every week nowadays
    • wieiw1 18 minutes ago
      I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.

      But thank you for spending other peoples money to give us the tech regardless!

  • frozenseven 1 hour ago
    Release the Kraken!
  • Pym 1 hour ago
    I saw it
  • Onavo 56 minutes ago
    The jump in scientific performance is non trivial.
  • unrvl22 1 hour ago
    someone screenshot?
  • Brainspackle 1 hour ago
    huh?
  • k9294 8 minutes ago
    [flagged]
  • Cachecartii 22 minutes ago
    [flagged]
  • ealready_value 1 hour ago
    I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
  • killerdog10 9 minutes ago
    [dead]
  • paxys 1 hour ago
    Why is this flagged ?
    • dang 1 hour ago
      The link was 404ing quite a bit and several previous submissions got flagged as well.
      • consumer451 56 minutes ago
        It's still down for me, in the EU.
  • bicx 1 hour ago
    Dead link for me
  • rvz 1 hour ago
    > GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

    Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:

    Did humans deploy the model, Or did the model deploy itself?

    It sounds like "AGI" just stands for "IPO" as it always has been.

    EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.

    • Supermancho 27 minutes ago
      > Did humans deploy the model, Or did the model deploy itself?

      > It sounds like "AGI" just stands for "IPO" as it always has been.

      People don't usually respond to noise.

  • jonplackett 42 minutes ago
    The launch video is incredibly cringe.