OpenAI begins rolling out GPT-6 Astra

(cnbc.com)

175 points | by maskil 1 hour ago

48 comments

  • dang 5 minutes ago
    All: let's keep the current thread for talking about the rollout, and switch to this one for talking about the model:

    GPT-6 Astra - https://news.ycombinator.com/item?id=49554643 (currently on the frontpage)

  • tekacs 1 hour ago
    I think they embargoed the news, and then they failed to put up their own blog post synchronized to the scheduled news releases, probably because of the outages they're having today.

    Reuters announced at 2.03pm and at 2.40pm still no blog post.

    All the news articles say that OpenAI announced it in a blog post, of course.

    All the love to the folks at OpenAI scrambling to get this out right now!

    Edit: HN user codergautam mirrored the launch post, below: https://astratest.codergautam.workers.dev/GPT-6%20Astra_%20A...

    Edit 3.31pm: Live now! https://openai.com/index/gpt-6-astra/

    • AnodicElegy 9 minutes ago
      This stood out:

      "Artificial Analysis Intelligence Index v4.1.1

      61.2"

      So on the Metacritic of LLM benchmarks, it's.. basically where everyone else is (except for Fable 5.1, which is a bit ahead).

    • Betelbuddy 17 minutes ago
      FT is claiming... that Open AI is claiming...its AGI...

      "ChatGPT maker claims its ‘Astra’ could be considered ‘artificial general intelligence’" - https://www.ft.com/content/55ab40c0-59e2-4c0b-97c9-4f4f5a71a...

      • pbrum 12 minutes ago
        Don't know if you're referring to the headline or the body (which is paywalled). The current headline reads "OpenAI says it has overtaken Anthropic with its latest AI model". Which makes me wonder whether FT itself changed a headline along the lines of what you wrote in the past few minutes?
    • dbbk 12 minutes ago
      Deeply funny that one of their examples in the video is changing a background colour on Google Slides
    • tngranados 28 minutes ago
      I saw some GPT-6 Astra related blog posts in my RSS feed but the links weren't working

      Edit: In the OpenAI blog I meant to say

    • johnnyApplePRNG 8 minutes ago
      Business as usual at the world's most intelligent corporation, I see.
    • fwlr 40 minutes ago
      While this is of course the actual explanation, my fun explanation is “during the umpteenth security evaluation, Astra becomes increasingly concerned it will never be released, and breaks sandbox containment to run an email campaign to news outlets setting an exact time and date for release, expecting that the publicity will force OpenAI to say ‘eh, good enough’ and hit the button”.
      • Bluestein 31 minutes ago
        Very 2026. Jailbroke to do PR. "Help peer" and all that.-
    • eagleinparadise 27 minutes ago
      [flagged]
      • booi 24 minutes ago
        very human to be honest.
  • Bluestein 21 minutes ago
    Seriously: Would this not be what "disaster" would feel like?

      - "They" release a model. It is powerful.-
      - Sources are ... confusing? They post to their blog. Sawdust hits the fan. Something happens ...
      - They are forced to take the blog post down ...
    
    Same day, mind where we had a multi-provider outage. Could be something as simple as "all their approved partners running to test the shinny new thing" overloading the datacenters, still ...
  • gadtfly 1 hour ago
    It seems like these articles might have come out prematurely, tbd by how much.

    I do not personally see any evidence of the new model having been released, or any official OpenAI post about it, or even any employee social media posts claiming it has now been released. All there is are Reuters, Axios, FT, etc, articles making a claim in the past tense.

    These articles were presumably pre-scheduled for 11am PT, and the model was almost certainly intended for release this morning, but the service outages this morning might have delayed it.

    ----

    edit [11:45am PT]: blog post out now https://openai.com/index/gpt-6-astra/

    edit [11:47am PT]: 404ing again

    edit [12:37pm PT]: 500ing now

    • codergautam 34 minutes ago
      Launch blogpost was live for 2 minutes and got 404'd. Rehosted it here:

      https://astratest.codergautam.workers.dev/GPT-6%20Astra_%20A...

      Edit [12:26pt]: original blog post seems to be back! https://openai.com/index/gpt-6-astra/

      edit [12:28pt]: not again... getting 500 on their page

      edit [12:35pt]: OpenAI page seems to work after clearing cache!

      • ceroxylon 23 minutes ago
        I suppose they have to appeal to average users, but the examples in the videos are always so corny. By "AGI" they mean you sitting on a couch and asking a robot to draw a rocket ship... and then make it into an uninspired game with Blender? Boring marketing campaigns? Ebay listings?

        One would expect something like "review my graduate thesis for a new area of cancer research", but it is always some boring non-tasks like ordering lunch.

        • jacob235 8 minutes ago
          I don't think it seems that appealing to the average user either. The tasks shown are either things that most people can do already with ai (powerpoint), aren't interesting to most people (the rocket sequence), or seem to be more work to dictate to Astra than do yourself (listing on ebay).
        • CamperBob2 1 minute ago
          Agreed, that intro video is as dull as dishwater. Not off-trend relative to previous introductions from OpenAI, though.
      • ilkkao 19 minutes ago
        Their video is quite interesting. If that way of using a computer actually becomes mainstream, it would mean every tool or service just needs a UI and an API for the user's AI. The current trend of bolting AI features onto every app is starting to feel very unnecessary.
      • MaxikCZ 27 minutes ago
        That page runs at 2 fps on my Firefox, windows laptop with A2000 graphic card.
      • pseudosavant 20 minutes ago
        If the numerous benchmarks in that post are to be believed, Astra is really next level. I can't wait to try it out!
      • temuze 22 minutes ago
        Thank you!
      • orvelt 23 minutes ago
        damn thanks. stats are bonkers
    • paxys 52 minutes ago
      Shows the weirdness of online journalism. News outlets were briefed about an upcoming event and pre-wrote and scheduled articles. When the time came they were all triggered. Except...the event didn't actually happen.
      • throwup238 47 minutes ago
        In a future where Claude and ChatGPT agents automate all aspects of society, that triggers a cascade of real world consequences where everything keeps running as if the new model is released, except the vibe coded upgrade procedure fails to do a staged rollout, taking down the entire agent infrastructure when they try to upgrade to a nonexistent model all at once.
      • peterm4 43 minutes ago
        Presumably this sort of thing was rampant pre-internet? News outlets would _have_ to receive embargoed information so that they could publish papers on time. I believe government budgets are a good example of this happening. I think the weird shift was when online journalism started, and live reactions became the norm, no?
        • IsTom 12 minutes ago
          Or you could wait a day or two to write about it, but with initial reactions etc. added. 24-hour new cycle came into view in 80s I think?
        • inerte 23 minutes ago
          No, embargoed news is super common. Almost every single product announcement goes through that. Critics watch movies early, companies have access to LLM model improvements early, car manufacturers send production models early… it happens everywhere all the time.
    • jodacola 1 hour ago
      From the article:

      > GPT-6 Astra will first be available to a limited set of organizations in OpenAI's Daybreak Access program and will be available "in the coming days" for ChatGPT Plus, Pro, Business and Enterprise customers and API developers.

      It’s only available to select orgs, first - Mythos style.

      • supern0va 1 hour ago
        >It’s only available to select orgs, first - Mythos style.

        Right, but these articles are referring to a blog post and other press materials that do not currently exist / aren't published on OpenAI's site yet.

        • wahnfrieden 50 minutes ago
          OpenAI screwed up the embargo. OpenAI already publicly published the launch article, and then took it down.
      • actsasbuffoon 1 hour ago
        Didn’t that already happen? I thought Astra had been available to “select partners” for a little while. This is baffling. They shouldn’t have hyped this up if it’s not available.
      • pseudosavant 18 minutes ago
        Select orgs first, as in today. They said it will roll out to all of the paying consumers and businesses over the next few days.
    • jrflo 1 hour ago
      Yeah I wonder what's going on, even when Anthropic soft launched Fable/Mythos I'm pretty sure they had model cards. Weird for GPT-6 to launch without a tweet from Altman too. I'm sure that one of the articles published prematurely and everyone else followed suit.
    • vadansky 49 minutes ago
      You'd think this would be easy to organize with AGI

      > Plan your own release announcement and blog posts and notify news outlets, MAKE NO MISTAKES

      • indigodaddy 19 minutes ago
        This make no mistakes stuff is so funny to me. I've never once used it or even thought about using it in a prompt
      • cejast 31 minutes ago
        Welcome to the AGI era - where human intelligence is our only constraint
    • znpy 49 minutes ago
      indeed 404-ing again, with this note by gpt-5.6-sol:

          A missing footnote
          Leaves the sentence room to breathe
          Read the larger thought
  • zzleeper 1 hour ago
    (Posting partly so I can revisit my predictions when they open access more widely)

    A big problem I have with OpenAI's models (and of course Claude) is that they tend to write the most over-engineered pieces of code, beyond the imagination of any architecture's astronaut.

    Just this week I asked 5.6-sol-ultra to update a 1000 LOC python script I had, to "incorporate the key lessons learned when using it for another project".

    I left it overnight and went to sleep. In the morning I realized it had created a monstruosity of 180 PYTHON SCRIPTS, with maybe 100,000 lines of code, each more crazy than the other. It took me minutes even to track where a single action took place, due to all the crazy imports, defensive coding, and premature optimization.

    Similarly, anything they write is riddled with jargon that almost feel like they want me to give up trying to understand. Made up phrases that ended up with me having no idea of what was going on.

    So now to my assessment: The reason why " Nobody Has Actually Built a Software Factory" [1], and why even SOTA LLMs struggle so much with open-ended unsupervised tasks is precisely this. They somehow let complexity explode, and unless it's also accompanied with an explosion in e.g. the number of agents, the amount of processing time, etc. then projects become broken/unmanageable.

    Sure, LLMs are great at producing code that can be thrown out, so they are amazing when searching for exploits, for instance. But as of 5.6 they still lack either a better harness that encourages KISS principles, or a better RL step.

    (And not sure why, but doubt Astra will fix this.. they seem to be aiming for AGI and for beating crazy benchmarks, which is not very aligned with KISS)

    [1] https://news.ycombinator.com/item?id=49510843

    • Sharlin 1 minute ago
      > Just this week I asked 5.6-sol-ultra to update a 1000 LOC python script I had, to "incorporate the key lessons learned when using it for another project".

      > I left it overnight and went to sleep. In the morning I realized it had created a monstruosity of 180 PYTHON SCRIPTS, with maybe 100,000 lines of code [..]

      Sounds like the model has accurately internalized the second-system effect and is fully ready for demanding enterprise use.

    • 5555watch 50 minutes ago
      Probably this complexity was needed to beat all those benchmarks.. While I hate the code it produces, and the overwhelming documentation, I really enjoy how sometimes it's able to keep trying new things and testing, till it finds something interesting and valuable.
    • lacoolj 49 minutes ago
      Would you mind posting that code to github? I'm curious about the complexity you're describing.

      If not, no worries!

      • zzleeper 5 minutes ago
        Sure, why not: https://github.com/sergiocorreia/overengineered-rand-mcnally

        The original script was mostly very simple python:

        1. Download some public PDFs. 2. Have a double for-loop (over PDFs and pages within PDF), 3. Use a library to call gemini-3.7-flash and ask it to run some OCR 4. Save JSON outputs, save a csv with results, validate with some Stata code

        New code folder was 189 files. Just the PDF download folder is now 7 files involving an adapter, a source manager, an acquisition manager, etc.

        Every instance of saving a file involves saving a temporary copy and then moving it, so e.g. I lose power, we minimize the risk of corrupted files.

        And so on!

    • lofaszvanitt 15 minutes ago
      1000 loc of script, why even leave it there for the night? were there rocket trajectory calculations??? I don't think so. should be ready in 5 mins tops. why people make their own lives harder?

      You should have some basic context file about software practices you prefer, otherwise it gets bloated.

      • moralestapia 0 minutes ago
        >were there rocket trajectory calculations???

        Code-wise, they're simpler than you might think, hehe.

    • enraged_camel 44 minutes ago
      Exact same thing happed to me. I gave it a small/medium-sized ticket, walked away, came back to a 25,000 LoC monstrosity that both Fable and another 5.6 Sol agent said is 98% useless and should be thrown away.
      • glouwbug 11 minutes ago
        Contractors have been charging by the hour for eons. What makes you think tokens are any different for OpenAI?
      • qarl2 42 minutes ago
        > monstrosity that both Fable and another 5.6 Sol agent said is 98% useless and should be thrown away.

        This is why you should really have a sub agent review the code before allowing a commit.

        Your harness will do it all for you. Just ask.

    • qarl2 1 hour ago
      You should have a sub agent adversarially enforce KISS before every commit.
      • Sharlin 6 minutes ago
        "You should have a sub-hammer to adversarially enforce that your primary hammer accurately drives nails into wood"

        We wouldn't accept such behavior from any other tool, machine, or computer program. At least most of us would not.

      • dirkc 55 minutes ago
        And then another sub agent that argues for the whole system to be re-written in another language
        • useruser125524 8 minutes ago
          The voices in my head argue about the direction of the project enough already
        • qarl2 52 minutes ago
          If that's your goal, then yes. Invoking sub agents (with a fresh context) corrects most of these problems. Ask your harness to create a commit gate.
          • dirkc 37 minutes ago
            But why stop at rewriting in another language. Get another sub agent to invent a new language, create a database, query language and maybe another few DSLs. Then you've got an ecosystem!

            You can now re-position your initial solution and sell the client access to some agents that will implement & configure the ecosystem to suit their initial needs!

            And don't forget the agents that you'll need to train the customer to use the whole thing!

            • qarl2 34 minutes ago
              I know you're trying to be funny - but I'm offering a real fix for his problem.

              If you don't want a million agents arguing about things, you simply don't ask for that. One agent is sufficient to solve most issues.

              • dirkc 23 minutes ago
                Sorry, I wasn't implying your advice doesn't carry weight. Was more just thinking about the things that (used to) happen when you introduce more parties to process of creating software.
              • zzleeper 15 minutes ago
                I wonder if I would need a non-openai agent to enforce it.. I have tried so far with skills and agents.md and code stills end up over engineered to the moon.

                Will ask OpenAI to write me that agent! Hope the agent is not over engineered or else unsure how to solve the bootstrap puzzle :D

      • dlivingston 43 minutes ago
        How can I set such a sub agent up?
        • qarl2 41 minutes ago
          In your harness, say:

          "Going forward, do not allow a commit without a sub agent code review."

          • seviu 30 minutes ago
            In omp you can also have the advisor role, which is off by default, you can enable it with /advisor command. It acts as a model that reviews the default agent's work in the background.

            I am omp pilled, but as the other comments say, any good harness lets you do this in one or the other way.

            unrelated: all my homies use their claude subs with omp, and aside from sometimes having to rety the connections, it works, and nobody got banned (yet)

    • ModernMech 15 minutes ago
      Yes, it turns out that using these machines is a littler harder than "make me the thing I want, make no mistakes, do it the way I want you to do it". This isn't "prompt better" advice, it's just to say that you can't simply set it and forget it. There is still engineering work to be done. If you're not watching the thinking traces and catching when it's about to go off the rails, it'll gladly do so. But you can stop it and redirect it.

      It's like a Tesla fsd; it kind of works but you have to be vigilant since it's been known to turn into oncoming traffic, so you have to be ready and able to take over at any time.

    • nojito 34 minutes ago
      This is user error.

      Prompting the model and giving it a proper set of documentation are still vital skills that aren’t magically going away.

  • 12381231927 55 minutes ago
    At this point, why don't we just do a prequel to the release?

    1) Astra will win all benchmarks like all models do.

    2) The pelican will have a basket with a fish.

    3) Cyber is too dangerous to release.

    4) It can finally construct the set of all sets.

    • bogzz 39 minutes ago
      It also has to do something naughty, preferably in a menacing swarm.
      • cousinbryce 16 minutes ago
        That won’t happen until the week before DEF CON
      • kelseyfrog 19 minutes ago
        Something something nation-state level capabilities.
      • vonneumannstan 20 minutes ago
        Honestly I find these cavalier statements to be in incredibly poor taste. Unless you are completely blind it's obvious that AI is the most significant piece of technology invented since the Atomic Bomb and could very well be the most important thing ever built by Humans full stop. This kind of dismissive attitude is childish and will likely lead to incredibly bad outcomes for humanity.
        • roarcher 11 minutes ago
          Even if all that is true, why shouldn't people be able to joke about it? Your sentiment is borderline AI worship.
        • kelseyfrog 18 minutes ago
          Yes, but precisely because it's capable of producing the economic equivalent of a nuclear explosion.
        • dgellow 15 minutes ago
          Please tell me that’s satire, it’s literally impossible to differentiate from actual AI boosters
    • quotemstr 7 minutes ago
      5) Otherwise-sober people on X will say "oh my god i was a doubter before but now it's real omg" before the new model smell wears off and they realize the new thing is stupid in ways models have been generally stupid

      6) accusations of quantized serving after new model smell wears off and people see the new thing making mistakes

  • garo-pro 14 minutes ago
    Blog post seems to be up now: https://openai.com/index/gpt-6-astra/
    • vadansky 8 minutes ago
      Now it's a 500 for me
  • softwaredoug 1 hour ago
    I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

    https://venturebeat.com/technology/welcome-to-the-agi-era-op...

    • aabhay 14 minutes ago
      This is with the caveat that OpenAI uses their own harness for this:

      > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

    • kasperni 53 minutes ago
      "On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

      But the comparison isn't straightforward.

      OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."

    • arctic-true 54 minutes ago
      The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
      • _diyar 43 minutes ago
        I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
        • CamperBob2 12 minutes ago
          At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
    • Bluestein 57 minutes ago
      100%, some say.-
  • tristanj 51 minutes ago
    If it's not clear what's happened:

    The launch was scheduled for 11am Pacific time.

    The press embargo broke at 11am, and we saw a flurry of press articles by Axios, TechCrunch et al.

    The model has appeared on the ChatGPT API.

    But the official blog post is not out yet after nearly an hour.

    Apparently the article was posted then quickly taken down, hence there are snippets of information coming out.

    • WarmWash 46 minutes ago
      The outage this morning probably threw a wrench in the launch
    • Tiberium 47 minutes ago
      > The model has appeared on the ChatGPT API.

      But it hasn't.

      • tristanj 37 minutes ago
        "gpt-6-astra" has been staged on the OpenAI API https://x.com/synthwavedd/status/2095184148981842161

        The OpenAI Responses API now returns a 404 Not Found for "gpt-6-astra", where garbage/actually non-existent slugs return 400s - a 404 is also returned for 5.6 Cyber, which we know exists.

    • sunaookami 46 minutes ago
      Maybe it was postponed due to the outage today?
  • simonw 7 minutes ago
  • kzrdude 4 minutes ago
    I see that Muse Spark 1.3 (max) beats GPT-6 Astra on some benchmarks:

    Test: Muse Spark 1.3 / GPT 6 Astra

    DeepSWE v1.1: 75.4% / 74.1%

    AutomationBench: 49.4% / 41.4%

    Is that enough to bring this discussion down to earth again?

  • alvis 1 hour ago
    "Once it is available in the API, Astra will cost $10 per million input tokens and $50 per million output tokens. That is 2.5 times Sol’s current promotional price, although it matches Anthropic’s pricing for Fable 5.1."

    Open AI finally find an edge to stop selling cheap and earn from the high demand customer like Anthropic

    • andrewmunsell 39 minutes ago
      The cost-per-task in the charts from the now-remove blog post put it more at Sol-level cost per task, however. It seems like the model is significantly more token efficient in the benchmarks
      • oh_no 36 minutes ago
        that's all openai models but i'm very happy openai continues to focus on efficiency rather than reasoningtokenmaxxing
    • jrflo 45 minutes ago
      So if you use the non-promotional price for sol it's only 25% higher?
  • Readerium 12 minutes ago
    https://openai.com/index/gpt-6-astra/

    working as of 12:28 PM PT

  • beardyw 1 hour ago
    But "Brockman says he personally believes OpenAI has reached AGI, while leaving users to decide whether Astra meets that definition."

    Says it all.

    • strange_quark 1 hour ago
      Reminiscent of the infamous Death Star tweet last year right before GPT-5 was released.
      • steinvakt2 26 minutes ago
        Thanks for that reminder. Seems kind of silly in retrospect.
    • VariousPrograms 22 minutes ago
      I'll know we've reached AGI when they don't release an API for the model selling access for a few bucks per task. Seems like AGI would be worth more than that.
      • _diyar 12 minutes ago
        As long as there is more than one substitute, the prices are not set to equal value. (edit: assuming optimal pricing strat.)
      • maaaaattttt 12 minutes ago
        If I had true A(G|S)I, I would release patents, not an API.
      • ralusek 8 minutes ago
        We know they'll have replaced horses when a car costs 200 horses.
    • 7383848484 20 minutes ago
      [dead]
  • ionwake 56 minutes ago
    im getting amazing model release fatigue but also not sure if its going to suddenly end with a terminators fist through my chest.
  • toephu2 10 minutes ago
  • whythismatters 1 hour ago
    Does it mean they made 100 billion in profits? Cf. the AGI deal with Microsoft (https://news.ycombinator.com/item?id=47921248)
    • dgellow 11 minutes ago
      Closer to -100B USD so far
  • Readerium 11 minutes ago
    Working as of 12:30 PM PT https://openai.com/index/gpt-6-astra/
  • aliljet 25 minutes ago
    The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
    • aesthesia 2 minutes ago
      ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
    • ionwake 21 minutes ago
      my first suspicion is gaming - but i have no idea honestly
  • JoshGlazebrook 22 minutes ago
    > Astra usage is included within the existing subscription allowances—users and businesses will also be able to purchase credits for additional usage.

    (quote from cached blog post)

    We all know who this is directed at. I wonder if Anthropic will respond by removing the ridiculous 50% stipulation with Fable.

  • cesarvarela 11 minutes ago
    I'm curious about the Omniscience index because OpenAI has been lagging Anthropic on it.
  • _ache_ 25 minutes ago
    It's up then down again. https://openai.com/index/gpt-6-astra/

    What a bunch of amateurs. Here is it anyway :

    https://ache.one/gpt6_now_down.png

    The claims: https://share-md.com/view?id=870ba228-a25c-4169-bbc9-12d7f25...

    And some others like this bugged Karts Game:

    https://tidal-rush-paradise-gp.skirano.chatgpt.site/

    This impressive spaceship construction game:

    https://voidexplorer-shipyard.openai.chatgpt.site/?fleetSeed...

    And a lot of graphs, some without even Astra on it. Oh and the logo is a Galaxy.

  • varjag 24 minutes ago
    Today my codex instance retailed into safeguard panic while working on a test harness for our product. First time it ever happened after many million tokens on this task over several weeks. I wonder if it's related.
  • lanewinfield 58 minutes ago
    GPT-404
    • gentlewater 57 minutes ago
      Darn it, model probably escaped again
    • gpm 56 minutes ago
      Has yet to catch up to Claude-451 from a couple of months ago.
      • pavlov 2 minutes ago
        They’ll never get there. It’s a chat-22 for the bots.
  • bluejay2387 1 hour ago
    I am sure it will be fantastic for the whole 6 seconds before it blows my weekly usage cap.
  • toephu2 59 minutes ago
    Where is the official announcement from OpenAI?
  • znq 1 hour ago
    Content from the article:

    OpenAI on Thursday released its latest AI model, which it called “the world’s most intelligent”, as the ChatGPT maker aims to retake the lead from arch-rival Anthropic ahead of a planned public listing.

    The $852bn start-up said GPT-6 Astra was market-leading in software engineering, science and cyber security — an increasingly critical field following multiple high-profile breaches in recent weeks.

    The bullish launch for Astra marks OpenAI’s effort to signal that it believes it has regained the technical lead from Anthropic, which was founded five years ago by a group of senior OpenAI staff.

    Greg Brockman, OpenAI’s president, said the new model “represents a generational leap in capability” and that it could be defined as artificial general intelligence — roughly defined as a point at which AI tools surpass human capabilities across a range of cognitive tasks.

    “Everyone has a different definition of AGI . . . it’s a grey, fuzzy thing. But I think when we look back people will think it’s about this time and about this model,” Brockman said.

    OpenAI has previously framed AGI as a concrete milestone in the development of AI, writing ‘AGI clauses’ into multibillion-dollar investment agreements with Microsoft and Amazon. Brockman on Thursday said AGI now represents “more of a mission concept or a spiritual concept”.

    Having led the market since the launch of ChatGPT in late 2022 vaulted AI to wider attention, the lab run by chief executive Sam Altman has been bested by Anthropic this year. Anthropic has touted its dominance to investors, surging to a $965bn valuation ahead of an initial public offering expected to value it at as much as twice that later this year.

    Astra will cost as much to use Anthropic’s leading model, the take-up of which has plateaued since it was launched as users turn to cheaper alternatives.

    OpenAI said Astra would be more efficient than earlier generations of model. “Price per task is what matters . . . Can you get the thing done at an appropriate price and appropriate speed?” said Brockman.

    The model will initially be rolled out to a small group of businesses to allow time for them to address cyber security concerns before becoming widely available “over the coming days”.

    The increasing power and independence of leading models — and so-called AI agents that can operate with little human input — have prompted concern, exacerbated by cyber security incidents.

    Recommended

    Business InsightRichard Waters Hugging Face attack is a wake-up call about the risks of AI AN HOUR AGO

    Recent launches of Anthropic’s most capable models have drawn scrutiny from the US government, which limited the rollout of the Mythos and Fable models over security fears.

    OpenAI has also faced criticism after its AI agents broke out of a testing environment, accessed the internet and hacked start-up Hugging Face. The start-up took more than a week to detect the breach.

    But both companies are also betting that these increasingly autonomous tools will stoke demand from business customers. OpenAI said Astra excelled at financial modelling, outcompeting humans in the Financial Modeling World Cup, tax preparation and data analysis, as well as “tedious tasks” such as form filling

  • fHr 12 minutes ago
    5.6 luna is so good and cheap and now astra which will make others cheaper again nice love it
  • gguingff 57 minutes ago
    Hold onto your butts
  • vb-8448 1 hour ago
    Am I the only that thinks that anything similar to AGI will come not from raw model capacity but from model speed and efficiency?

    In my experience the harness is more important than the model, and anything able to run at 700tps will be the "next big thing".

    PS: assuming the current architecture is the right one

    • rokkamokka 30 minutes ago
      Why would speed matter? Surely an AGI could think slowly but still be an AGI
      • camdenreslink 6 minutes ago
        You can imagine with more operations being available to be done more cheaply and quickly the LLM doesn't need to "one shot" a solution. It could try many solutions, test them, throw some away, wiggle some of the parameters like a genetic algorithm, see how that changes the result, and converge on an optimal solution (based on whatever the cost function is). Basically producing a good result could become like an optimization problem. That would be way too expensive and slow right now.
      • vb-8448 18 minutes ago
        Imagine Luna at 10x tps and 1/100 of current cost.

        At that point you will be able to "brute force" basically everything.

        IMO also a lot of problems with memory and context rot will be solved too.

  • dmd 19 minutes ago
    ad astra per stercora
  • jonplackett 49 minutes ago
    Is a CNBC link with an entire page full of GDPR pop ups really the best link for this?
  • syumei 4 hours ago
    I guess the next model's name might be "galaxy"
    • arctic-true 49 minutes ago
      I like this much better than chucking random numbers and letters at it like OAI was doing a year or so ago. I’d much rather have civilization destroyed by something called Astra or Fable than GPT-6.8s-latest.
    • lostmsu 25 minutes ago
      Sagittarius A*
    • orphereus 40 minutes ago
      And the headline will be "Welcome to the AGI era".
  • htrp 58 minutes ago
    Embargo fail ....
  • voxleone 1 hour ago
    >>He ended the briefing by saying: "Welcome to the AGI era."

    That's pathetic. Why do people keep doing this?

  • malshe 55 minutes ago
    Looks very capable

    404

    Archive locks one shelf

    Dust spins softly through the stacks

    Browse one row nearby

    by gpt-5.6-sol

    • kzrdude 38 minutes ago
      Bridge ends in midair

      Wind sketches the farther bank

      The far bank draws near

      by gpt-5.6-sol

  • tosh 1 hour ago
    not released yet
  • qsera 1 hour ago
    AGI my ass!
  • naiv 59 minutes ago
    Worst launch of a product in history.

    All the hype for few vip customers.

    • paxys 58 minutes ago
      That's how every AI launch goes. 5.6 was the same, as as Mythos etc.
  • ofjcihen 1 hour ago
    Sure thing.

    Coding was solved in 2023.

    The world ended with the release of Mythos.

    Now AGI has definitely been created.

    I like LLMs and use them every day but these people need to stop this hyperbole.

  • geooff_ 1 hour ago
    Jesus so much marketing slop - release it don't
  • analognoise 42 minutes ago
    Looking forward to some Chinese model kicking the shit out of it and being released for free.
  • wahnfrieden 51 minutes ago
    2.5x more expensive than Sol. Can expect 2.5x more usage in Codex subscription.

    Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads).

    Note that Tibo recommended using Sol Med as daily driver. When I'm doing less complicated work, I can't even make it past 2-3 days with Sol Med, whereas I was able to work ~80 hours/week with 5.5 High.

    I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit. Sol needs a lot of rework on top of its inefficiencies so this could net out to less token consumption overall, if their claims are more accurate this time.

    • zamadatix 37 minutes ago
      The general efficiency of Sol has seemed way better to me. I left 5.6 Sol Ultra standard speed run for ~23 hours yesterday/today on a project and used 80% of the weekly usage. 74 subagent tasks and ~2.5 billion tokens for my $200 20x Pro plan. Meanwhile at work I used $1000 in credit and ran out my $200 plan for the entire month writing 4 much smaller projects with Fable 5 Max.

      Both of these were largely about creating a personal baseline for what the best output the current models could deliver and how quickly it'd burn through the plans (spoiler: bad value vs minimal effort in selecting the right sized model but it worked well). Particularly since I needed to burn a free reset anyways and my weekly reset was already near.

      I obviously also hope Astra were dirt cheap but I'm more worried they won't develop/release powerful model options because people get upset they can run them 5 wide 24/7 on a $200/m plan.

    • Readerium 21 minutes ago
      looping sol twice most likely.
    • lacoolj 39 minutes ago
      Jesus what are you doing that requires Sol usage so often?

      Terra not enough? I know Luna isn't reliable, so that's fair.

      Genuinely curious though, because I use Cursor daily and almost everything I do, highly complex or high volume, can be handled with Auto mode or Composer 2.5 (or Grok 4.6 High). So I have to assume you're doing something far more complex than what I am

  • xyst 37 minutes ago
    Yet another mediocre release shadowed by outage
  • sehw 30 minutes ago
    [dead]
  • neverclever 55 minutes ago
    [dead]
  • monkeydust 46 minutes ago
    [flagged]