9 comments

  • atkrista 14 minutes ago
    I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
    • nbardy 2 minutes ago
      I think it's weirdly just a choice of deciding to cut releases.

      We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

    • petesergeant 6 minutes ago
      > largely 2-horse

      The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

      • bayindirh 0 minutes ago
        Gemini is also pretty nice for researching things. It turns out that having the whole indexed and having unlimited access to YouTube is a force multiplier of some kind.

        Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.

      • Bluestein 0 minutes ago
        ... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
    • TeMPOraL 12 minutes ago
      AGI will make one, about humanity, after we're all gone - "They were so dumb, they just deserved to die".
      • cindyllm 9 minutes ago
        [dead]
      • f6v 5 minutes ago
        The sooner the better, brother.
  • madsgarff 10 minutes ago
    It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
    • nextaccountic 7 minutes ago
      > It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode

      But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)

  • egeozcan 16 minutes ago
    Worries of AI going rogue take so much attention that no governance body seems to care about the shady subscriptions and limits business.
    • semiquaver 7 minutes ago
      Do you mean the subsidized subscriptions which allow individuals to pay a tenth of the normal API cost for tokens?
  • max979 3 minutes ago
    Wild pace. Guess they found a critical bug or a quick win to push it out so fast. Astra is a high bar.
  • Pythagon 3 minutes ago
    Is this a duplicate thread of this? https://news.ycombinator.com/item?id=49896586
  • moomin 20 minutes ago
    This has got to be a panic move from OpenAI, right? They’ve had some bad press lately because from their billing changes, and Anthropic have finally released a fast, relatively cheap Opus with improved written English.
    • nba456_ 18 minutes ago
      Didn't they announce the billing changes the same day?
  • Dinuda 57 minutes ago
    After 5.5, I basically don't notice a jump in model performance, other than my usage ending sooner.
    • zero1009 20 minutes ago
      Felt the same until I started using Luna. I feel like I get similar performance, but faster, and my usage lasts so much longer.
    • ModernMech 2 minutes ago
      Same. 5.5 got work done then 5.6 was also fine then 6 was maybe not quite as good. Now with 6.1 they are cutting usage and raising prices and introducing ultra fast mode, but things were good enough 5 months ago.
    • tom1337 16 minutes ago
      Kinda same but I miss my 5.3 Codex. Thing lasted forever on my $20 subscription and with detailed prompts was able to pretty much implement everything I requested it to do with a acceptable quality.
    • ndbe 17 minutes ago
      [dead]
  • 77rushi77 4 minutes ago
    [flagged]
  • tobin1994 14 minutes ago
    [dead]