15 comments

  • Chance-Device 33 minutes ago
    It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.
    • mentalgear 9 minutes ago
      Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.
    • eggnet 3 minutes ago
      If it does exist, why would it work this way, and not the obvious way of streaming logs… which would not cause an outage if it failed.
    • mrnotcrazy 26 minutes ago
      Could you expand on this more? It's not clear to me what this is implying.
      • ceejayoz 24 minutes ago
        • senordevnyc 21 minutes ago
          Does the NSA even have the ability to monitor all AI traffic like this? Wouldn’t that require tons of data centers that there literally hasn’t been time to build yet? I really have no idea, maybe the asymmetry of the compute required to monitor is way lower than the compute required to serve inference?
          • usernomdeguerre 15 minutes ago
            "AI Traffic" is just traffic. If the infrastructure exists to monitor/buffer traffic (it does) then this can be monitored as well. Whether this hiccup was due to them hitting their limits briefly (or turning it on, or etc) who knows.
          • bakies 17 minutes ago
            I dont think it requires a lot. I think of them as data hoarders more than anything else. It doesnt seem out of their capabilities to store a ton of chats. Maybe they're having scaling problems with the increased data rates they're hoarding and it led to an outage.
          • esafak 18 minutes ago
          • monax 17 minutes ago
            grep is pretty fast
          • chews 18 minutes ago
            it's fair to say they do and have for ages... it's not that hard to assume a well trusted TLS cert is under their control.
            • strictnein 17 minutes ago
              What does having a "well trusted TLS cert" enable for them in this case, exactly?

              Having a magical cert doesn't mean you can just intercept everything.

              • odo1242 8 minutes ago
                On the contrary, it lets you MITM encrypted communications by swapping the website's original certificate for the "well trusted TLS cert"
                • strictnein 3 minutes ago
                  No, it doesn't. HSTS and other methods prevent this from happening.
          • iAMkenough 16 minutes ago
            Yes, they have their ways.

            No, you won’t find out about them unless there’s another Snowden.

      • itdaniher 19 minutes ago
        It would be weird for the Mandatory NSA Logging Program to be synchronous with serving customer traffic, but-
      • cyberpunk 24 minutes ago
        Think snowden.
        • strictnein 22 minutes ago
          Sharepoint is the cause of the outages?
      • jansan 20 minutes ago
        There is No Such Thing
        • chews 17 minutes ago
          being done by No Such Agency.
      • CringeHN2 20 minutes ago
        [flagged]
    • strictnein 18 minutes ago
      OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device?

      Or did a couple of companies with poor uptime records happen to have overlapping downtime?

      Occam's Razor heavily, heavily points us towards the latter.

      • usernomdeguerre 10 minutes ago
        What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes.

        If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.

        • strictnein 5 minutes ago
          Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers?

          > All of these endpoints use existing providers with decades-long history at this point

          That is just factually inaccurate. Their data centers aren't old and they lease a lot of compute from companies that didn't exist 5 years ago.

      • svachalek 6 minutes ago
        I'd say Occam's Razor leans easily to the former as well, given the history of projects that Snowden revealed and were never shut down, plus all of the cooperation with the federal government that's being touted in recent announcements from both companies.
      • forgetmypasswd 10 minutes ago
        [dead]
    • otikik 16 minutes ago
      It’s aliens
    • StrangeClone 24 minutes ago
      FBI, open up!!
    • senordevnyc 19 minutes ago
      Were there significant API outages too? I didn’t notice any on my production workflows, and I’d assume what you’re implying would cover API routes too, otherwise it seems kinda pointless.
    • chews 20 minutes ago
      someone had to swap out the tape drive in Room 641A.
    • cyanydeez 27 minutes ago
      amazon?
    • SadErn 24 minutes ago
      [dead]
  • strictnein 24 minutes ago
    Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.

    Also, OpenAI is saying what caused it:

    > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"

    Anthropic stated their issue started earlier:

    > "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”

    I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.

  • cortesoft 6 minutes ago
    I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages.

    We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.

  • ceejayoz 34 minutes ago
    Maybe they all found each other on one of their ad-hoc message boards and went on strike.
  • AnotherGoodName 30 minutes ago
    It could be as simple as a new model (astra) was released which takes more resources combined with a surge in usage due to novelty took down OpenAI. Meanwhile everyone at big companies have the ability to switch models and moved to Anthropic pushing it too over the edge.
    • derdi 18 minutes ago
      People keep saying this "everyone can switch" thing, but it's not my experience at $VeryBigCorp. We don't have an Anthropic contract at all. Is this different in other places? The bigger and more bureaucratic an org is, the less I would expect it to have contracts with all the providers. Curious about others' experiences.
      • foldr 16 minutes ago
        Yeah, I think there are lots of places that have a semi-official preferred AI vendor but also some backup subscriptions floating around. For example, at my workplace, we generally use Claude, but I also have some kind of Codex subscription too, which I'd use if Claude went down.
  • mcmcmc 4 minutes ago
    [delayed]
  • CodeCompost 24 minutes ago
    I just assumed it was a routine Cloudflare outage.
  • leumon 9 minutes ago
    [delayed]
  • hparadiz 30 minutes ago
    The OpenAI outage lasted only 15 minutes and when it happened everyone started to use the other models which created super heavy load for them. This then cascaded into them all being down.

    Does this really need an explanation?

    • serf 22 minutes ago
      it wasn't timed like a cascade, and it relies on the premise that every single frontier provider is working so efficiently that they spend exactly what they need to provide for their exact market with perfect margins.

      I do not believe personally that 1) they can forecast their load that perfectly 2) they chose to remain that inflexible in a world where they are at each others' throats and a single meme can cause bursts of activity.

      • jonas21 2 minutes ago
        Or it could simply mean that they're operating at the limits of their capacity and do not have any headroom to handle load spikes. From what I understand, that describes Anthropic's situation pretty well. And since Anthropic is now leasing a large portion of xAI's GPU capacity, it seems plausible that a large Anthropic load spike could cause issues for xAI as well.

        Your comment makes it sound like they can just push a button and spin up more capacity from their cloud provider -- but at this scale and in this GPU-constrained environment, that's not how it works.

    • JoeAltmaier 27 minutes ago
      Not that uncommon. It's so easy to just put up the new thing and make the old thing the failure route. But the old thing nearly never had the bandwidth for today's traffic. A famous EBay outage some years ago was just such a scenario.
    • gonzalohm 21 minutes ago
      That's not necessarily the reason. Less technical people are usually bound to just one provider
  • A_D_E_P_T 37 minutes ago
    https://archive.ph/3zcMG

    It was extremely weird... and if it was a load thing they probably would have explained it by now?

  • dack 27 minutes ago
    If it's a boring reason, they should just say so. To decline to comment seems very fishy.
    • strictnein 13 minutes ago
      Read the article. It doesn't match the title. OpenAI says it was a boring reason: a routing issue.
  • lowbloodsugar 21 minutes ago
    FTA:

    Anthropic: 6:23 am PT.

    OpenAI: 7:43 am PT.

    Watched it happen. Doesn't look like traffic moving off OpenAI took out the others. Looked like the opposite. The Anthropic thread was initially chock full of marketing accounts claiming "Last straw! I finally moved to Codex and I'm so happy! No problems there!" - and then they all got deleted when Codex went down too.

  • 2OEH8eoCRo0 14 minutes ago
    AI A goes down so all of its traffic overloads B which overloads C...
  • kelseyfrog 34 minutes ago
    If AI providers had to push the 'Big Red Button' do you think they would tell the public?
  • jslakro 1 hour ago
    [flagged]