How Complex Systems Fail (1998)

(how.complexsystems.fail)

96 points | by shortcrct 2 hours ago

11 comments

  • tptacek 29 minutes ago
    I'm a broken record on how important I think this document is, and that it's hard to appreciate it until you've had extended experience with complex systems actually failing.

    The most commonly cited subtext or thrust of it is that "root cause analysis", at least on complex systems, is a fools errand. Something goes wrong, say, in a distributed lock system, and your whole deployment system enters a metastable failure state. Naturally, the "root cause" seems like lock system resiliency. But definitionally a metastable failure is one that persists after the inciting condition is resolved. Now you have two "root causes", the lock failure and the metastability of the deployment system fault. Keep looking and you'll find more.

    But to me the biggest brick to the forehead in this piece is further observation that random things are failing all the time in any complex system. "Complex systems run in degraded mode". Resilient components are good, but it's the resiliency of the overall process that orchestrates the whole system that determines whether things are going to blow up.

    All practitioner actions are gambles. I should have that inked somewhere.

    • ErroneousBosh 16 minutes ago
      > I'm a broken record on how important I think this document is, and that it's hard to appreciate it until you've had extended experience with complex systems actually failing.

      If you want a bit of cheese to go with that wine, this article pairs nicely with The Grug-Brained Developer: https://grugbrain.dev/

  • icantevenhold 1 minute ago
    One of the great documents of our civilisation
  • jedberg 1 hour ago
    > Failure free operations require experience with failure.

    This is why we created Chaos Engineering. By constantly forcing failure, it made us always create systems in defense of that failure, and gave us great data on where the tipping point is for different systems within a particular failure mode.

    • AlotOfReading 50 minutes ago
      I've always struggled to apply this to the systems I work on. If the system fails, someone potentially dies, though in practice they've never been more than hospitalized. To avoid that, huge amounts of effort are expended on failure modelling and testing, but that doesn't eliminate unknown unknowns. That discrepancy has made front page news a couple times.
    • obscurette 26 minutes ago
      It's much more universal and complicated than that. One of the big issues schools and pedagogy in general struggle with is that our environment is far too safe in too many ways. For kids there is too few ways to learn from failures. Attempts to solve the issue look often like "Hey, kids, let's fall over now all at once in safest way possible and learn from it!". But it doesn't work at all. Real failures have to be unexpected, related to your decisions and really hurt so that you can learn from them.

      PS. Btw, I am certain that this is the main cause of the mental health crisis amongst young people.

  • feyman_r 1 hour ago
    I may have shared this before on a different submission: John Gall’s books are really good on this topic: General Systemantics [https://en.wikipedia.org/wiki/Systemantics]
    • littlecranky67 1 hour ago
      Gall's law is amongst my favorite ones and with decades of experience in software development, I have to say it holds absolutely true:

      > “A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works and cannot be patched up to make it work. You have to start over with a working simple system.” — John Gall, Systemantics (1975)

  • squirrel 32 minutes ago
    The definitive work on this topic is Normal Accidents, with a modern retelling in Meltdown.

    https://en.wikipedia.org/wiki/Normal_Accidents

    https://en.wikipedia.org/wiki/Meltdown_(Clearfield_and_Tilcs...

  • rowyourboat 1 hour ago
    All of this sounds just like any air crash investigation I ever read
    • tptacek 26 minutes ago
      Richard Cook was a UChicago anaesthesiologist who took up safety systems research after studying patient safety; some of his work is rooted in Three Mile Island, and some of it comes from aviation safety.
    • shash 56 minutes ago
      And industrial accident investigation (except the ones with low regulation or whatever). And market or supply chain collapse, and civilization collapse (late Bronze Age anyone?)
  • yipinwong 1 hour ago
    I think there are a few common themes to the failure reasons, but cannot get my hands on it.

    This seems like a list of reasons while I am looking for more abstract directions on how to prevent them.

    ---

    I am trying not to use AIs to just do that for me to tinkle my neurons.

    • shash 57 minutes ago
      I think, part of the point is that it’s not possible to have a recipe to prevent failures. They are cascades of many events coming together to fail in an a priori non obvious way.

      Or so I read the [site? article?]

  • mohamedkoubaa 1 hour ago
    I can't tell if the article is describing how complex systems fail or if they are using failure characteristics to define complex systems.
    • shash 55 minutes ago
      It’s more about failure. It’s right there on top.
  • haemdahl 36 minutes ago
    [dead]