The AI Race Just Got Awkward

(insufferable.dev)

223 points | by allisdust 56 minutes ago

30 comments

  • cmiles8 23 minutes ago
    Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

    • jedberg 13 minutes ago
      What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

      An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

      • JackFr 8 minutes ago
        But the analogy still holds.

        The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.

        • jdonaldson 0 minutes ago
          Yeah, the whole thing seems like a human centipede of rug pulling. Probably the same as it's always been. Curating AI knowledge should be something that we put our best researchers towards, but realistically I think we wind up with 2-3 highly biased nationalistic models that are constantly copying off each other's notes.
        • nonethewiser 0 minutes ago
          He didn’t say it was an analogy. He said both are distillation.
      • cmiles8 8 minutes ago
        I get that angle but it’s a weak argument as Anthropic is doing the same to others. Also while there’s certainly a lot of computing power needed to do what Anthropic does, it’s increasingly clear there isn’t much secret sauce involved. Everyone knows how do to the core work it’s just a question of who wants to burn billions on compute to do it.

        Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.

      • bushbaba 3 minutes ago
        And the communal work of humanity is orders of magnitude more work than what anthropic pays for their scraping of content. I got no check from them for my contributions
      • faangguyindia 9 minutes ago
        Isn't it better for planet? By not doing the wasteful transformation work again
        • jedberg 7 minutes ago
          Absolutely. I'm not taking a side here, I'm just pointing out why Anthropic might have a valid complaint.
          • isolay 3 minutes ago
            That complaint is invalidated by the argument of tu quoque. Complaining about something they are doing themselves.
      • baxtr 3 minutes ago
        Wait, wasn’t 95% of the work creating the content in the first place?
      • arctic-true 9 minutes ago
        They spend 95% of the money, perhaps, but burning compute is not the same as doing the work.
        • jedberg 7 minutes ago
          I'm calling "work" here the conversion of energy to LLMs.
      • watwut 10 minutes ago
        You are going to be surprised to hear how many resources were necessary to create all the data Anthropic is digesting
        • ipsod 8 minutes ago
          From one perspective, the 5% estimation is near-infinite orders of magnitude off, since they've trained on something approaching the sum-total of human knowledge.
      • mrwh 12 minutes ago
        I mean, 95% of the work if you don't factor the work to create the training data in the first place...
        • cyanydeez 10 minutes ago
          also, the actual work is the _copyrighted material created by the world_.
        • off_with_their_ 6 minutes ago
          [dead]
    • nonethewiser 1 minute ago
      >Aren’t Anthropic’s models not just distilling down other people’s work?

      Can you elaborate on that? I mean my direct answer would be no, of course not. But what is your reasoning?

    • jrflo 1 minute ago
      Because cost of original training >> cost of distilling. It's the same thing that happens with Chinese knockoffs of physical products - it takes a lot of money and R&D time to design a new product, but it's basically free to buy the product, reverse engineer it, and resell it. All the data they originally trained on was available for free on the internet. If the original work was so valuable, it shouldn't be up on the internet for free in the first place imo.
    • sergiotapia 9 minutes ago
      "That’s called competition. You’re allowed to test somebody else’s products all you want." - Jensen Huang https://x.com/wallstengine/status/2104604118937735553
    • jorblumesea 15 minutes ago
      $$$$

      it's not complex. there's hundreds of billions of investor dollars counting on vendor lock in and walled gardens

      • 2OEH8eoCRo0 12 minutes ago
        It ain't gonna happen. At work I have a dropdown menu in vscode with a dozen models to use interchangeably. They're all essentially commodities and will compete on price and squash almost all profit margin.
        • dpweb 6 minutes ago
          That's not their business model. They won't win on price, but they won't compete on price. Their business model is making the current state of the art.

          If I'm a business and I need something done today, and bc Anthropic has the best model, there's a 99.9 chance it will be completed successfully for $1000. And using Deepseek there's a 70% chance it will, for $10 - you or me will go for the $10. Big businesses don't. Bc 1000 per task is nothing to them.

          • HWR_14 1 minute ago
            Yes, large corporations frequently pay orders of magnitude more for slightly better software. That's why Oracle produces the best stuff on the planet.

            The real issue is that Deepseek has a 99.7% chance. So I can run it 10 times until it works and still pay 1/10 the money.

          • bushbaba 1 minute ago
            Actually opposite occurs. Big businesses are ok with a mediocre but cheaper result. Very few are willing to pay such cost. Just look at tech wages and the distributions
      • teaearlgraycold 11 minutes ago
        Sorry but it’s looking more and more like the top American labs won’t have any kind of moat.
    • cyanydeez 11 minutes ago
      Because no one outside the AI scientists understand what distilling means. They probably all think about Mash and a vodka still, and a completely unrelated association.

      The word itself is the pivot, not anything else.

      • hn_throwaway_99 8 minutes ago
        > They probably all think about Mash and a vodka still, and a completely unrelated association.

        I don't know anyone with even a passing understanding of how LLM training works that thinks that is the appropriate analogy.

    • dominotw 12 minutes ago
      [flagged]
      • cmiles8 1 minute ago
        Then Anthropic should stop saying it. So long as they try to play victim here folks are going to call out their BS.
    • nater5000 15 minutes ago
      [flagged]
  • slowin 29 minutes ago
    I'm also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

    Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

    • 10xDev 18 minutes ago
      An authoritarian regime is not your friend and will pullback the moment their own models become highly capable.
      • pksebben 12 minutes ago
        Oh no, they might stop doing the thing that benefits me and that they were never required to do in the first place.
      • computerex 1 minute ago
        As opposed to what? The US? You think the US is any different? Literally our pedo president publicly admits to insider trading. You think the US government gives a rat's ass about the American people?
      • slowin 14 minutes ago
        There's no "pulling back" things that have already been open sourced.
        • abirch 0 minutes ago
          Unfortunately most of the Chinese models are open weights and not open sourced.
        • HappyPanacea 10 minutes ago
          Intelligence wants to be free
      • horsawlarway 15 minutes ago
        Yes, we already discussed the US.
        • Avicebron 10 minutes ago
          It's crazy how articles like this get spawn-camped by people like this trying to throw this zinger in. Both AI conpanies in the US and the chinese companies with ccp desks in the corner can be bad. The good path forward is locally hosted AI models, that's known.
          • horsawlarway 6 minutes ago
            Right, which is why I'm happy to see China continue to innovate in the open, and increasingly wary of the US stance given articles like

            https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

            It's VERY clear that the US companies are trying to push for regulation to kill open models and open weights. I see this as much more hostile and authoritarian response than what we're seeing come out of China right now.

            So is China going to always publish in the open? No clue. But right now they're modeling much better behavior.

      • feverzsj 6 minutes ago
        It's already happening now. They just banned engineers of their top AI labs and their families from leaving the country.
      • randbyte 12 minutes ago
        Just like how Anthropic and OpenAI is already doing?
      • nutjob2 12 minutes ago
        China is much more authoritarian than the US, but at this point it's like comparing two types of metastatic cancer.

        The point is get what you can from both to develop open models, data and tools.

      • CodingJeebus 15 minutes ago
        This is equally true for US AI
      • 4gotunameagain 16 minutes ago
        While your friend is Sam Altman, or US megacorps ?

        Or did they not pull back when their models allegedly became highly capable, with the whole mythos debacle ?

      • ahriad 13 minutes ago
        Chill, Buddy. Why are you so anti-American?
    • ducktective 18 minutes ago
      >distillation datasets

      You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?

      • forshaper 14 minutes ago
        There are several? And there exists companies whose entire business is just providing them? iirc
      • atherton94027 7 minutes ago
        Given the amount of people complaining about crawlers in the past 2 years, I think it's the latter
      • frabcus 13 minutes ago
        It's particularly important we all scan and destroy our own unique books!
    • joe_the_user 1 minute ago
      The Chinese models are to an extent that distillation data.
    • skybrian 22 minutes ago
      There’s a libertarian sentiment that that doesn’t sit well with “AI is harming people” sentiment. If AI has harmful uses, and I think anyone sensible would have to agree that it does, then giving everyone unrestricted AI is likely to make it worse.

      It’s sort of like gun nuts arguing that more guns is the answer. I mean, ok, maybe you’re a responsible gun owner or AI user but relying on personal responsibility doesn’t fix systemic problems. There are bad people out there.

      • bronson 17 minutes ago
        If food has harmful uses, and I think any one sensible would have to agree that it does, then giving everyone unrestricted food is likely to make it worse.

        You can do this with cars, tools, computers, ... whatever you want. So, no, I think your point is wrong.

        • skybrian 14 minutes ago
          We do in fact have car and food safety laws. Regulation is normal.
          • hyperlinerapp 8 minutes ago
            Guns are highly regulated. Try getting a gun in liberal California, where Reagan screwed us.

            Now, what I want to regulate are accordions.

        • bobmcnamara 8 minutes ago
          Nice try Philipp Mainländer!
      • hamdingers 16 minutes ago
        I simply don't trust the people who would decide who gets AI (or guns) to make good choices.
        • skybrian 13 minutes ago
          This is a common populist sentiment, but if you don’t trust anyone then nothing can be done. Is it just game over?
          • hamdingers 4 minutes ago
            I didn't say I don't trust anyone. Don't put words in my mouth, it's a sign of bad faith.

            Other countries have governments that have earned that level of trust. I believe the US could get there eventually, but it will take a very long time because it has a very long way to go.

      • hyperlinerapp 9 minutes ago
        Gun owner enters the conversation. In high trust societies, armed people are very polite people. I don’t want the bad guys out there being the only ones with guns. Besides I like to shoot just like you like (whatever you like to do that is legal).

        Replace “gun” with anything and you will see how your comment falls apart.

        What’s next? A registry for food purchases? Your beer gut is starting to show.

        • skybrian 1 minute ago
          If you mean high-trust societies like maybe Switzerland, I agree that it can work, but the US has lots of guns and doesn’t seem very polite, so how we’re doing it doesn’t seem to be working very well.

          We do have lots of food safety regulation, which has more to do with selling food.

      • iamnothere 20 minutes ago
        All concerns balance against competing concerns, and in this case freedom of computing and knowledge wins over safety. Especially since it’s trivial to copy and share open models.
        • skybrian 17 minutes ago
          Okay, you’re asserting that but I disagree. Why should anyone else be convinced? Why can’t we get the good uses without the harms? It doesn’t seem like an unavoidable tradeoff.
          • iamnothere 4 minutes ago
            Well then do your best, people like me will keep sharing models just fine. (Not like I even do anything with them, but I’m a compulsive data hoarder.) It’s not like the copyright industry has been able to stop sharing either, and there’s serious money at stake there.

            I predict that the AI scaremongering will fizzle out when the bubble bursts. There will still be die-hard believers but the public will lose interest.

          • short_sells_poo 12 minutes ago
            I agree that it isn't an unavoidable tradeoff in principle, but looking back at our (as in humanity) track record, it is 99% likely to be.
            • skybrian 7 minutes ago
              Populist doomers will tell you that nobody can be trusted, no global problems will ever be solved, nothing can be done, game over, don’t even try.

              People did use global agreements and regulation to fix the ozone hole, though, so I think there’s a chance.

        • tonyedgecombe 15 minutes ago
          Actually I think profits win over safety.
      • nutjob2 7 minutes ago
        > giving everyone unrestricted AI is likely to make it worse

        Worse for whom?

        The only effective defense against predatory corporate and government AI is personal protective AI.

        Anything else is unilateral disarmament. It's the only way individuals can survive in the worse case scenario.

        > gun nuts

        Guns are different. They can't protect you against the government, contrary to gun nut claims.

      • slowin 12 minutes ago
        I would say that there are very few people on earth that I trust less than Sam Altman, Dario Amodei and Elon Musk. Also my own government claims to have used Anthropic models to bomb a girls' school in Iran. If you combine US regulations with sociopathic private companies, you get into a worst case scenario for humanity imho. Again, I'm thankful to China or any other entity pushing open models, local models and even distribution of this technology.
  • listless 11 minutes ago
    I'm beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We're going to need more powerful models than we have now if we're gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

    I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.

  • reticulates 21 minutes ago
    “So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.”

    I don’t think it is intentional but this is actually quite bad for the western labs.

    The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.

    The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.

    • DangitBobby 13 minutes ago
      I don't see why revenue has to fall even if marginal costs drop off a cliff. As long as they have the best models (perceived or otherwise) and can make security and IP guarantees that satisfy enterprise, and no firm with similar guarantees undercuts them on price (why would they want a race to the bottom?) they can have high revenue and high margin.
      • reticulates 7 minutes ago
        unless the major players collude they don’t get decide if they are in a race to the bottom. The best model was compelling 6 months ago when everyone was too impressed to care about price but that has worn off now and clients are paying attention to price. The best model is no longer a license to charge any amount.
      • bobmcnamara 6 minutes ago
        Costs dropping opens you up to competition on price.
      • dominotw 9 minutes ago
        There isnt a lot of money in enterprise ai. Also my enterprise company gives me glm.
    • altcognito 13 minutes ago
      It is funny that so many comments vascilate between "It is so expensive these companies can't make money and will go bankrupt in seconds" and "Inference is so cheap that these companies can't make money and will go bankrupt in seconds".

      I never take them seriously, I just assume they are coming from countries that don't understand how capitalism works or are operating out of bad faith. The underlying reality of the market is always changing and needs are always changing. Some AI companies will fail, that is a given. Remember alta-vista? Yahoo? Did search go away? How about Microsoft phones? Nokia? Motorola?

      OpenAI and Anthropic are not in the inference business. That is a commodity. They need to sell products and solutions.

      • reticulates 0 minutes ago
        They’re not contradictory positions. Inference is too expensive now to make money because the industry is immature and hasn’t yet optimized for financial success while customers don’t care much about price because they’re more concerned about not missing out.

        Inference will be too cheap long term to make money because it is being commoditized and customers will start to care about results and not just be wowed by impressive technology.

  • reedf1 25 minutes ago
    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
    • bix6 16 minutes ago
      $9k for a 5090 now? Sheesh.
      • bitexploder 2 minutes ago
        Well, I have a $750 card that runs at about 50-60% of that token rate :)
      • rubyn00bie 2 minutes ago
        In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

        Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

        Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

      • off_with_their_ 14 minutes ago
        $9k is a small price to pay to experience the rapturous glory of AGI. I'd easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.
        • literalAardvark 0 minutes ago
          Except you can do that cheaper by renting compute
    • oidar 20 minutes ago
      What are you thoughts on it's performance compared to 4.6?
      • reedf1 9 minutes ago
        Indistinguishable or very mildly better. But it's considerably faster. Some portion of that is also probably down to improvements in model harnesses, I've been using opencode.
    • teaearlgraycold 10 minutes ago
      Frontier from 9 months ago? I don’t know about that. But it sure punches above its weight.
    • redanddead 22 minutes ago
      Well how’s it been so far
  • bwest87 12 minutes ago
    The best explanation is that it's a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to "commodotize your compliments".

    I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

    • bobmcnamara 5 minutes ago
      It's also a huge propaganda opportunity to influence the distribution of groupthink.
  • eggbrain 32 minutes ago
    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    • kennywinker 23 minutes ago
      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
      • lumost 20 minutes ago
        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.

    • londons_explore 11 minutes ago
      For nearly all tasks, I want the fastest and smartest AI model.

      It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.

      Therefore, I believe we are nowhere near 'good enough'.

      I never drive my steam engine to work these days. It isn't good enough.

      • danielmarkbruce 3 minutes ago
        I find this too. In fact, recently I've been pushing more and more to the latest and greatest model every time there is an update. It just saves me so much headache.
    • ericol 28 minutes ago
      > people will soon *stop paying

      Think you missed a word there.

  • rglover 16 minutes ago
    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    "Thus the expert in battle moves the enemy, and is not moved by him."

    They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and 'falling behind'—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.

    Checkmate.

  • why_only_15 2 minutes ago
    Why do you think the Chinese labs figured this out before the western labs? No reason to believe that whatsoever.
  • georgeburdell 3 minutes ago
    To answer the author’s question of why Chinese labs give away their work for less than cost, the answer is involution. China is struggling with overcompetition in other areas of its economy as well, such as electric cars, and perhaps ironically its labor share of income is substantially lower than the U.S.
  • amelius 36 minutes ago
    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Any ideas?

    • curuinor 31 minutes ago
      There are 1000 Chinese labs. They are involuting, they cannot coordinate and the state won't let them coordinate because the state wants domination, not actual profits for anybody. So the market forces are allowed to dominate.

      Because of the basic huge recession going on in China, you can't actually make money in China doing China things. So they gotta gird up their export stuff and try to export. That entails strong relations with American companies, American PR, English stuff, etc.

      If you want an essay about this from a VC, read this one

      https://earnedintuition.substack.com/p/involution-without-ex...

      • dabedee 25 minutes ago
        What a strange way to put it. Market forces are a good thing in a market economy. Only someone who secretly wants or hopes for monopolies would you say something to the contrary (a VC).
        • DrewADesign 12 minutes ago
          Pouring 100% of your tech optimism, and probably portfolio, into a product that some say is about to get Wile E Coyote flattened by market dynamics would probably inspire serious market skepticism.
        • curuinor 20 minutes ago
          The party has done something about enormous involution in solar panels, for example (https://www.csis.org/analysis/chinas-solar-industry-upheaval...) and previously steel. They're planning something for cars. They don't on LLM because of the newness and the wish for preeminence.
        • ajkjk 19 minutes ago
          "only X would say Y" is a rhetorical device (the no-true-scotsman fallacy, if you want) that should basically never be used ever.
        • iamnothere 16 minutes ago
          China has a different perspective on it, they believe that there is such a thing as harmful competition and they are willing to step in to stop it.

          Enshittification and related problems can be a result of market forces just as much as they can be a result of monopoly/duopoly or a small cartel. Excess competition sometimes results in all firms scraping the barrel to squeeze out pennies, especially with technology (such as large online marketplaces) making pricing more transparent.

          Marx actually predicted that ever-intensifying competition would destroy markets through overproduction, although he did not use the term involution.

      • luke5441 17 minutes ago
        I'd call it overcapacity instead. A lot of investment without capital discipline making sure there is actually return of investment leading to too much supply.

        Not that OpenAI, Anthropic or SpaceX aren't doing the same.

      • googaar 15 minutes ago
        Nice read. American media does a terrible job of covering this.
      • bilbo0s 9 minutes ago
        >you can't actually make money in China doing China things

        Do you do business in China?

        I'm curious what you mean by this? Because in my experience, you can only do business in China by doing "China" things.

        I'd be interested in picking your brain as to how you get around those issues?

        • curuinor 6 minutes ago
          I don't do business in PRC anymore, haven't for an amount of time that means I don't know anything anymore, basically.

          I'm talking like, getting 100x, VC sized returns. Of course you can sell widgets in China, it's a major world economy.

      • thrawa8387336 25 minutes ago
        LMAO recession, China? You've been reading too much Brad Setser
        • curuinor 22 minutes ago
          The youth unemployment rate is at 19% with employment counting as 1 hour a week...
    • carbonguy 25 minutes ago
      My immediate midwit take is: doesn't matter if it helps Anthropic/OpenAI if it helps DeepSeek more, relatively. Making open-weights models even cheaper and easier to run expands that "market" and increases competitive pressure on the Big Two, who still have to charge money.
    • teekert 15 minutes ago
      Idk, but it's doing a lot for my view of China. Maybe that's a point? Maybe they just want their own innovation to go as fast as possible and they don't care that other countries also benefit? A rising tide lifts all boats? They are already know for the best manufacturing, they're just adding software dev to the list? Maybe they just want to undermine the US in a non-aggressive way?

      Why did we (the west) ever start open sourcing anything? Maybe we just like sharing? Maybe humanity only grows on pre-competitive layers like Linux and clean water. Maybe, the chinese government is closer to their people, and does not let large companies influence them and just doesn't like closed private hyperscalers with a lot of power?

      (Some points assume the government has a role in the openness, which I think is likely)

    • audunw 9 minutes ago
      I think it’s fairly simple: they’re forced into this situation by being late and worse in terms of capabilities. They’re not far behind, but as long as they’re behind they’ve needed to give people some reason to try and use their models. Cost is one factor. But it probably wasn’t enough. Being open has given them a lot of attention. Free marketing. Good will.

      Put another way: if they were not cheaper and open, they would simply not be competitive. They would already be dead.

      I don’t think this ends well for the Chinese labs. This is going pretty much like I thought. Western labs is just copying their improvements (I don’t think publishing the techniques matter here.. they’d just hire to gain the knowledge or figure it out themselves), and they have access to more GPUs and have better branding, so in the end where can the Chinese labs compete? Even lower cost? Open weights? I’m not sure open is a sustainable way to compete either. Eventually there will be some fully open source AI models that cuts out that avenue of competition as well.

    • pj_mukh 29 minutes ago
      Occam's razor: Going to closed-source just to hide KV-cache optimizations seems silly?
      • twoodfin 21 minutes ago
        This looks like a speed run of the history of analytics DBMS’s.

        Once upon a time, everyone had a secret sauce in network or data encoding or query optimization, but in the last ~10 years computational physics and economics have basically decided the “correct” architecture and everyone (including OSS) has converged.

    • HeavenFox 6 minutes ago
      The post makes an assumption that US labs did not already possess similar optimization. It's also very possible that they did, but are simply not telling anyone in order to maintain obscene margins on cached read, similar to AWS' absurd pricing on bandwidth.
    • seydor 13 minutes ago
      The chinese don't view AI as metaphysical, they view it as an engineering challenge they consider good for their state and want to dominate the global market like they do with batteries/EVs/photovoltaics. They want to proliferate them as much as possible and traditionally they don't care much for IP. They also want hardware makers to make optimized chips specifically for these models.
    • Catloafdev 14 minutes ago
      Yes - the Chinese labs serve a market that rely on open-weight models and managed deployments, and the labs gain competitive relevance by releasing those models. The cache optimization feature they came up with required new software to utilize on the inference-end, meaning that open source software would need to be specifically updated to work with these models. It wasn't the type of advancement that they could even theoretically keep secret.
    • feverzsj 15 minutes ago
      It's just their usual national strategy like what they did to solar pane and EV. The solar panel industry is mostly dominated by China and their profit rate is basically ... negative. The EV industry in China is in similar condition, where the average profit rate is only 1.5%. Their upstream suppliers are also hold as hostages that most of them won't get their money back within 6 months.

      The weird ideology here is to dominate the market at ANY COST, even it benefits the opponents.

    • thefourthchime 13 minutes ago
      Because it's entirely possible that Western labs already did this optimization but didn't publish it, and then the Chinese figured it out and decided to brag about it.

      We don't know either way, so I find the whole thing silly to speculate on.

    • TrackerFF 23 minutes ago
      My guess would be that if they "help" western labs becoming better, then any break-throughs they (western labs) make after that, is also a benefit to the Chinese labs - if they can distill the models.

      Basically, western labs are in it for the money / commercial monopoly. Chinese labs are in it for the tech? As long as they can keep distilling models, and get access to research other ways, they benefit. And if they can push western labs forward, they'll benefit from that themselves.

    • corford 22 minutes ago
      "A week in Beijing and Shanghai with the people building AI in China": https://earnedintuition.substack.com/p/involution-without-ex... does a decent job of exploring some possible reasons
    • ambicapter 32 minutes ago
      They want the Western labs to keep digging themselves into a hole? You can do so by egging on the true believers on your side to egg on the labs on the other side.
    • Windchaser 17 minutes ago
      > Any ideas?

      Unpopular, maybe, but what about the normal reasons? The researchers are looking to make a name for themselves, and/or they genuinely care about AI advancement.

    • jollyllama 23 minutes ago
      Where do you think most of the hardware is manufactured, and do you think the hardware manufacturers will keep getting paid if labs start going under?
    • chrismarlow9 18 minutes ago
      AI fundamentally insecure. Vulnerable to forcing hallucinations via search results. Vulnerable to invocation of commands in data stream. More AI means more vulnerabilities.

      I can't even fathom the trend these days of "we don't review the code" from security team perspective.

      Just my guess though.

    • foul 31 minutes ago
      Market manipulation or slowing down demand for chips for a bit/moving the offer elsewhere temporarily?
    • mpalmer 23 minutes ago
      They would like to see Western civilization keep getting dumber, and if that means bolstering the success of Western firms, that's okay.
    • transdev12 22 minutes ago
      [dead]
  • sigbottle 9 minutes ago
    This is insanely cool, what the hell.

    How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?

    I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.

    This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol

  • moooo99 4 minutes ago
    In all honesty, all these distillation complaints brought forward by Anthropic etc make me enjoy the cheap Chinese models even more
  • skerit 8 minutes ago
    > It’s beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on.

    I'm glad people are saying this out loud, because that is what they want. Not for the good of the world, but for the good of their pockets.

  • NewEntryHN 9 minutes ago
    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Because contrarily to the author's assumption, all labs, Western or not, have sufficient skills to discover the optimizations anyway, and publishing or not is not actually that important?

  • wren6991 9 minutes ago
    The doublethink required to simultaneously believe "our safeguards prevent our models from doing unsanctioned cybersecurity tasks" and "distillation is why Chinese models are getting better at cybersecurity tasks" is genuinely quite funny.
  • impossiblefork 13 minutes ago
    Yeah, and Anthropic probably got inspired to this new fast read-in thing for making agentic stuff make more sense from the latest DeepSeek model. Maybe it was in the pipeline, but it clearly has the same effect and DeepSeek had published it by the point Anthropic dropped their prices for reading tokens in, so they may well have copied it.
  • LogicFailsMe 23 minutes ago
    Watch any interview with the Chinese AI leaders and compare it to the unending doomer word salads from America's mightiest paper billionaires. We're losing because we have a loser late stage capitalist scarcity mindset. They're winning because they're sharing notes and one-upping each other just like we used to until 2015 or so. They have a healthy ecosystem of competing small AI startups. We have two bloated unprofitable pigs both striving to be too big to fail. My money's on China for the immediate future.
  • emtel 28 minutes ago
    As far as I can tell, neither of the frontier US labs have referred to distillation as "stealing", but someone please provide a link if I'm wrong.

    They do claim that it violates their ToS, which we can assume is simply correct, since they get to put whatever they want in their ToS.

    Given all that, I don't know what the fuss is. Are they supposed to not use the advances that were openly published by Chinese labs? The entire industry is built on a discovery made at Google, which was published openly. Should Chinese labs therefore not use transformers? Should US labs not try to prevent distillation of their models?

    • jerrygenser 19 minutes ago
      I'm not sure if they don't refer to it as "stealing" but they refer to "distillation attacks"
    • dgellow 25 minutes ago
      I don’t think there is fuss, just the author sharing the information and mentioning how they find it a bit ironic that US labs expenses can be reduced drastically thanks to the Chinese companies they continuously frame as adversaries
  • Reptur 17 minutes ago
    Open releases are just the obvious move when you're not the incumbent. You commoditize the thing your competitors charge for and get distribution you could never buy.
  • amichae2 5 minutes ago
    I am not a fan of Anthropic but this article offers no concrete evidence that Anthropic actually ripped off Deepseek. It is all circumstantial.
  • nater5000 0 minutes ago
    >The new game in town is adopting Chinese labs’ advances. Note how I call this adoption instead of the more vitriol-infused “stealing” that Anthropic tends to use.

    I mean, there's a pretty big difference between labs publishing their research openly and a competitor utilizing it versus a lab breaking TOS to... hmmm, what's the word? steal data from a competitor?

    >That’s because, unlike the Western companies, the Chinese are pretty much giving away their recipes.

    Yeah, Western AI companies have never published their research. It's crazy how the Chinese had to independently develop the foundational technology that powers LLMs because Western companies simply never publish their research (I mean, as long as you ignore stuff like this <https://arxiv.org/abs/1706.03762>).

    >The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that DeepSeek has generously shared with the world.

    Thank you, generous corporation. I'm sorry that other corporations don't provide you free publicity for your selfless contributions to the world.

    >Now I don’t know why they would freely give away such a breakthrough, but they just did

    Well I'm glad the author finally got to their point. A very insightful analysis.

    >They do seem to be a little embarrassed by the copying. Hence the silent releases without much pre-announcement for both Claude Opus 5.5 and GPT-6.1 Sol.

    You have to be in pretty deep to infer this kind of emotion to these kinds of corporate activities.

    >So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Then why write this article? Why point out these things just to have no conclusion?

    This article sucks. Even if you hate US AI labs and are all aboard Chinese labs producing open models, there's nothing of substance here. This is the loose draft that you hand to your LLM to finish for you, but it seems the author just forgot to do so.

    Even if you're willing to characterize US AI labs as evil and selfish and Chinese AI labs as righteous and generous (which is already completely trivializing these dynamics to the extent that anybody over the age of 14 can likely identify is lacking nuance), you can at least put some effort into producing some hypotheses about why these dynamics are occurring. Of course, odds are if the author did try to articulate some hypothesis, they'd likely quickly realize that the narrative they're painting just doesn't hold up.

  • open592 24 minutes ago
    > If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.

    Ah brings back Halo 2 memories

  • LunicLynx 16 minutes ago
    The clue is: Bursting the bubble
  • revexos 4 minutes ago
    too much pace
  • adamrezich 27 minutes ago
    OpenAI is jobbing (in professional wresting terminology) hard right now.
    • brcmthrowaway 19 minutes ago
      So who is the kayfabe?
      • adamrezich 15 minutes ago
        Did you not see the meeting with the President yesterday? All of the “safety discourse” was kayfabe.
  • nba456_ 33 minutes ago
    Appropriate domain name
  • Handy-Man 35 minutes ago
    Just assumptions, nothing backing it. So maybe I'd sit out calling others out.

    Edit: Apt domain.

    • squidbeak 23 minutes ago
      Deepseek's innovations are published as research. There's nothing 'assumed' about this. The common slur repeated in the West that Chinese labs are parasitic distillers is totally absurd when so many genuinely valuable advances and contributions to the field are published openly by China's labs.
    • slowin 28 minutes ago
      How is it just assumptions? They provide the data to back up their claims.
      • Handy-Man 3 minutes ago
        I am talking about correlating Anthropic/OpenAI cache prices going down with Deepseek publication - neither of those labs have said that's what they used for example.

        And the only data they are showing is that cache prices went down for new Claude/OpenAI models but that's proving nothing, IMO.

  • jgrahamc 30 minutes ago
    The first sentence is: "If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse."

    Since I have no idea what "spawn-camped" means I gave up reading the rest.

    • emilecantin 27 minutes ago
      It's from video games, where a player "camps" near the spawn point and kills newly-spawned players, presumably with better equipment.

      It's not that niche, if you've been online a little bit you'd know this expression.

      • tejohnso 23 minutes ago
        > It's not that niche, if you've been online a little bit you'd know this expression.

        No way. You'd need to be pretty well versed in gamer lingo. Even more specifically, combative, likely FPS gamer lingo.

      • jgrahamc 18 minutes ago
        I dunno, man, I first got on the Internet in 1986 and was Cloudflare's CTO for years. I've been online quite a bit.
      • mjc26 11 minutes ago
        Also, it's possible that jgrahamc is the former CTO of cloudflare
      • cassianoleal 24 minutes ago
        > if you've been online a little bit you'd know this expression

        I've been online since circa 1995 (earlier if you count BBSs), and I can't say I did. It's possible to infer its meaning but assuming everyone is on the same circles as one is, is silly.

    • bsoqk 21 minutes ago
      We live in the era of LLMs, which can produce definitions for any word, further explanations, and limitless examples.
    • jamiek88 6 minutes ago
      So you had the opportunity to learn a new phrase but instead discarded the whole text because of something you hadn’t encountered before?

      Do you have many mini tantrums like this per day? Probably makes you very difficult to work with Mr I was a CTO.

    • Intragalactic 24 minutes ago
      If you tend to stop reading every time you encounter something you don't understand, I can't imagine you learn very much

      "spawn-camping" is the process of taking out your enemies at the point they spawn (or appear) in a game without giving them a chance to regroup. In this case I think the writer is saying that the news implies that western models are getting distilled on release. Not the perfect analogy but it gives some color.

    • khazhoux 26 minutes ago
      In first-person shooters, when you die you regenerate (“respawn”) somewhere on the map. If those regeneration points are know to opposing players, they can wait next to them, and kill you again the moment you respawn.

      The premise in this article is: Western companies do a ton of expensive work building new models, meanwhile the Chinese companies just wait for a Western release and then they immediately grab and distill it and announce it as their own model. That’s the spawn-camp.

    • some_furry 29 minutes ago
      It's gamer terminology.

      In a PvP (player vs player) game, if you kill a player the moment they spawn into the game arena, that's called "spawn-camping".