Seeing a lot of tricks similar to how ridesharing companies tried to be "profitable" before going to IPO. Caveat: Thing have materially improved but really Uber is carried by its insane Ads margins
The idea of removing model training from your costs is a little wild tbh.
The profitability of being able to serve a query wasn't really under question (nor is the margin expected to be anything less than 80%+) I think.
> The idea of removing model training from your costs is a little wild tbh.
Yeah, I didn't believe they'd claim something like that. But yes indeed, from the article:
> Anthropic's gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), opens new tab, and the cost of training its model,
Is this how all AI companies calculate if they're profitable or not, by removing the highest costs? What a circus.
> Anthropic has told shareholders that its adjusted operating income will be positive for a second straight quarter, the Financial Times reported on Sunday, citing multiple people with knowledge of the matter.
Note this claim is about „operating profit“, which commonly is the revenue - operating expenses (COGS, rent, payroll). This does not include RnD cost.
>Anthropic's gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), and the cost of training its model, the newspaper said.
Gross margin is typically (revenue - COGS) / revenue. Thus, both statements above seem generally in line with commonly accepted accounting standards.
Typically R&D is not part COGS. It’s absolutely part of the bottom line when you will typically recognize the cost over some period of time to try to get a true picture of the business.
It’s also easier to strip it out of the picture to think about how much it costs to serve the next token. If you can have great economics to serve the next token (profitable) you can always figure out ways to further reduce your R&D costs.
Now they are absolutely intertwined but I don’t think this is ever as big of an issue that people make it out to be. Replace token with any widget, this is how businesses measure themselves.
The way I read it, they are convincing the investors that they can fool the larger population convincingly. At the end of the day the investor term is misnomer for big institutional investors, given that these people are managing other people money where they always make out a certain percentage of fees despite the outcome.
I think part of the big push to "slow down AI development" is to add some sort of regulatory pressure that will give them sort cover to train less models and slow their burn rates
I still question it. It's very light on details by just casually dropping they have 80% margins. I smell bullshit on this.
It is certainly not above them to play accounting tricks to pretend to be anywhere near profitable.
If you create a machine that can turn a dollar into 5, you don't dillute ownership of the machine, you use your fabulous profits to expand production. Anthropic, on the other hand, raises money like crazy, and seems desperate to IPO.
I can’t count the number of times I’ve heard variants of ‘they’re losing money on every query’ and ‘I’m getting 10k worth of tokens for $200’ over the last year. People clearly believed serving margins were -ve
> idea of removing model training from your costs is a little wild tbh
It's one of several metrics and tries to estimate steady-state profitability. It's the only one being leaked because it's the most sensational one. But don't assume cash-flow profitability is negative just because you don't know it.
Well then you haven't listened to Ed Zitron or any of the other AI bubble doomers. His contention is that its worthless and they lose money on every query.
But isn't he taking all the costs into account, that created the experience? Rather than just literally the inference/serving infrastructure? Bananas way of calculating things if so, doesn't match reality at all.
Zitron is worthless–lying about numbers and not correcting the record when you're called out means you aren't trustworthy. Worse than that if you directionally agree with him, which I do.
Is it really capex though? New models are being constantly trained, released at least quarterly, while old ones become obsolete. Training costs vary, but never disappear.
New models are constantly being trained, but they don't have to be constanatly trained. If OpenAI or Anthropic 'just' wanted to be a profitable business, they could ease up on that, but they're both racing to create a machine to can automate all or most human labour.
If they stop racing, the wave of open models will pass them and their inference margins will drop to zero. A realistic model of their operating costs surely must include ongoing training.
>If they stop racing, the wave of open models will pass them and their inference margins will drop to zero.
1. ChatGPT's ~billion weekly active users aren't going to give a shit about some open source model, and neither would most of Anthropic's Enterprise cutomers.
2. Open AI and Anthropic are in a race between themselves, not open source model trainers. There's a reason those models are consistently several months behind and often perform much worse than benchmarks indicate. In the first place, they're only as close as they currently are from the distillation attacks on Anthropic and OpenAI. If they slowed down, they would slow down too.
1. They absolutely would, eventually, if open models surpassed Ant & OAI's flagships. Keep in mind "surpassed" encompasses both output quality and cost-saving architecture innovations like DSA, which may not be possible to apply to old models (and may take advantage of new hardware!)
2. "They're only as close..." is not natural law. You really think open weights couldn't catch up to a fixed target if Beijing makes it a priority? And what happens to their valuations if they abandon the goal of building AGI? There is no strategic alternative to constant training for these companies, which is why they're, uh, constantly training.
1. Mainstream users don't care about benchmarks or whether some open weight model has technically surpassed GPT-X on a leaderboard. They care about whther GPT does what they want it to do. Capable Open source models already exist, and that hasn't caused ordinary chatGPT users to abandon chatGPT for them. Hell Anthropic exists, and that didn't cause that either. OpenAI still dwarfs Anthropic in the consumer space. Obviously, sufficiently large differences in capability can eventually matter like when Anthropic blew everyone away in coding at one point, but that's very different from saying OpenAI has to train a frontier model every few months or inference margins go to zero.
2. Nobody said anything about a fixed target. Not sure why you interpreted 'slow down' as 'freeze current models forever'.
>And what happens to their valuations if they abandon the goal of building AGI?
The capabilities these companies already have, combined with their growing userbases, revenue and distribution are plausibly enough to sustain trillion dollar businesses already. OpenAI is a company with a billion active users that has started running ads that reached ARR of $1 billion in the first 2 months and Anthropic is a company that hit $11B+ in revenue last quarter after a pretty massive jump.
Enterprises absolutely care about benchmarks (especially internal ones, but the headline benchmaxxed ones too), and the Ant coding thing is a great example. How long did that last again? A few months? Illustrates my point perfectly. Switching is easy. Why would a business have any loyalty to one text->text endpoint over another? The consumer market may be less responsive to quality, sure, but it is more responsive to cost which I mentioned. It's also just not as big.
Well, I'm not really talking about freezing models forever either, I'm saying that nonstop training is a necessary part of their business. I don't think slowing down is untenable, I just think it's silly not to expect & account for ongoing training costs. That's all my original comment meant.
I also don't understand why you think the open labs couldn't catch up to a given level of quality. If something's been done twice already, why can't a well funded team of experts somewhere else do it a third time? Sounds like wishful thinking.
Isn't training necessary for the end-product? How is it not a cost to generate the output if you can't generate the output without having done the training? Seems more like saying that a car is profitable product when you don't have to account for the steel that its made from. Seems completely disingenuous.
Ok, but a more honest comparison here would be removing the cost of hiring architects and hiring lawyers to draft contracts, not all the per-house costs.
Yes, but --- using something they call "adjusted operating income".
This is reportedly a sort of "Enron" accounting which excludes some really big expenses like revenue sharing, the cost of model training and hardware deploymments which are kept off the corporate balance sheet using "special finance vehicles".
>This is reportedly a sort of "Enron" accounting which excludes some really big expenses like revenue sharing, the cost of model training and hardware deploymments
source? this seems false. reportedly the adjusted profitability includes inference and amortized training costs
> before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), opens new tab, and the cost of training its model, the newspaper said.
"We are profitable when we ignore our costs".
I wonder what other funny strategy they may employ to claim 80% margins.
> Anthropic's gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), opens new tab, and the cost of training its model, the newspaper said.
Yes the company known for famously training 1 model
The “opens new tab” sometimes drives me insane when using TTS to listen to Reuters articles. I’m assuming they are using the wrong CSS/markup for the purpose.
Or they've hit diminishing returns that will collapse their valuation, so saying "this is a threat to humanity" sounds better than "This is about as good as the tech is going to be for a long time" to investors.
That's a bit hyperbolic. It's closer to using EBITDA as your "earnings" and bucketing model costs in a rapid depreciation model (which is fair, I'd assume a model is good for more than just 1 year...
The models don’t exist without training. I don’t see how excluding the training cost from the thing they are selling(inference) is perfectly reasonable and not just an accounting trick.
If I am building a widget and have a widget factory that cost money to build and operate, is it reasonable to only use the cost of shipping my widgets to my buyer as the costs for my gross margin?
Generally speaking, IDK about Anthropic specifically, they don't train purely from scratch though. A good chunk of the setup is reused previous models can serve as a basis for the next model. Plus there are methods that also use the previous model as a warm start.
Just taking a wild guess, but I'd assume the .5 releases are built on the previous and the Majors (3, 4, 5) are more extensive retrains?
> If the article is to be believed they aren’t including their training costs
In this metric. Revenues aren't a scam because they don't include costs.
We're getting a partial picture. But concluding adversely based on selective disclosure that isn't in control of the person disclosing doesn't make a lot of sense. Especially when it's a compelling narrative.
> models don’t exist without training. I don’t see how excluding the training cost from the thing they are selling(inference) is perfectly reasonable
Planes and cars take a lot of capital to develop. That's a fixed cost. Building them consumes labour and resources. Those are variable costs. It's useful to know the slop of the variable-cost function.
This concern trolling on behalf of people that can't read to the end of a sentence, about a conversation that they aren't even part of.
Anthropic is clear on what they are communicating. If every message had to be dumbed down to the level of the least attentive person to run across a message third hand, we could communicate and nothing but grunts.
I would like them to communicate something else though (ie reveal what their training costs are, how much revenue share they have and with who) so I could have some idea of their actual profit margins. They're not public yet so they're within rights to withhold that, but it does mean this isn't very useful information for a lot of questions one might have?
The idea of removing model training from your costs is a little wild tbh.
The profitability of being able to serve a query wasn't really under question (nor is the margin expected to be anything less than 80%+) I think.
Yeah, I didn't believe they'd claim something like that. But yes indeed, from the article:
> Anthropic's gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), opens new tab, and the cost of training its model,
Is this how all AI companies calculate if they're profitable or not, by removing the highest costs? What a circus.
Note this claim is about „operating profit“, which commonly is the revenue - operating expenses (COGS, rent, payroll). This does not include RnD cost.
>Anthropic's gross margins are above 80% before accounting for revenue shared with distribution partners, including Amazon (AMZN.O), and the cost of training its model, the newspaper said.
Gross margin is typically (revenue - COGS) / revenue. Thus, both statements above seem generally in line with commonly accepted accounting standards.
It’s also easier to strip it out of the picture to think about how much it costs to serve the next token. If you can have great economics to serve the next token (profitable) you can always figure out ways to further reduce your R&D costs.
Now they are absolutely intertwined but I don’t think this is ever as big of an issue that people make it out to be. Replace token with any widget, this is how businesses measure themselves.
In one sense, yes, but I do see people question it regularly.
It is certainly not above them to play accounting tricks to pretend to be anywhere near profitable.
If you create a machine that can turn a dollar into 5, you don't dillute ownership of the machine, you use your fabulous profits to expand production. Anthropic, on the other hand, raises money like crazy, and seems desperate to IPO.
HN had long debates about whether AI inference could even be affordable from a compute perspective.
It's one of several metrics and tries to estimate steady-state profitability. It's the only one being leaked because it's the most sensational one. But don't assume cash-flow profitability is negative just because you don't know it.
And either way, the training of new base models will eventually slow from the current frantic pace.
Which unfortunately probably hides the real truth. That large labs do have potential problems with long term profitability.
This is an open question!
How is this different from planes or cars? If Ford or Boeing zero line their R&D...well, we know what happens.
1. ChatGPT's ~billion weekly active users aren't going to give a shit about some open source model, and neither would most of Anthropic's Enterprise cutomers.
2. Open AI and Anthropic are in a race between themselves, not open source model trainers. There's a reason those models are consistently several months behind and often perform much worse than benchmarks indicate. In the first place, they're only as close as they currently are from the distillation attacks on Anthropic and OpenAI. If they slowed down, they would slow down too.
2. "They're only as close..." is not natural law. You really think open weights couldn't catch up to a fixed target if Beijing makes it a priority? And what happens to their valuations if they abandon the goal of building AGI? There is no strategic alternative to constant training for these companies, which is why they're, uh, constantly training.
2. Nobody said anything about a fixed target. Not sure why you interpreted 'slow down' as 'freeze current models forever'.
>And what happens to their valuations if they abandon the goal of building AGI?
The capabilities these companies already have, combined with their growing userbases, revenue and distribution are plausibly enough to sustain trillion dollar businesses already. OpenAI is a company with a billion active users that has started running ads that reached ARR of $1 billion in the first 2 months and Anthropic is a company that hit $11B+ in revenue last quarter after a pretty massive jump.
Well, I'm not really talking about freezing models forever either, I'm saying that nonstop training is a necessary part of their business. I don't think slowing down is untenable, I just think it's silly not to expect & account for ongoing training costs. That's all my original comment meant.
I also don't understand why you think the open labs couldn't catch up to a given level of quality. If something's been done twice already, why can't a well funded team of experts somewhere else do it a third time? Sounds like wishful thinking.
In your car example the training is much like setting up the manufacturing line.
I think the issue here is that the capex depreciates super fast since the models obsolete really fast.
Active competition requires constant reinvestment and does not allow them to milk their trained models long enough (except poor Haiku maybe).
This is reportedly a sort of "Enron" accounting which excludes some really big expenses like revenue sharing, the cost of model training and hardware deploymments which are kept off the corporate balance sheet using "special finance vehicles".
https://www.msn.com/en-us/technology/artificial-intelligence...
source? this seems false. reportedly the adjusted profitability includes inference and amortized training costs
Listed at the end of my post.
this seems false.
Source showing this in accordance with GAAP (Generally Acceptable Accounting Practices)?
edit: def not gaap profitable or they would have said that to investors. and their stock-based comp is surely astronomically high on paper.
"We are profitable when we ignore our costs".
I wonder what other funny strategy they may employ to claim 80% margins.
They'd be in their quiet period...
Yes the company known for famously training 1 model
If I am building a widget and have a widget factory that cost money to build and operate, is it reasonable to only use the cost of shipping my widgets to my buyer as the costs for my gross margin?
Just taking a wild guess, but I'd assume the .5 releases are built on the previous and the Majors (3, 4, 5) are more extensive retrains?
If the article is to be believed they aren’t including their training costs.
Also lol at reporting it as above 80% without accounting for the revenue sharing as well.
I bet anyone’s finances look great if you just start ignoring all the money they owe.
In this metric. Revenues aren't a scam because they don't include costs.
We're getting a partial picture. But concluding adversely based on selective disclosure that isn't in control of the person disclosing doesn't make a lot of sense. Especially when it's a compelling narrative.
Lol and truth.
Planes and cars take a lot of capital to develop. That's a fixed cost. Building them consumes labour and resources. Those are variable costs. It's useful to know the slop of the variable-cost function.
Anthropic is clear on what they are communicating. If every message had to be dumbed down to the level of the least attentive person to run across a message third hand, we could communicate and nothing but grunts.
Then wait for the S-1. They aren't communicating internal financials to you.