Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
These shouldn't even exist without a negotiated contract.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Monthly electricity bills are based on usage. The difference is the relative orders of magnitude you can be charged for these services you can go from 20$/month to 200k/month without warning.
> In an ideal world, our agents could help with this.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry
Having a monthly summary or estimate of how your spending is going would be really useful, too.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”
Counterpoint: if you can automate API calls on the client side, why can't you automate billing caps? If you want a machine that can run 24-7 and make money for you while you sleep (which let's face it is the motivation for a lot of AI takeup), isn't the onus on you to install cicuit-breakers?
Because many cloud services have incredibly complex or opaque pricing structures that make it difficult to impossible to determine how much something is going to cost you ahead of time, especially if it's usage-based a la network egress (and the usage statistics don't update frequently enough to make such circuit breakers possible to implement client-side).
They might not be able to predict your bill but how much time do they need to add up what you already spent to minimize your overage? And TBH how much time should be acceptable to exceed your cap before it's their fault for the lag in their software.
I would just not sign up for a service without price transparency, or pre-calculate my liability based on available information before pushing the (metaphorical) Deliver Now button.
I'm not sure where you got the impression that I'm making excuses for cloud providers. I'm just stating the way things are, not the way I think they should be.
Making incredibly complex and opaque pricing structures is not necessary for the providers to charge for and make a profit on their service. And being technically difficult is a lazy excuse. Cloud platforms have to solve many, much more difficult challenges to offer their services at all, they just don’t want to invest the time in more customer friendly billing because they expect it will result in reduced revenues.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.
I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
Ubicloud does not have hard budget caps, which I only realized this morning after moving all my CI over to them over the past few months. Fortunately I didn't learn the hard way.
I understand this is snark, but if you think about it, this is already implemented in electrical infrastructure. If I use too much power, the circuit breaker trips to protect me and protect the electrical grid. OP is about a billing breaker, but the parallels should be obvious.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
Why do people think new laws are needed to solve every last problem in the world?
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
Yes and no. I suspect many of the hard limits were set arbitrarily, and we'll see a relaxation of limits as people get frustrated with the limited use they get out of them. And some services will genuinely need to be re written to support higher rps or risk losing customers
I'd support this provided we have the converse as well: if the customer doesn't pay their bill on time, the service gets shut down immediately. (Disclosure: I sell SaaS services to people who don't pay their bills on time).
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
the premise seems a bit faulty to me. why should we be giving next token predictors access to spend our money? like what great benefit do we get from this that we should allow them unfettered access, but with safeguards in the form of hard budget caps?
I don't think Simon means you should hand off the spending to agents/LLM (which would also make me uneasy) but that if you're probing one for hosting/SaaS providers they should default to recommending ones with budget caps
Because it’s unlikely they’ll actually be able to collect that million dollars from a lot of those customers. Rephrased: why would your vendor want to make it harder to accidentally give you a million dollars of services in exchange for debt of dubious quality?
https://cloud.google.com/blog/topics/cost-management/new-ear...
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
[1] https://github.com/tkgally/je-dict-1
[2] https://github.com/tkgally/eex-dict
It works for some non-B2Bs.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
Seriously. You all asked for this.
I can’t think of anyone saying they would hate for AWS to support hard spending caps.
That said, the limits are accurate with Claude to a handful of cents over the last few months. I just don’t see it as a big deal.
Our Gemini API usage is also similarly accurate.
The LLM costs are way more accurate than the limits on our GCP env as a whole.
If you don’t have a mechanism for enforcing hard caps, you don’t get to send customers a bill for unlimited amounts.
Then the irony of whether or not your capped SaaS proxy SaaS product itself has hard caps.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
Or if there was info that was after potentially horrible expensive operation.
That was blocking automated checking.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
What on earth is Simon whittering on about?
This may happen also without vibecoding.
Article seemed clear on that?