> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.
I remember when Claude identified as ChatGPT a long time ago. It proves nothing else than that there is a lot of training material on the internet with Claude as the AI.
I don't understand why labs don't include correct model name in the training process. Almost nobody seems to be willing to tell their models who they actually are.
If the model works correctly, you don't need to include the model name in the training process. You just add in the system prompt "You are X, trained by Y" and the model will claim to be that. That also allows users to "white-label" the outputs sort to say.
This approach is basically how all "models know who they are" (they don't actually typically "know" that at all), it's just a system prompt instruction in the platform you use.
That only makes sense if you don't sell access to the model for 3rd party product development. One of those products is chat personas/agents. The company integrating the model will want it's own identity.
You can put anything in the prompt. And if you have bare model that you don't give any prompt from the start and want to find out which model you are talking to, you are out of luck.
I think Qwen teaches its models that they are Qwen. Most others don't bother.
Because they do not know the name of the model before they train it. There is also distillation, where multiple models will be trained from a larger one. E.G. Sonnet was promoted to Opus at one point after it surpassed expectations.
Models don't have an inherent identity. It should be obvious by now that every models trains on public AI chat session transcripts. I've seen Claude say it's Qwen.
when open source models are banned maybe i will become a cartel kingpin smuggling open source weights into the usa. find a sufficiently shifty street corner, 'what do you need', "kimi". are you familiar with my product? pure mxfp4, $500 per TB.
(joke and walter mitty, i know the government reads my messages)
I have used Telnyx for, on the opposite end of cool-ness, their fax API. Curious if their AI pricing and quality are actually competitive or if this is just a rapid pivot to try to ride the AI wave.
That is not going to be cheap for long if K3 prices does the same as GLM5.2 prices. If nothing else open weight models are great to get providers to compeete on price.
I think telnyx is a good product, with the only stain to its name being the supply chain attack on their python library.
But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.
And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.
But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.
AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.
Large profitable corporations with very poor customer support have existed long before AI, so why should it be different? Using AI to provide poor customer service is just an implementation detail.
Companies changing their scope or opening side products is not at all unusual. Microsoft is Windows, what's this Xbox thing? Google is search, why do they have email? Y Combinator is a startup accelerator, why'd they make their own Reddit?
Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.
OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.
Seems like a good riddance, telnyx is for business usecases, not for personal use.
Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..
If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.
No? Firstly you can't do a Sybil attack when each identity costs actual resources, and secondly you can still have lots of phone numbers, they just know who each one belongs to.
We estimate that the true blended price per million tokens for running Opus 4.7 on agentic tasks at $0.99 despite the sticker price being $5/$25 per MTok.
They're not saying anything about Anthropic serving costs in that quote, just calculating what MTok price is for running agents. Next sentence after your quote:
> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.
90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.
Please correct me if I'm incorrect, but it seems to me those numbers are describing a fairly different situation to this one. Anthropic serving their own model to their own users at that scale gets cache hit rates and machine utilisation that someone standing up another company's 2.8T model in four regions isn't going to get, and this thing needs 64 accelerators minimum before it will run at all, so a rack sitting mostly idle through a quiet hour costs the same as a busy one. The margin figure quoted is also just the price against the cost of producing the tokens, it doesn't have buying the hardware in it, or depreciation, or maintenance, or the money they'd have made renting those machines out instead, which with rental prices up 40% since October isn't nothing.
I largely agree things are overpriced, I just don't think that article is the right basis for saying it about this one.
Why is this provider specifically on the front page?
Also why is this not available over openrouter?
€2.693/M input €13.464/M output Surprisingly, cache is not mentioned
Upd: Tensorix joined the fray, with the same prices, with cache at €0.673/M read
Or does openrouter have like a specific contract you have to sign with them and requirements or smth? https://openrouter.ai/moonshotai/kimi-k3#providers
Looks like they are first vendor to undercut in price.
> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!
https://imgur.com/a/jqpc2Jc
and
> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.
https://imgur.com/a/AKxeysH
This approach is basically how all "models know who they are" (they don't actually typically "know" that at all), it's just a system prompt instruction in the platform you use.
I think Qwen teaches its models that they are Qwen. Most others don't bother.
(joke and walter mitty, i know the government reads my messages)
The top text says MENA as well, but the regions page doesn't mention it.
But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.
And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.
But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.
AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.
OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.
Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..
If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.
No? Firstly you can't do a Sybil attack when each identity costs actual resources, and secondly you can still have lots of phone numbers, they just know who each one belongs to.
according to Semianalysis those prices would be far above actual costs.
> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.
90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.
300:1 input/output blend: (300 * $0.95 + $25) / 301 = $1.03 per MTok.
They don't say exact cache hit rate they calculated for ("90%+"), so close enough.
I largely agree things are overpriced, I just don't think that article is the right basis for saying it about this one.
Not the **literal** cost