Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
The Claude models definitely felt more susceptible to moods, like you could leave them for a few hours, come back and it suddenly was unable to do things which it was doing just earlier, which tellingly is never an experience I've had with an open model.
Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.
Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
Dollars to donuts, they are speculating, and not privy to inside information on the topic.
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
If you follow reddit forums for claude code, its common to see people, on the same day, claiming that Opus/Fable is especially smart today, and especially dumb today.
I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.
If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.
But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc
I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".
That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.
They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.
But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.
For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"
I had to tell it to ssh into the server and run journlctl to check it
Anecdotal, I know, but they all seem to be less capable with time.
Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.
As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.
Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.
Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
And here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.
They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
Reminds me of how slot machine users swear the odds have changed on a machine.
also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order
The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.
I think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output).
Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.
But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets.
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.
Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.
Buy decent coffee instead of your $200 subscription and sidestep all the scams.
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?
Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.
Well sorry, still have to get decent AI somewhere. Productivity without AI is about 5x less. I am not comfortable with paying Chinese companies, and no Western companies provide subscription-based pricing for open models.
It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.
Makes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.
So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?
This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks.
Not just openai & anthropic, popular openrouter models too.
Tests their intelligence, not their diligence.
Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.
The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
I work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid.
It’s a slot machine for what they’re actually giving us behind the opaque paywalls.
The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
The ROI just isn't there. It feels like Fable is in the same place Opus was early last year; at best marginal improvement that's barely noticeable over the lower model, for 10x the cost. I'm sure it'll take over as the workhorse as Opus did once they get it down, but right now it just doesn't make sense
I stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better.
This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam.
If I pay for Fable, I should get full, not nerfed model at honest pricing.
Regulators should investigate them.
OpenAI is no different. Astra has basically the same problem.
Good point. I suppose watching the number go up is useful information in itself.
I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)
I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
I also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.
This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.
The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open.
(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)
Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too.
I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.
EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.
Opus 5.5 is being served under opus 5 right now.
On what basis are you claiming this?
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
Wow. Anecdotally, it's been munching tokens like doomers were right (ie. there was "no tomorrow" ...)
I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.
If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.
But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc
That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.
But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.
For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"
I had to tell it to ssh into the server and run journlctl to check it
Anecdotal, I know, but they all seem to be less capable with time.
_edit_ I use the same reasoning level of `medium`
As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.
Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.
Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.
I wonder what their official explanation for this behavior is.
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...
But it's still happening: https://github.com/anthropics/claude-code/issues/81759
At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order
The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.
Buy decent coffee instead of your $200 subscription and sidestep all the scams.
Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.
1: "Our model will bring about the end of all things. Flee, flee for your lives"
2: "Our model is basically AGI"
3: "Our model will be available in limited release next week"
4: "Everybody who subscribes at the $200 level gets access now"
5: "Everybody who subscribes at the $20 level gets access now"
6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"
I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.
Tests their intelligence, not their diligence.
Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.
[0] https://aistupidlevel.info/
[1] https://marginlab.ai/trackers/claude-code/
Its similar to other data services I see around my F500 company.
It’s a slot machine for what they’re actually giving us behind the opaque paywalls.
Yes, I’m on a business subscription plan.
This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam.
If I pay for Fable, I should get full, not nerfed model at honest pricing.
Regulators should investigate them.
OpenAI is no different. Astra has basically the same problem.
I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)
(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)
I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.
EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.