> Claude Opus 5.5 is our first release since we called for pacing the frontier.
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
I think it's perfectly clear. It means improvements shouldn't simply advance as fast as possible and more specifically, it's a reference to a past statement of theirs to that effect. At that level of generality it's as clear as it needs to be.
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
Is the meaning clear? Nobody would use "pacing" in this way. I only know what it means because I've seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
Granted I am not a native English speaker but I have no idea what "pace the frontier" means. When I read it I just assumed "frontier" refers to "top of models" and "pacing" is that they are getting there quick.
Is this the meaning or do I have it wrong? I have not checked.
also non native.
but since "pace yourself" means to control your speed, energy, or workload so you do not get too tired or stressed before you finish ->
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
"Pace yourself" is a semi-common English idiom (rarely conjugated, usually an imperative.) It is a gentle way of telling someone not to run/work/eat too quickly, and is typically said when you are concerned they may hurt themselves due to acting hastily.
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
It makes sense to me. If you 'pace your running', you're setting the speed intentionally. The phrasing doesn't describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal.
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and 'the frontier' is pretty much the shortest, clearest way to say 'state of the art development of AI'.
Right, except they mean the exact opposite in this case: they're actually advocating for slowing down the pace. Hence the criticism of the language, because your interpretation would be completely valid.
What does the pace car in a race do? Aka safety car? Or the pacesetter for a marathon. It's a perfectly cromulent use of the word. It's like when LLMs using six dollar words like delve. Some people have better diction than others, and it turns out that AI has read the whole dictionary. Anti-intellectualism is alive and well so we have to dumb things down to sound human rather than come across as smart/AI.
I suppose it depends on your life experiences. In running, someone who paces the group or a pace car is meant to keep the pack progressing at a constant, predictable speed. So in that way, it makes sense to me
Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.
Agreed. This feels like pointless navel-gazing and a strange new language policing, and not something I look forward to. It's tiring and it lacks curiosity (e.g. "what's behind the affinity for the words elevated from wherever it is they come from?").
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
The words have the very practical problem of not communicating anything of substance. I guess "we want regulatory capture" didn't have quite the same ring.
How do they not communicate clearly? A few weeks ago they stated that they would like to regulate the pace of frontier model development/releases, and they reminded us of that in this post.
You must be pretty new to (not just) online discussions if you consider language policing to be a new phenomenon :)
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
> I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
What? How am I doing that. Claude writing style and the discussion at hand are two entirely separate issues that I haven't conflated at all. How am I "acting like people are being grammar nazis"?
What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
It is that, but it's also your speed. When running, you typically pick a pace that's slower than your max, so it can be sustained for the full distance.
Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm.
Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
What does model naming have to do with pacing or not?
This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.
"The improvements are underwhelming for a model that, according to previous claims, should have replaced all knowledge work twice by now, but you know, it's just because we're pacing the frontier."
They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).
When they talk about pacing, they are referring to their dangerous competitors, particularly those evil open-sores and Chinese ones, not their lovely safe models because you can trust them to look after your interests.
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.
You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
>and potentially about Anthropic future profitability too
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
Speaking for myself, I have not been able to use Opus as much as I’d want because its verbose prose makes human reviews of its assumptions, architecture proposals etc. more painful than its predecessors. If they’ve solved that, I’ll be accelerating through my backlog that much faster, and using tokens accordingly.
I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive
In my mix it's usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I've seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude/Codex, that's a reasonable approach for someone like me, a "gentleman code farmer", lol. And even lower tier stuff, I have the local models taking care of.
Now that Jev is out I can finally have a true AI sysadmins managing my "cloud in the basement" homelab at the cost of electricity, which is not cheap btw
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I've don't think Astra is a better model, but it's the first OpenAI one that seemed good enough for me. Definitely keen to try Opus 5.5 and see if this claim is real.
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow:
1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT.
2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot.
3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?
Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.
It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.md
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
I love the contrast with yesterday's open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.
It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.
It's less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra's hero/scrolling animation. I'm pleased to see they didn't make the entire page that like OpenAI did, but I'm expecting to encounter this pattern more often on these announcements now.
Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.
Parallax scrolling effects were very cool ~2010. By 2015 or maybe earlier it already felt like me-too design that's unoriginal and a little annoying. By 2020 everyone and their mom has it and it's super tiresome. Now it just screams slop design (among a million other signals).
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.
The problem with 5 wasn't just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I'm okay with verbosity if it's actually readable.
It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use.
IME, if you eyeball how many words the thing they're trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic's grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It's terrible.
> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%
So longer threads get cheaper and one-shots stay the same price.
Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.
Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.
Benchmarks often don't survive contact with reality.
Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
That's not my experience at all. Opus is an extremely capable coder on high or xhigh effort. It can read academic papers, implement algorithms from the description in the paper alone and reproduce results without breaking a sweat. This is remarkable because it is pure reasoning on unseen material; in some cases the paper was just published and there wasn't an implementation to learn from in the training data.
did you actually verify that it's output in those scenarios is good ? in my experience opus has been a disappointment and constantly trailing behind actually solving hard problems versus the OpenAI models.
I'll say that both have terrible writing style though.
Mines pretty much inverted - my colleagues and I noticed almost zero difference between the quality of code in Opus vs Fable. Occasionally I'll switch to Fable for an arduous debugging task but that's about it.
typically, when any AI company says a model performs as well as fable, all they're really telling us is that the benchmarks that exist for measuring AI capabilities aren't very good.
Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.
It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
I should find information about the user's concern instead of just assuming.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
And how was the Xbox 360 naming choice a “debacle”, exactly?
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.
Maybe this model can finally figure it out for them.
I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.
> Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
I'm confused how they have been able to create so much public negative perception around distillation. It seems pretty clear that they are the only ones who lose out, and everyone else benefits. I don't have any ethical issues with it, nor is it illegal: at worst it's a ToS violation.
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
I agree. Hopefully Anthropic has fixed Opus' ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.
I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate
(perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?
No, it's just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
> also, what "if you're reviewing this rather than just reading it" even means?
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.
How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.
Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.
Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient
They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we
recommend accessing our models with reduced cyber blocking classifiers via our Cyber
Verification Program. Claude Opus 5.5 will be available through this program in the near
future.
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).
Yay, yet another model I can't use for anything interesting, even with CVP.
So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.
People aren't the target audience of that part of the post. They're hoping saying enough safety stuff will ward off the looming regulatory sledgehammer.
Anthropic is one of the most valuable companies in the world. Their comms are designed to appeal to a very wide readership.
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.
I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.
I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.
> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.
We can't test it properly because it knows it's being tested.
I think this release is really going to give them a hard time selling Fable:
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
"benchmark margins have become a less
reliable guide to real-world differences"
sounds like a big problem.
My guesses:
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
Opus 5 fucking sucks. Like it's horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don't want to burn as much quota.
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
My bet is anthropic has NN people org who work hard to distill open models in addition to trying to find what other useful materials they can download from shady torrents.
"bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.
Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!
my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.
> In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
> It’s good at finding and fixing inefficiencies in software
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
Great that they listened! The improvement in communication style looks fantastic. Opus 5 was insufferable and I was on the verge of cancelling my subscription.
So it beats Fable 5.1, by quite a bit, on every metric? Interesting.
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
Why can't they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.
is OPUS 5.5 still not reading CLAUDE.md, failing to follow told tasks, inventing and hallucionating, just refusing to read files ("read the whole file" -> read 2-lines -> infere its wrong -> destroy the codebase), needing constant babysitting just because its so UTTERLY DUMB! i cant imagine going back to OPUS 5 - i'll rather jump out of the window as to use it EVER AGAIN!!
Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
Is this the meaning or do I have it wrong? I have not checked.
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal.
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
sir, this is a hacker news thread
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
"Pace yourself" specifically means "slow down".
But stating it plainly like this would make the contradiction too obvious.
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
doesn't sound like a razor at all
that, they fully intend to 'pace'. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
Simply make them something that derives a text response from its training data.
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
It does work out to be a similar cost per task though
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
For long running tasks it is. That's what made Deepseek so cheap.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I've don't think Astra is a better model, but it's the first OpenAI one that seemed good enough for me. Definitely keen to try Opus 5.5 and see if this claim is real.
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
Giving moral lecture is different than reality i guess.
ChatGPT 6 Pro answered it without issue.
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
I agree that it's sort of stupid, not a fan.
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
Everyone that doesn't gets served some animated bullshit.
Yep, you're right. I tried on my phone and got the scroll through image.
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
God I hope so
https://github.com/AminBlg/SimpleEnglish
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
So longer threads get cheaper and one-shots stay the same price.
Nice. I was starting to think that Haiku got abandoned.
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
Benchmarks often don't survive contact with reality.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
Can you give more details here? This sounds intriguing.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
This is news to me. Excited to try it out! Thanks.
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
Maybe this model can finally figure it out for them.
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
[1]: https://github.com/p-e-w/heretic
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.
Claude is just comically bad nowadays.
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
Not efficiency in writing, clearly.
Thank you.
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
Yay, yet another model I can't use for anything interesting, even with CVP.
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
[1] https://kizi.to/claude-talks-too-much/
I tried Opus 5 and Astra.
Nice. I was starting to think Haiku was going to be abandoned.
We can't test it properly because it knows it's being tested.
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
Infomercial at its best.
No wonder we are hammered with ai announcements.
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
Less companies involved means less pressure to go fast.
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
Yup. As unusable as Fable 5.1, for assembly on mid-80s 68k platform. Awful.
[1] - https://news.ycombinator.com/item?id=6372466