Claude Haiku 5.5

(anthropic.com)

161 points | by sfkgtbor 30 minutes ago

24 comments

  • minimaxir 26 minutes ago
    Pricing is...a bit weird.

        Input
        $0.10 / MTok for prompts up to 100,000 tokens
        $0.50 / MTok for prompts over 100,000 tokens
    
        Output 
        $0.50 / MTok for prompts up to 100,000 tokens
        $2.50 / MTok for prompts over 100,000 tokens
    
    100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

    In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

    • Tiberium 13 minutes ago
      There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

      You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

    • Eridrus 17 minutes ago
      It's actually existing flat per-token pricing that is weird.

      Neither encode nor decode are linear in compute, so providers need to price for average expected length.

      This is just getting closer to the true cost of generating tokens.

    • dannyw 15 minutes ago
      Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

      For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

      These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

      In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

    • HarHarVeryFunny 8 minutes ago
      Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

      For this application 100K token input is plenty.

      Of course Anthropic and OpenAI at $0.10/M are still 2.5x the cost of Jev's $0.04/M.

      • martianvoid 4 minutes ago
        I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification
    • tr4656 22 minutes ago
      Luna does as well, but just at a higher limit.

      From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

    • giancarlostoro 23 minutes ago
      I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

      I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

    • insanitybit 15 minutes ago
      I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
    • AustinDev 14 minutes ago
      encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

      There are plenty of workflows like translations where you'd easily be under the cap.

    • enraged_camel 14 minutes ago
      >> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

      Your vibes don't appear to be supported by facts. From the announcement:

      >> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

      • Philpax 10 minutes ago
        People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.
    • esafak 14 minutes ago
      It's their creative way of 'matching' Luna's prices.
    • j45 19 minutes ago
      It could be to incentivize people to not be lazy users of tokens.
    • system2 15 minutes ago
      Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?
      • user43928 10 minutes ago
        Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.
  • jjcm 1 minute ago
    Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

    Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

    Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

    Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

    One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.

  • charlesabarnes 10 minutes ago
    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users

    This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes

    • tech234a 1 minute ago
      OpenAI will probably add this to their plans within a week
    • geek_at 3 minutes ago
      This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay
  • wyrdcurt 2 minutes ago
    About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
  • seaal 20 minutes ago
    The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.

    Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.

    • 0gs 14 minutes ago
      yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
  • garo-pro 8 minutes ago
    > Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.

    Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5

    • jstummbillig 7 minutes ago
      Maybe they think it's of interest.
  • tpoacher 18 minutes ago
    Good to see Anthropic back alternative OSes.
  • TheAmazingRace 27 minutes ago
    I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?
    • onlyrealcuzzo 14 minutes ago
      Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years.

      You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.

      We haven't yet seen that at any size AFAIK.

      • thefourthchime 10 minutes ago
        Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.
    • istjohn 6 minutes ago
      According to Epoch AI:

      > The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]

      0. https://epoch.ai/publications/the-plunging-price-of-thought

    • bravetraveler 24 minutes ago
      I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.
    • ChaseRensberger 24 minutes ago
      reminds me of this blog post: https://campedersen.com/singularity
    • himata4113 27 minutes ago
      double the information density every 2 days?

      serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.

      • qeternity 22 minutes ago
        Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.
    • dyauspitr 25 minutes ago
      Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.
  • swalsh 12 minutes ago
    Top of the page in 17 minutes? Now I know what y'all do while your agents are working.
  • iagocc 22 minutes ago
    Where is Pelican? ehehhe
    • rvz 17 minutes ago
      This is what AI psychosis has reduced readers to on this site.

      We are witnessing the acceptance of average and accelerating more of the same low quality slop.

      • swalsh 9 minutes ago
        I think it's a joke at this point, but also the visual benchmark is a remarkably dense method for demonstrating how good a model is.
      • InsideOutSanta 13 minutes ago
        So do you have the pelican or no?
  • afrnswrth 18 minutes ago
    The important question though...how does it do making a pelican on a bicycle?
  • sroussey 22 minutes ago
    It’s about time Haiku got an update!
  • patrickwdaly 19 minutes ago
    How are y'all using Haiku though? I rarely select it.
    • swalsh 11 minutes ago
      I've been using GPT-6 Luna in some capacity for nearly all my agent workflows. It's just a really good model, and the pricing is cheap. If Haiku 5.5 is better, and the same price (under 100k context... which is a big caveat) i'd probably swap it.
      • dannyw 6 minutes ago
        It’s absolutely better than Luna. It feels closer to a “sonnet 5.2” if that makes sense.

        Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.

    • mariocesar 13 minutes ago
      I have a zsh functions that calls claude code with haiku to suggest commit messages, is faster and the instructions are two lines.

      I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...

      I use haiku for things that needs to be quick, have really clear instructions.

    • gghootch 11 minutes ago
      I was waiting for this.

      Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them

      ( https://evalship.com )

    • svachalek 16 minutes ago
      Opus often picks it when it's doing a "find me something" subagent. But largely it's been held back by being fully a year old at this point, and priced at a much higher price than models that are far more capable.
    • steve_adams_86 16 minutes ago
      It's great at parsing documents inexpensively. For the few skills/plugins I've made, I usually instruct Claude to use Haiku for low-reasoning grunt work.
  • TomGarden 21 minutes ago
    From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces
  • margorczynski 21 minutes ago
    How does the price compare to Luna? At least looking at the numbers it is noticeably better at most tasks.
    • onlyrealcuzzo 19 minutes ago
      IMO, this is better. Luna is super cheap, but it's not that capable. At higher levels of reasoning, it's not that fast.

      This is more expensive, but it also looks like it's better enough that it's far more useful.

      I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.

      Luna will still be a great option for doing non-engineering tasks super cheaply.

    • TomGarden 18 minutes ago
      For prompts under 100k tokens, it's priced the same as Luna - $0.10 in, $0.50 out.

      For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.

  • sfkgtbor 27 minutes ago
    Happy about the Sonnet cache read price cut.
    • minimaxir 20 minutes ago
      That was effectively required to match GPT-6.1 Sol (costs and caching prices are now equal). Sonnet 5.5 made zero sense to use over Opus 5.5 under the old cache prices.
  • maz1b 24 minutes ago
    Wow, the rate of improvements in the AI era is staggering.

    GDPval-AA v2.1 as of now: 1620

    GDPval-AA v2.1 for Haiku 4.5: 735

    The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.

    Nice release, congrats to Anthropic.

  • simianwords 14 minutes ago
    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, see our Help Center article.

    Did anyone read this? We get free API credits on some plans now

  • crooked-v 8 minutes ago
    The important question is, does it talk in incomprehensible Claude-ese like the other Claude 5.x models?
  • AtNightWeCode 12 minutes ago
    Probably the same scam as the last Haiku update I guess. Uses more tokens to compensate for the lower price.
  • caaqil 15 minutes ago
    > Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.

    If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?

    Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.

  • simianwords 20 minutes ago
    I remember a friend asking me why LLMs suck so bad. She was using Haiku 4.5 and that poor model couldn't keep track of the context within 3 messages.

    She said she was using Haiku 4.5 because she was advised to be careful with the spending.

    I hate that model so much lol.

  • aidiveyt 9 minutes ago
    [flagged]
  • vickyonlinecont 9 minutes ago
    Does anyone still use Haiku model?