16 comments

  • rdli 1 hour ago
    It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

    9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

    • Waterluvian 15 minutes ago
      Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
    • rdli 18 minutes ago
      I would also say: I think this works great because the goal is well-defined and measurable.

      I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.

    • slaser79 26 minutes ago
      Agree Opus 5.5 is incredible and efficient with Claude Max Usage.What a jump after the writing slop you got from Opus 5. I had moved to using Fable for my orchestration workflow mainly due to the communication issue (I think Opus 5 was capable enough but spoke in riddles so you lost confidence quickly)..With 5.5 it needs less steering now and communicates well, and I have had the same CC session running for the last week (obviously compacting with durable plans etc as the post), with it PM'ing my home built agent orchestration of the other coding agents(antigravity, codex, pi etc) and making decent decisions and all the recommendations are normally usually good.
    • tamimio 41 minutes ago
      It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
    • chewchewchew 1 hour ago
      9 hours?!
      • rdli 1 hour ago
        Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
        • Tade0 28 minutes ago
          I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
          • rdli 21 minutes ago
            I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.

            (Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)

            • tripleee 0 minutes ago
              God I hope the prices drop quick. Once they stop subsidizing it these kinds of workflows will be unaffordable for anyone who isn't already wealthy
            • atif089 6 minutes ago
              Does it resume automatically on higher subscriptions?

              I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.

    • whatsThisBtn4 47 minutes ago
      Meh... There's a reason Opus 4.6 is still an option.

      Pros know these are lower cost models.

      • oidar 14 minutes ago
        I do like Opus 4.6, but I think 5.5 on medium or low is a better value. I have to steer 4.6 more and build more scaffolding around the tasks. 5.5 just does what I ask. Visual spatial reasoning greatly improved in 5.5 as well.
    • GroksBarnacles 47 minutes ago
      Is "wall-clock" an actual term you used before Claude? I had never heard it before the model used it and I can't stand it.
      • hazard 44 minutes ago
        It's a pretty old term, to distinguish from e.g. CPU time. This was in common usage even 30 years ago.

        Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...

      • dolebirchwood 38 minutes ago
        Trying to make us feel old? Very common term among people around my age and higher (40+). Please don't turn "I've never heard that expression before" into "must be AI saying it".
      • kgwgk 39 minutes ago
      • sdthjbvuiiijbb 42 minutes ago
        It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
      • woodruffw 44 minutes ago
        “Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
      • Klathmon 42 minutes ago
        I've used wall clock for many years, normally when compared to CPU time when talking about parallelizing some process.

        CPU time might go up while wall clock time goes down

      • pertymcpert 15 minutes ago
        Wow...it's super common. For example, timing a program you can get user CPU time and then wall clock time, which can be two very different numbers.
      • jghn 18 minutes ago
        Lolwut? You seriously have never heard anyone say this before?
    • Betelbuddy 26 minutes ago
      >> It’s a really good model.

      In the meantime, I have cancelled my Anthropic subscription...

      I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.

      I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.

      Then I ask each model to review and critique the other proposals.

      By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.

      Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...

      • karp773 9 minutes ago
        It's not even funny any more. Chinese model, Chinese vendor, Chinese, Chinese... Did I say Chinese? Chinese!

        Nobody in his right mind will use a Chinese clones when you have models like Opus 5.5 for peanuts.

        • Betelbuddy 1 minute ago
          Yes ...I was so impressed I cancelled my subscription. I could not stand all the winning. I offer cheap hourly rates of $1000 for debugging AI slop.

          Contact me at : prompt.plumber@gmail.com

        • verdverm 6 minutes ago
          > Nobody in his right mind will...

          let a few valley elites decide how humanity can use this technology

          open and transparent is the way, China is showing how

          • christophilus 1 minute ago
            I’m rooting for open models, but SOL 6.1 and Opus 5.5 are absolute workhorses on a $100/mo sub. I share your fears, though, and really hope an open model catches up and can somehow compete with the subscription prices of the big 2.
  • kingcauchy 10 minutes ago
    I've had troubles with it getting stuck "waiting" for a day on some hook or something in CC that never completed and during a task that was waiting on an orphaned process. That maybe saves token money on checks waiting for long-running processes but makes it hard to trust for long-horizon work.

    It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.

    It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.

  • hibikir 28 minutes ago
    It's much better than 5, but I've had a couple of situations this week where it was too interested in being independent, making calls that went directly against my recommendations. It can also do fun things like convince auto-mode to go way past what I have autorized. For instance, specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries. And a few of the times it got the calls very wrong, by assuming it understood systems it didn't. It'd even argue with me when corrected, as it assumed similar names were referring to the same thing, when they weren't.

    So asking it to do things on its own for a long time? Given last week, absolutely not.

    • istjohn 12 minutes ago
      Otoh, I've had it push back against my false assumptions when I was confidently incorrect, finding proof unasked.
  • ToJans 45 minutes ago
    Superb model indeed.

    I've given it some big tasks and asked it to parallelize as much as possible etc.

    It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)

    I've now come to a point where I even delegate my discovery for new features to it.

    You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.

  • alwinaugustin 5 minutes ago
    Can Anthropic introduce a $50 USD Plan ? I dont want to spend 100$ , but $20 plan is not enough for me.
  • magicalhippo 25 minutes ago
    Been very impressed with my most recent project. I wanted to simulate some older electronics circuits. I handed it a folder with scans of old service manuals which contained circuit diagrams. It managed to correctly interpret the circuits, including figuring out some were the same topology despite the diagrams being quite different, or some that had some subtle but very important differences despite looking almost identical at a glance.

    In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.

    It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.

    I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.

    It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.

    I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.

    I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.

    Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.

    I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.

  • TomGarden 14 minutes ago
    Incredible model. I don't see why they can't just include the recommended workstyle as a guided approach into the claude code harness though, and let the people who want to diverge just ignore it
    • verdverm 8 minutes ago
      > and let the people who want to diverge...

      the company is run by holier-than-thou, we know what's best... who apparently don't read claude's output and blindly trust it

      the mythos "hacking" of the linux kernel, as finally told from the linux side, is eye opening

      https://www.youtube.com/watch?v=NnV_cWeoo5Q

  • briga 28 minutes ago
    Is it still necessary to ask Claude to spin up sub-agents? If Opus 5.5 decides how carefully it needs to think after each question, surely it can also decide whether it needs to spin up sub-agents? GPT 6 series models at least seem to do this agent management automatically.
  • alansaber 50 minutes ago
    "Don’t ask it to show its reasoning in the reply" “Explain why you chose this approach in three sentences” says it all really
  • epistasis 26 minutes ago
    I really hate long tasks. Claude never gets things right, at least for me, and wastes tons of time when a simple question would have gotten me to the right result rather than several turns of correcting bad decisions in addition to the long amounts of wasted thinking time.

    What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.

    Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.

    Ah well, the Chinese models will still work well, I guess.

    • ricardobeat 1 minute ago
      Plan mode has become pointless since Opus 5 came out, they know when to switch between planning and execution now. But that iteration/discussion is still necessary unless you're building completely blind - the model cannot read your mind.

      I've had it running 8h+ of non-stop optimizations, chasing a performance target or rewriting systems. All it needs is a clear goal.

    • istjohn 8 minutes ago
      Look up the grilling skill[0]. To make it even better, tell Claude to modify it to use the AskUserQuestion functionality. It's so much better than plan mode.

      0. https://github.com/mattpocock/skills/blob/main/skills/produc...

  • danbrooks 1 hour ago
    Agreed on Opus 5.5 being a great model. It's the first one that I trust for long running (>1 hour) tasks.
    • jester997 48 minutes ago
      It’s amazing at debugging too. I had it running in Powershell controlling an lldb session in MSYS2. The way it can read addresses and so root cause analysis is amazing! It takes a lot of mind power to do those things.

      I like to watch it work though because it honestly teaches me some tricks.

    • whatsThisBtn4 46 minutes ago
      I'm convinced gpt 4.5 was the best model humanity has ever made.

      Now we are getting downgraded models that do 100x COT because it's cheaper.

  • tebrun 46 minutes ago
    Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)
  • sergiotapia 33 minutes ago
    The best model I've ever used easily. It's incredible. I've done so much in the past week. About three months of work I reckon.
  • Handy-Man 1 hour ago
    Phenomenal model, not sure what they did, but I have been able to do so much with my $20 plan!
    • amelius 58 minutes ago
      I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
    • yfontana 37 minutes ago
      A significant part of it is that with Opus 5.5, cache reads are priced at 5% of inputs, instead of 10% for previous models.
    • 233mhz 1 hour ago
      If enough people keep saying it I'm sure they will nerf it
  • franze 1 hour ago
    [flagged]
  • i_love_retros 1 hour ago
    [flagged]