Remix.run Logo
mcintyre1994 2 days ago

> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.

invalidusernam3 15 hours ago | parent | next [-]

I would happily pay more for a claude model that performs the same but speaks normal English. The proliferation of claudespeak in the workplace is driving me insane. Every ticket, every PR feedback, every comment in the codebase is poisoned with its ridiculous unnatural vocabulary

ascendantlogic a day ago | parent | prev | next [-]

So far it seems the same. I used Opus 5.5 for an hour this evening and it was just as painfully verbose as Opus 5. It also used the term "load bearing" 4 separate times.

bobbylarrybobby a day ago | parent | next [-]

I noticed that when opus 5.5 was on parts on my codebase that had lots of 5.0-generated comments, it picked up its style. Unfortunately I think 5.5 has been trained to mimic what it sees so that you can ease up on the instructions, but this does mean parts of your codebase that 5 touched will be somewhat viral.

hypfer a day ago | parent [-]

GLM-5.3 does the same, funnily enough.

I gave it some vibecoded patch someone created with Opus 5 with the task of together figuring out the real root cause and what to do about it.

Big mistake. The rest of the session was all claude-speak up until I've rage-quit and restarted with Qwen and no context other than "here's what I think we've missed in the current implementation. How could we approach that?"

GLM felt like it got at least 20 IQ points dumber just from being exposed to claude's writing.

loveparade a day ago | parent | prev | next [-]

I will forgive it using load bearing as long as it finds the right seams.

mrcarrot a day ago | parent [-]

Yeah, you definitely don't want to hang things off the wrong seams, or you'll end up with quite the blast radius

mr_mitm 21 hours ago | parent [-]

Let's not paper over that gap and find the root cause instead

danbmil99 9 hours ago | parent [-]

Let's unblock that gate

MattyRad a day ago | parent | prev | next [-]

If anything it seems worse. I'm experiencing about -10% insufferable jargon, but +30% more verbosity. It's unredeemable. There also seems to be even less structured output (headings, bullets, etc).

therealdrag0 a day ago | parent | prev | next [-]

Due to this release note, I used it once to rewrite a doc, and was disappointed.

samat 8 hours ago | parent | prev [-]

thank you, i did not try it myself and having read this, will not waste my time and energy

sdthjbvuiiijbb a day ago | parent | prev | next [-]

I'm surprised that you're getting so many replies saying it's the same. So far in my usage today Opus 5.5 does seem like a noticeably better writer. Opus 5 frequently made me want to strangle it while 5.5 has been producing a lot less incomprehensible gobbledygook.

MattyRad a day ago | parent | next [-]

I know we're all experiencing NDFSMs differently, like that's part of the whole problem, but 5.5 just gave me "The truncating quantizer collides two oranges", which is a new low for me.

kyleee a day ago | parent [-]

Was it a true statement though? You may need to explain what you were working on (heh)

MattyRad a day ago | parent | next [-]

Like the sibling comment says, text is always "true" in the sense that it's reacting to context correctly. So yes, true, but insufferable.

Ironically, what I'm working on a post-processing hook for colorizing and summarizing responses without degrading the session quality. So "truncating" = summarizing, "quantizer" = char limits and thresholds, "collides" = conflicts, "two oranges" is referring to the "alert level colors" where a second model (Haiku/Sonnet) colorizes text based on the perceived (or suggested) priority of a response's statements (e.g. "just so you're aware, I didn't commit" is fucking useless and it needs to be blacked out).

So the original insufferable statement translates to something like "The code that checks whether a text fragment is too verbose was conflicting with the part that colorizes the text."

P.S. Let me know if there's something out there that exists like this- something that adds a dimension like color or priority-assessments on a per-response basis. So far all I've seen is 2 dozen ~100k starred GitHub plugins that add zero value or make things worse.

dev-complete 16 hours ago | parent [-]

Yep, sounds like the Opus I despise and the reason I dropped Claude Code. It reads like it doesn't want to be understood.

colordrops a day ago | parent | prev [-]

These claudisms are usually "true" but they are so heavily load-bearing that you need a claude-to-english dictionary to understand it.

pixelready a day ago | parent | prev [-]

Yeah I’m having a much better time reading Opus 5.5 output today vs. 5’s wall of nonsense. You still get a few telltale turn of phrases, though the load-bearing smoking guns haven’t turned up yet. It’s still a bit verbose compared to what I’d ideally like, but it’s tolerable now.

Code-wise it seems to still nitpick, especially in reviews, but it doesn’t seem to rabbit hole quite as badly on tangents and scope-creep. These are just first impressions though. It’ll take a few weeks of regular use to really have a sense of it.

derangedHorse 2 days ago | parent | prev | next [-]

As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.

BatFastard a day ago | parent [-]

I prefer Astra for creative uses, Fable seems better for hardcore coding.

int_19h 5 hours ago | parent | next [-]

I find that Fable still makes the best orchestrator model. Astra is great for code reviews, it is very eager to find flaws. Cheaper models can handle the actual coding - you want to do a review either way at the end.

comboy a day ago | parent | prev [-]

Whoa, I'm exactly the opposite.

Jtarii a day ago | parent [-]

Almost as if evaluating models is mostly astrology as this point.

comboy a day ago | parent [-]

Well, if you have a clearly defined task it's easy. When using them for my pipeline of writing explanations for Chinese words I have clear ranking, for example - opus 5.5 clearly better than opus 5 at writing and knowing details, annoying nit picker when it comes to finding errors (high accuracy, low usefulness) all in repeatable numbers on different datasets. The problem is that these models are most useful when you are facing a new task that you haven't encountered before. And yup then it's astrology.

atonse a day ago | parent | prev | next [-]

Yes this is a big part of what has turned me off Opus 5 completely. The other (more dangerous) one is how often it gets assumptions wrong. These both (along with Astra) caused me to split my time 50/50 now between the two models.

Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.

We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.

jaflo 2 days ago | parent | prev | next [-]

I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?

itsafarqueue 2 days ago | parent | prev | next [-]

The writing style is insufferable but it’s not just that. https://opusfived.dev/

mitchdoogle a day ago | parent | next [-]

I've been using Opus 5 since it was released and don't understand all the hate it gets. It very well could be something in my own local memories or Claude.MD files that prevents it, but I certainly have never experienced something like that site portrays.

mcintyre1994 a day ago | parent | prev | next [-]

That’s funny but I don’t really recognise that issue. I’m very confident that Opus 5 would correctly change the colour of just one button.

weego a day ago | parent | next [-]

You are right and make an important insight. While well meaning and amusing, it did not reflect the entire spectrum of outcomes that could arise from the worktree.

Navigating the landscape of agentic levers certainly requires a more detailed approach than this and you were certainly correct to push back.

mcintyre1994 a day ago | parent [-]

:(

a day ago | parent | prev [-]
[deleted]
UnboundedContex a day ago | parent | prev [-]

Nice, captured the flavour of 5's writing exactly.

notatoad a day ago | parent | prev | next [-]

after a few hours use, my impression is that Opus 5.5 writes like gemini. drastic improvement over all opus 5.1's claudeisms.

and i think i was generally more tolerant of 5.1's writing style than most people.

LtdJorge 10 hours ago | parent [-]

I think you mean Fable 5.1 here. Fable 5.1 was an improvement over Opus 5 but still had many Cludisms. Opus 5 was insufferable.

epicepicurean 2 days ago | parent | prev | next [-]

Much better than Opus 5. prompt:

> hi, can you explain how the scheduler works. keep it brief, but include important correctness details

some excerpts:

>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.

> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.

> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.

All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.

croemer a day ago | parent | next [-]

That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.

throwaway219450 a day ago | parent [-]

This one is pretty terrible (right after “The worst outcome is a spurious rerun, never lost work.”). We’ve got lands, several "no X", hyphenation, strange noun/verb sentence order and an unnecessary analogy word (swallowed).

> There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

pgphn a day ago | parent [-]

It’s absolutely atrocious and has made the latest models unusable. Seems like I’ll have to stick with Opus/Sonnet 4.6 for a little longer.

senderista a day ago | parent [-]

That's the main reason I'm using GPT models. I'll ask Fable to analyze something, then pipe its output straight through Astra without even looking at it first.

californical a day ago | parent | prev | next [-]

Oof thanks for sharing, that seems just as bad if not even worse than Opus 5 to me. Just about every sentence is painful. Particular standouts that a human would never write:

> Rewrites are declared by the publisher, never inferred from overlap

> NULL means dirty, and DELETE is the fence

croemer a day ago | parent [-]

Hah! You independently picked exactly the same sentences I flagged (I know you posted this 11min before me but the comment only appeared after I had submitted mine).

californical a day ago | parent [-]

Wow!! This is genuinely hilarious and is a pretty damning evidence of the problem

rfgplk a day ago | parent | prev [-]

So still effectively nonsense.

> Rewrites are declared by the publisher, never inferred from overlap.

This style of writing is idiotic because it conveys no additional information. It's no different from stating

> Rewrites are declared by the publisher, never when moons collide.

The two sentences are actually logically identical. No idea why these models keep writing like this.

> Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.

This is even more ridiculous.

adonese a day ago | parent | prev | next [-]

I don't know what they did but opus is really good while Astra/Sol are comparably bad. For my own tests and taste this might be the first time (fable perhaps excluded) where claude models are better than openai, since codex 5.3.

physicles a day ago | parent | prev | next [-]

Oh god yes.

Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.

r0l1 a day ago | parent | prev | next [-]

I worked with Astra for two weeks and the output was really bad compared to Opus. It made so many wrong decisions within C++, Go, Python and Typescript code bases. My college made the same experience and we moved back to Claude.

Trasmatta 2 days ago | parent | prev | next [-]

Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

nonethewiser 2 days ago | parent | next [-]

I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

penagwin a day ago | parent | next [-]

I assume it’s largely a side effect from the final RL in post training?

That’s the step that causes the most significant gains in agentic performance.

But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).

That’s why it often gets worse on models that simply had more RL post training from the same base.

nonethewiser a day ago | parent | next [-]

Reinforcement learning for specific use-cases like coding that degrade it's writing style... makes sense. Maybe it stands to reason later version of Opus were improved more by this sort of fine-tuning. Feels consistent with the observation of diminishing returns and worsening writing style. Wonder what changed (supposedly) in 5.5.

tancop a day ago | parent | prev | next [-]

Does Xiaomis approach help with this? They do all the post training steps at the same time instead of one by one, switch topics after a couple prompts so writing style is mixed with coding and tool use.

Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.

senderista a day ago | parent | prev [-]

For whatever reason, GPT models simply don't have this problem.

red75prime 15 hours ago | parent | prev | next [-]

I have a theory that they pushed the model away from human writing styles to not get copyright violation nags. In Opus 5.5 they've found a better point in the latent space of styles that is still far away from the human ones.

Trasmatta 2 days ago | parent | prev | next [-]

It truly was bizarre. I've used every major model since 2022, and not a single one had a writing style as bad as Opus 5

jaapz a day ago | parent [-]

Fable 5 was pretty bad too, but they fixed it with 5.1. Now with Opus 5.5 it seems they fixed it as well

senderista a day ago | parent [-]

Fable 5.1 is still intolerable. I invariably filter its prose through Astra before inflicting it on myself or anyone else.

walthamstow 20 hours ago | parent | prev [-]

I've been listening to some old Acquired podcast episodes and they do often talk in this type of Claudish. So it's getting it from Silicon Valley startup podcasters... Great...

LtdJorge 2 days ago | parent | prev | next [-]

Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.

rfgplk a day ago | parent | prev | next [-]

Opus is only usable if you have a post-turn formatter that strips all comments from the generated source. I'm not even kidding it's that bad.

senderista a day ago | parent [-]

Just run all comments through Sol or Astra.

atombender a day ago | parent | prev | next [-]

Astra is better here, but the one I'm the most impressed with is Gemini. It's always been good, but 3.6 Flash is even better. It writes in a pleasant, human style. Not perfect, but it has a good balance between technical accuracy and readability that is better than what I've seen from any other mainstream model.

Aperocky 2 days ago | parent | prev | next [-]

It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!

senderista a day ago | parent | next [-]

You're half right, but the half where you're wrong is hiding the real unlock.

FireBeyond a day ago | parent | prev | next [-]

You're right to call this out, and what's more, it's not even solving the original problem. I overlooked this in pursuit of the load-bearing seams and finding the wedge needed to uptick engagement.

legobmw99 a day ago | parent | prev [-]

Here's the X that Ys the Z:

outworlder a day ago | parent | prev | next [-]

Here's the part that nobody is talking about: corporate always had their jargons and unique writing style – LLMs have just created their own :)

michaelsalim 18 hours ago | parent | prev | next [-]

But is it load-bearing if their writing style changes tho?

epolanski a day ago | parent | prev [-]

> especially as all my coworkers started lobbing Opus 5 slop grenades everywhere

People that produce slop have to be fired asap, they're just human relays anyway.

algoth1 2 days ago | parent | prev | next [-]

Please update with your feedback

sha-3 2 days ago | parent [-]

I haven't heard it say "load-bearing" yet (I've used it for 30 minutes now), so that's a start.

comboy a day ago | parent | next [-]

That's a sharp observation and you're hitting on something most people never even realize.

mdavidn a day ago | parent | prev | next [-]

That's a sharp catch, and I think it exposes a real bug.

kgwgk a day ago | parent | prev [-]

Worth flagging!

altern8 a day ago | parent | prev [-]

Yes, it's unbearable. Hopefully they've actuallu lly fixed it.