Remix.run Logo
anotherhue 2 hours ago

I'm glad Claude is so recognisable, it lets me bounce right off the empty calorie language very efficiently.

Maybe this thing is great, but it cannot be determined with this presentation.

CodeBeater 2 hours ago | parent | next [-]

Opus 5 specifically seems to be affected by a severe case of turbo-encabulationitis, almost as if it was trained to spit out as complex of sentences as it can.

And may your deity of choice help you if you decide to venture on subjects which you don't have a deep understanding of, as half the sentences it generates will be (barely) cohesive.

JustFinishedBSG an hour ago | parent [-]

> almost as if it was trained to spit out as complex of sentences as it can.

It probably was, at least inadvertently. Very easy to imagine that one of the post-training step (RLHF, DPO etc) reinforced the "sounds clever" behaviour.

juujian an hour ago | parent [-]

"It tested well with the focus group..."

pertymcpert an hour ago | parent | prev [-]

You can actually tell from the very first real sentence in the README:

> Efficiency is a 162-run controlled benchmark (same agent, same file tools, only the context differs).

It's so bad that I can tell it's Claude, specifically Opus 4.8/5.0, from the first 6 words.

shrishdwi an hour ago | parent [-]

Opus 5 is very paranoid on giving proofs for every thing it claimed. so I let it keep this one line.