Remix.run Logo
book_mike 6 hours ago

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

okamiueru 6 hours ago | parent | next [-]

How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.

f6v 4 hours ago | parent | next [-]

My definition is that I can be much less precise with AI the more intelligent it is. It can extract the intent from my fuzzy description of the problem. Which means I can offload some of the thinking effort.

It wasn't possible a couple years ago. I used to make fun of people who were trying to get ChatGPT to think about the problem when all it could do was write code from the pseudocode you provide.

But now I can say: "Look at the latest log and make a plan to fix". And it takes it from there.

bikemike026 5 hours ago | parent | prev | next [-]

If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

okamiueru 5 hours ago | parent | next [-]

I'd have to ask for you to be more specific, otherwise, to take your answer at face value, it comes across as a contradiction.

> [Opus 5's output] is beyond the comprehension of virtually all engineers and developers

That would make it pretty bad? The key defining quality of good software, is clarity, and the ability to simplify a complex problem to the point of it seeming trivial.

> Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

The bar here should absolutely be to judge this against the expert level within each domain. I have time and time come across LLM output being woefully underwhelming in every single request where I am an expert. For all areas that I am not, it sure seems plausible. It is far more likely than not, that it is equally inadequate in the areas I lack the necessary knowledge to tell.

If the AI is being subpar in every field and category compared to an expert in said respective field, then, what a strange gauge of a tool's usefulness. Are we attributing higher value because a single model is "attempting to solve all knowledge and fields at the same time", why is that of any importance, or excuse?

We should not define "intelligence" as how effectively it can convince a non-expert of something being plausible. That sounds like the absolute worst tradeoff. You'd have to waste the experts time in filtering and refuting incorrect postulations that are cheep to generate. The perfect storm for bullshit asymmetry.

bikemike026 2 hours ago | parent [-]

I disagree with points 1, 2, and 3. Point 4, AI is better than average, and sometimes it's better than excellent. Point 5 is irrelevant.

hgoel 5 hours ago | parent | prev | next [-]

I don't think that's because of its "intelligence". It speaks obtuse techbro-ese: stringing together words that sound smart to obscure the simplicity of the thing it's describing. In many ways it's the opposite of intelligence.

Opus 5 and Fable 5 in particular suffer from this issue at worse level than most models in the same class.

greenchair 2 hours ago | parent [-]

yep, it is so bad i had to create rules to cut down on the techbro language and domain slang.

hgoel an hour ago | parent [-]

I just canceled my Claude subscription outright. The models are all gairly fungible, it's easy enough to just switch to another provider.

logicchains 5 hours ago | parent | prev [-]

You mean Fable 5 right? Opus 5 makes lots of stupid mistakes about anything that requires any domain knowledge.

odig 4 hours ago | parent | prev [-]

so?????

17 minutes ago | parent | prev [-]
[deleted]