Remix.run Logo
hellohello2 2 hours ago

This whole thing immediately reads as Claude generated, making it hard to take seriously.

Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent

seizethecheese an hour ago | parent | next [-]

These results don’t just contradict more serious benchmarks, they are wrong on an entirely different axis. This is a saturated benchmark. Haiku gets 96%. The results here are “not even wrong” and this being #1 on HN right now is a massive smell of either bots or massive ignorance or both.

urams an hour ago | parent [-]

People REALLY want the open models to be better than the frontier labs'.

geek_at 11 minutes ago | parent [-]

investors REALLY want the closed models to stay frontier forevery. mmw the future of AI will be local and offline

Tepix 2 hours ago | parent | prev | next [-]

Note that in AAs report, Kimi K3 was at #1, then they updated their criteria and published a new report on the same day where it was no longer at the #1 spot. They may be under pressure not to declare a chinese model as #1.

rdsubhas 2 hours ago | parent | prev | next [-]

I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling.

Look at the later comments, they have substance oriented discussions.

HN Mods - can you please consider a policy against such comments, it's overflowing the site and is diluting discourse and value. If people don't like an article, they can simply ignore it. These articles are reaching the top because enough people consider it of value.

But when an article comes to the top, and the topmost comment and discussion thread is an unqualified witch hunt, it's getting sick.

Glyptodon 2 hours ago | parent | next [-]

I don't know about Claude specifically, but I think people are starting to internalize a sense of when prose reads as having "AI-smell" and I do agree phrases like "Same driver, same track. The LLM is the star." trigger it for me too. That said, that doesn't mean a ton about the whole thing - could be anything from humans starting to echo AI style to someone writing "give my results a headline summary" to an AI to someone saying "here's the data, write an article."

rdsubhas an hour ago | parent [-]

AI is trained on human data. And high quality human data at it's best.

Can we assume everything we think as AI - must have had a high-quality human pattern behind it, and there is no way to 100% prove which is which - unless the author shows a screencast of them typing the artice?

This is not healthy. The right thing to do is – if someone doesn't like an article, they should ignore it – they shouldn't so confidently brand it AI without any proof at all, just because it fits their mood and style.

Glyptodon an hour ago | parent | next [-]

I think it's more complicated than that. Training does a lot more than just make models imitate the highest quality training data, and even what high quality training data includes can be subjective depending on your goals and tastes. And a lot of effort does go into making sure the models don't have "bad personality" - I'm sure a lot goes into making sure they trend towards appropriate reading levels and various other things that aren't strictly about being a Mark Twain or JRR Tolkien level writer.

adastra22 an hour ago | parent | prev | next [-]

At this point AI is trained on AI data, and it is moving AI output into super attractors that have no correlation with human text.

dgellow 42 minutes ago | parent | prev | next [-]

> And high quality human data at it's best.

LLMs are trained on Reddit…

avazhi an hour ago | parent | prev | next [-]

> AI is trained on human data. And high quality human data at it's best

And if there was any doubt that you don’t know how LLMs work, this line sorted it out lol.

MallocVoidstar 42 minutes ago | parent | prev [-]

AI writing is not high-quality writing. And I don't think it's trained on particularly high-quality writing, I'm pretty sure it's trained on SEO slop. The old outputs of gpt-4o that loved the word "delve" read just like SEO spam blogs.

thegeomaster an hour ago | parent | prev | next [-]

For what it's worth, before I hurl such an accusation I always check the post in Pangram (https://pangram.com). It always detects the text at 90+% AI generated.

Notably, Pangram is very conservative, and it's not difficult to manually get an LLM generated passage of text to turn human-written. So a score of near-100% AI generated means the writer didn't do even very light editing for a large part of the text.

There are extremely good reasons to be skeptical of fully LLM-written content. Our attention spans and our online platforms were built in a time where a long, data-supported article with references was expensive to produce. The time to write it was vastly longer than the time to read it, which means you could usually rely on some good faith, baseline level of accuracy and thinking on the writer's part.

With LLM-generated content, it's very difficult to know if 5 minutes, 5 hours or 5 days went into writing of the content. On the surface, it all looks similar, but the 5 minute version usually communicates very little or very shallow ideas, makes factual errors, and is generally lacking a lot of context. It's fast food writing.

These low effort versions of content take way more to read than they take to write. And combined with the obtuseness of the writing style, it all places undue burden on the reader to figure out the underlying message, because a lot of it has been mangled by the writing process.

I think LLMs are hugely helpful for writing, but to use their proper potential one needs to use them for feedback and engage with them at a level deeper than simply "write an article about X" or "rewrite this paragraph", and the text then doesn't obviously read AI generated as a bonus - I think nobody really has a problem with this.

dtech an hour ago | parent | prev | next [-]

I don't know if it's an intentional joke or something, but your comment smell incredibly like it's written as AI.

Per the writing, reading AI writing is like having something taste "chemical". Not very specific, but still a very recognizable and bad taste that makes it hard to enjoy and marks the thing as low quality.

voidmain an hour ago | parent | prev | next [-]

The first thing I do when I see an interesting article title is click on the comments to see if people have noticed that it's AI generated. If I didn't have this option because of the policy you desire, I would probably just stop reading HN entirely.

lelanthran 8 minutes ago | parent | prev | next [-]

> I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style.

Other than having all the AI tells, what proof can there be? Are you seriously asking that, because there is no way to provide hard proof that we should just stop pointing out obvious AI tells?

dgellow 44 minutes ago | parent | prev | next [-]

FWIW this part of your comment does look generated by Claude:

> I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling.

Was that the case? Genuine question. Here the em dash is an actual em dash symbol, where in the other paragraph you used what looks like a minus symbol. It’s also the type of construct used by Claude.

timmytokyo an hour ago | parent | prev | next [-]

The "About" page basically admits the entire site is written by AI.

https://reinvently.co.uk/about/

itishappy an hour ago | parent | prev | next [-]

> [...] there is one unquantified, unproven comment at the top saying it's 100% [...]

The parent was clearly not stating anything confidently.

> Look at the later comments, they have substance oriented discussions.

There are two sentences in the parent comment. Why is the top response (yours) pointedly ignoring the sentence with substance?

lopatin an hour ago | parent | prev | next [-]

The poster's username is "ed-is-ai" so I don't know how unproven it is.

Farmadupe an hour ago | parent | prev | next [-]

It's a valid shift to move onto actually trying to read the article critically (which I don't mean in an insulting way -- If you assume a writing has something worthwhile to tell you, reading it critically is how you learn the worthwhile thing)

In this article, _I_ get unstuck right at the very first paragraph:

> Same driver, same track. The LLM is the star. Seventeen leading models driven round the identical 28-realworld task lap — one harness, same verbatim prompts, deterministic grading — and the results go on the board.

It jsut doesn't make much sense to me. At best, I think it can be glossed as... "I made an arbitrary benchmark which I'm not going to explain, and I plotted the results."

------

Getting my own opinions out, this is blatant slop. It claims to be "deterministic grading", but then almost the _entire_ webpage is editorialization. Examples:

* "If you only run one model, run glm-5.3"

* "opus-5 posts the best rubric on the default panel"

* "deepseek-v4-pro is nominally cheaper still at $0.0029 [...] treat it as a batch-only option."

* " It performed well on what it completed"

jsnell an hour ago | parent | prev | next [-]

It's not nearly all articles. It's predominantly for the AI-written articles. And no, it's not unsubstantiated allegations or "virtue trolling". The tells are painfully obvious, and can be verified with high quality AI-detectors like Pangram.

And this really has to be policed. Once some tipping point is reached and too much of the HN homepage is AI slop, the site is dead.

hellohello2 25 minutes ago | parent | prev | next [-]

Respectfully, the linked website reads exactly like the Claude artifacts I read all day, i.e., it is low effort. I do not mind AI generated writing at all, but I do mind bad writing.

Further, you ignored the actual contents of my comments, to latch onto a superficial aspect. Please tell me: why do these results contradict existing attempts at benchmarking LLMs, which were designed with considerably more effort? Because the website certainly doesn't explain why in a way that is human-readable, which is why I asked.

EDIT: as explained by another commenter below, its because Fable refused to perform some of the tasks.

malshe an hour ago | parent | prev | next [-]

This article is pure AI slop. But if you want an objective metric, Pangram 4.0 says "100 % of this text is AI". In my experience, Pangram has a very low false negative rate and relatively high false positive rate for human writing. That means it errs on the side of humans. So if it says something is 100% AI written, I'm quite convinced it is.

But beyond that, can't you see how terrible the writing is? This is unadulterated AI slop.

seizethecheese an hour ago | parent | prev | next [-]

I’ve previously called for such a policy, but I am slowly changing my mind. This blog post is so obviously slop… and I suspect it’s the product of an upvote ring of some sort (Haiku is 96% on this benchmark.)

johnfn an hour ago | parent | prev | next [-]

This article is so blatantly, obviously, painfully AI generated. The real "worrying trend" is this getting upvoted to the frontpage in the first place. If you want "proof" just chuck it into pangram.

HDThoreaun an hour ago | parent | prev | next [-]

I generally agree, but the very first sentence of this post is "Same driver, same track." This prose is so AI coded, that even if its your natural writing style you would change it so as not to confused with AI. If this was in the middle, fine, but as the very first sentence it is quite strange.

ljm an hour ago | parent | next [-]

I've been watching old TV shows lately, the original CSI, all those 24 episode a season procedurals and so on.

Whatever we call 'AI coded' now has an awful lot in common with old TV screenplays where no word of dialogue was wasted. It's all the same style: punchy, plays on words, a bit of smart-ass in there.

HDThoreaun an hour ago | parent [-]

I think that makes sense. The "ai prose" stuff is a result of post training where the labs force the model to sound a certain way. It makes sense that the general chat models are trained to sound similar to mass market media, revealed preferences show the average person likes that kind of language.

lopatin an hour ago | parent | prev [-]

To my horror, it kept the metaphor going on and off the whole article. Half way through, there is a section called "Three cars failed the crash test".

an hour ago | parent | prev | next [-]
[deleted]
woodruffw 2 hours ago | parent | prev | next [-]

[dead]

desecratedbody an hour ago | parent | prev [-]

[dead]

SwellJoe an hour ago | parent | prev | next [-]

While I was inclined to push back on the results, with Fable and Sol being so low, I have to admit I've also run into refusals several times since the latest models have arrived, and I've had to use Kimi K3 or DeepSeek to complete the task. Usually security auditing type stuff, but Fable balks at all sorts of ridiculous things, sometimes stupid things. I've even had Fable fall back to Opus and then Opus refused the task as well. So, it actually is becoming hard to use US models for everything because they refuse to work on a pretty broad selection of security and security-adjacent tasks. I guess if you're not at a Fortune 500 or a member of a fascist government, you don't get to use the best models to protect yourself and that's just how it's going to be.

But, you're right. The prose is miserable Claude-speak, difficult to wade through.

hellohello2 13 minutes ago | parent [-]

Interesting idea, I had not considered refusals. I have ran into some as well although rarely. I'm not certain the 28 tasks described would trigger it though, if I understand correctly the security tasks are about avoiding prompt injection and not about doing security work.

EDIT: You were correct, Fable and Opus reject some of the coding tasks, which is why they score lower. Thanks for explaining.

EDIT2: I believe this benchmark is invalid, my Opus 5 runs the supposedly rejected tasks just fine.

koe123 2 hours ago | parent | prev | next [-]

Haha theres even an em-dash in the title

noname120 2 hours ago | parent | next [-]

It’s an en dash, not an em dash ;)

wolttam 2 hours ago | parent | prev [-]

HN will automatically do that to your title

ckocagil 2 hours ago | parent | prev [-]

Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.

hellohello2 11 minutes ago | parent | next [-]

Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.

irishcoffee 2 hours ago | parent | prev [-]

The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.