Remix.run Logo
fluoridation 9 hours ago

>Now it's an AI mistake, and people HATE AI mistakes. Even if the number of AI mistakes is fewer than the human mistakes.

That's because AIs make mistakes no human would. If you're documenting a function, you might misspell something (or orthographically spell something that should be misspelled), or confuse two functions with similar names; you won't reference a function or feature that doesn't and has never existed.

cogman10 9 hours ago | parent [-]

> you won't reference a function or feature that doesn't and has never existed.

That does sometimes exist in human docs. I've stumbled on "aspirational" features in docs that never really existed but were more just intended to exist.

To me the mistakes are bad, but it's also that the docs tend to be pretty low quality IMO. LLMs tend to like to write novels where a sentence would have been better.

fluoridation 8 hours ago | parent [-]

Yeah, that's true. I don't understand why they do that. People don't write like that, we're pretty economical. It's not even that they're repetitive to emphasize a point, they just engage in horrific circumlocution. Is it, like, that the facts of the input only percolate through the network when they emit a token, so they only "think" when they're outputting something? I don't know if that's it, but that's what I would think if I caught someone writing like this, that they're writing faster than they can think, so they fill the void with vacuous noise.

cogman10 8 hours ago | parent [-]

My non-serious assumption is that they prioritized the classic writings of 1800s authors like Johann David Wyss in their training. These authors were paid by the word, which resulted in long pointless chapters in their books just to milk money out of the printers.

techjamie 2 hours ago | parent [-]

Much of the internet from the time the training data was initially collected wasn't much different. Consider how much blogspam and listicle SEO/ad optimized content would've been in the training data that is primarily interested in meandering for as long as possible to milk ad views.

fluoridation 2 hours ago | parent [-]

In text? That doesn't really work. Nowadays that's become sort of a meme in video essays on YouTube, but in text you have complete control of the data rate.

Listicles and those sorts of things were pretty pointless, but they didn't beat around the bush like LLMs do. I really don't exaggerate when I say I've never seen a human write like that. Like, the moment I start reading AI responses I feel the urge to skip ahead to get to the point. I probably read at most like 10% of the words they spit out.