Remix.run Logo
zahlman 4 days ago

> generated an article that continually undermines its own main point.

I disagree that this accurately describes TFA.

deathanatos 4 days ago | parent | next [-]

The entire second on chess engines is, from the view of the entire thesis of TFA, is incoherent. Let's assume, for sake of argument, that I agree with the section: that an idealized chess move predictor isn't a predictor — which is not a thing that exists, as the space of chess is enormous, but let's pretend! — that's not what LLMs are? Even if we just restrict ourselves to the space of written English prose, the space is quite literally infinite. So, hopefully obviously, no LLM is comparable to an idealized chess engine. Similarly, incoherently, we wave away the "make_more_likely", when, at least to me, the entire meat of that argument would be in the reward function, and we just gloss over that entirely.

(I would also agree with the parent commenter on that the writing smells like an LLM.)

astrobe_ 4 days ago | parent [-]

The reward function seems indeed to be the protagonist there, still it stays in the shadows. One can only imagine that it is some kind of evaluator that scores the sequence based on grammar correctness, semantic consistency, etc.? To use the proposed chess analogy, maybe it could be a Stockfish engine that evaluates the submitted position that results from the move submitted by the LLM?

Planktonne 4 days ago | parent | prev [-]

I'm not sure what you want me to do with that information; clearly I do think that my description is accurate.

The article is littered with both AI tells and admissions that 'next token prediction' is what is happening. Hence my description.

garrinm 4 days ago | parent | next [-]

It was written by a human. There are AI edits but it’s very much a human composition. Perhaps a bit sloppy.

Planktonne 4 days ago | parent [-]

In my experience, people who do 'AI-assisted' writing tend to be very bad at noticing how much of their work AI has changed. I'm sure you put thought into it, but passing it through AI takes a lot of that out.

garrinm 4 days ago | parent [-]

I think that’s fair, I didn’t actually run the whole thing through an AI. it was more targeted edits, but each time it does erode at my writing. But at the same time, I don’t think it’s a good reason to dismiss this. Because I did spend several hours writing it, and I did put a lot of thought into it, and it was not in any meaningful way generated by AI.

Planktonne 4 days ago | parent | next [-]

Mostly I disagree with the article's ideas, if that helps; the AI was just a secondary factor.

> I don’t think it’s a good reason to dismiss this

AI-generated prose reads as sending a 'lack of effort' signal to a lot of people, just as no editing at all does. ; It's an effective heuristic that we've all learnt in the last couple of years.

In either case, it's not always fair: there are people who deeply care about their ideas but forget to fix basic errors, or pass it through AI.

In both cases though, the advice is the same: if you want people to take your output seriously, you need to signal that you are taking it seriously. That used to mean editing for spelling and grammar. Now it means not using AI.

zahlman 4 days ago | parent | prev [-]

Out of curiousity, do you ask your editing system for diffs? Seems to me like the best way to notice and review whether your "voice" is degrading.

Personally I would never let an LLM touch my prose (although I'd happily use it for research and paraphrase things it told me), but if I force myself to consider the idea, that seems like the first thing I'd want. Maybe upon reading a diff you'd even consider going a third way with the text.

zahlman 4 days ago | parent | prev [-]

> I'm not sure what you want me to do with that information

For example, you could cite specific things that you believe to be "AI tells" or "admissions".

Planktonne 4 days ago | parent | next [-]

It's a short article; you could read it. One example to get you started is the very first sentence:

> Strictly speaking, the statement “LLMs are next-token predictors” isn’t wrong, but it’s incomplete.

The article is about how 'next-token predictor' is the wrong mental model; it opens with the admission that it is not the wrong mental model.

zahlman 4 days ago | parent [-]

I did read it. People are allowed to disagree with your conclusions. Comment guidelines ask us all not to make such accusations.

To say that a statement is incomplete, but not strictly speaking wrong, is perfectly compatible with describing it informally as "wrong" in the sense used in the title (i.e.: "not the most appropriate possibility").

Planktonne 4 days ago | parent [-]

To informally describe something as wrong in an article focused on how it's wrong to informally describe something is incoherent.

There's a certain irony in pointing me towards the guidelines on the grounds that I have limited patience with your comments that violate them in various ways. I'm not sure that this is a productive discussion.

angoragoats 4 days ago | parent | prev [-]

Not the person you’re replying to, but I read the whole article as an admission that it’s still a next-token predictor. More specifically: what does RLVR fundamentally change that somehow makes the whole process no longer a next-token predictor? The article makes no attempt to explain this. Additionally, I find its framing of the term “next-token predictor” as meaning “predicting the next token only based on raw training data” in common usage to be a bit dishonest.

To summarize: yes, RLVR and other synthetic training methods exist! It’s still a next-token predictor, and it does not “learn” or “think” or “reason” in the human sense, like so many people seem to believe.