> the AI says things like “Interesting!”

My experience of those utterance is that it’s purely phatic mimicry: they lack genuine intuitive surprise, it’s just marking a very odd shift in direction. The problem isn’t the lack of path, is that the rhetorical follow-up to those leaps are usually relevant results, so they stream-of-token ends up rapidly over-playing its own conviction. That’s why it’s necessary (and often ineffective) to tell them to validate their findings thoroughly: too much of their training is “That’s odd” followed by “Eureka!” and not “Nevermind…”

▲

jackcarter 2 hours ago | parent | next [-]

It’s funny that this is probably due to bias in the training texts, right? Humans are way more likely to publish their “Eureka!” moments than their screwups… if they did, maybe models would’ve exhibit this behavior.

Now that AI labs have all these “Nevermind” texts to train on, maybe it’s getting easier to correct? (Would require some postprocessing to classify the AI outputs as successful or not before training)

	▲	embedding-shape 3 minutes ago \| parent \| next [-]
		I think it's more explicit than that, part of post-training to enforce the kind of behavior, I don't think it's emergent but rather researchers steering it to do that because they saw the CoT gets slightly better if the model tries to doubt itself or cheer itself on. Don't recall if there was a paper outlining this, tried finding where I got this from but searches/LLMing turns up nothing so far.
	▲	Forgeties79 an hour ago \| parent \| prev [-]
		My understanding is that it’s the result of these companies making sure to keep you engaged/happy less than the result of data these companies train with. I don’t know if it’s true or not but it certainly tracks given LLMs are way more polite than the average post on the internet lol

▲

sigbottle 2 hours ago | parent | prev | next [-]

I think that a lot of models have to sprinkle in a lot of "fluff" in their thinking to stay within the right distribution. They only have language as their only medium; the way we annotate context is via brackets and then training them to hopefully respect the brackets. I'd imagine that either top labs explicitly train, or through the RL process the models implicitly learn, to spam tokens to keep them 'within distribution' since everything's going through the same channel and there's no fine grained separation between things.

Philosophically, it's not like you're a detached observer who simply reasons over all possible hypotheses. Ever get stuck in a dead end and find it hard to dig yourself out? If you were a detached observer, it'd be pretty easy to just switch gears. But it's not (for humans).

▲

hmontazeri 30 minutes ago | parent | prev | next [-]

The new Opus 4.7 thinks quite often with: Hmmmm…

Haha anyone else seen this?

▲

2 hours ago | parent | prev | next [-]

[deleted]

▲

epolanski an hour ago | parent | prev | next [-]

Interestingly this is strikingly similar to how my mind would process something I find genuinely interesting.

▲

animal531 an hour ago | parent | prev [-]

I've somehow managed to train mine out of trying to fluff me up the whole time, its become very factual.

Overall it saves me a lot of time reading when it's just focusing on the details.