Remix.run Logo
▲ Izmaki 5 hours ago

...I know, which is why "as a Language Model and not a real doctor" is a pointless comment to start off with. It should simply not recommend treatment if it's not sure it is correct. I wouldn't blame it or anyone if they asked for help treating a stiff neck, and the LLM (or your neighbor or parent or spouse) suggested light exercises to help relieve it - and do not jump to the suspicion that you may have meningitis.

As a Human, I do not need to know it is a Language Model.

▲StilesCrisis 5 hours ago | parent | next [-]

LLMs are famously bad at determining "if it's not sure it is correct." They are always confident, because a confident tone ranks better in RL.

▲wxnx 4 hours ago | parent [-]

> They are always confident, because a confident tone ranks better in RL.

This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).

I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.

▲Forgeties79 4 hours ago | parent | next [-]

> This makes it sound like RL rewards a confident tone

Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.

If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.

We don’t really need to speculate, this is already a problem.

▲daveguy 4 hours ago | parent | prev [-]

> This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).

A binary response vs rating is not related whether it learns confident or hedged tone. Either will produce a confident tone because humans respond more positively to a confident tone, hence the conman's language. Binary or not humans reward the tone and very much bias the model.

But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident. When the prior is greatly biased, a random number generator biased to that prior does better. The difference with humans and machines is humans are less likely to respond if they are less confident because they understand not knowing, which is why the training set is biased. It is one of the many fundamental flaw of LLM training and confusion of LLMs with intelligence. And that will not be fixed within the LLM architecture.

▲StilesCrisis 4 hours ago | parent [-]

I feel like in real life, we're constantly exposed to "I don't know" as a valid answer, but obviously we don't write down all the I-don't-knows in expert literature so the training corpus is wildly skewed towards confident answers because "we studied this for a month and have no idea, it's confusing" doesn't get published.

▲daveguy 2 hours ago | parent [-]

That's a great point. A training corpus based on written text will be inherently biased toward confident and right. Then the RLHF exacerbates the problem because people respond more positively to confident and too often assume correct when they read a confident response.

▲jester997 5 hours ago | parent | prev | next [-]

Yeah it should just state thing it means. But then again, there’s psychological impacts on society that we must be careful. For instance, teenagers talking to AI. If the AI just talks, people already start to feel real connections to the seemingly human entity. Maybe it’s better to disclose the reality up front?

▲Izmaki 5 hours ago | parent [-]

Would it be so bad that lonely people can have a real friend that they can bring everywhere they go and even share its passion with through vision and audio? We don't want destructive friends encouraging us to do bad things, but a real 'buddy', somebody who always has our best well-being in its interests?

Would it matter if this digital friend is not a real human behind a computer screen, but a Language Model in a data center?

I guess it falls into a similar category as buying "special performances to satisfy certain urges". It probably feels close to the real thing (I wouldn't know, I've never tried - promise! :P), but it's never the same as love.

▲nvme0n1p1 5 hours ago | parent | next [-]

> We don't want destructive friends encouraging us to do bad things, but a real 'buddy', somebody who always has our best well-being in its interests?

GPUs are not people, and generated tokens can't have interest in a person's well-being. If you try to pretend otherwise, the results are not great. https://www.cnn.com/2025/11/06/us/openai-chatgpt-suicide-law...

▲Izmaki 5 hours ago | parent | next [-]

That's almost a year ago. One LLM year is like 10 human years. They're a bit like dogs in that regard...

I'm pretty sure you will be paid a large sum of money if you can make one of the frontier models urge you to commit suicide from normal interactions with it.

▲nvme0n1p1 4 hours ago | parent [-]

Ah yes, my favorite LLM fallacy. "You used the wrong model! The newest fanciest model is perfect and makes no mistakes, have you tried it yet?" Let's force all of society's most vulnerable people to pay extra $$$ to Sam Altman, then surely all our problems would be solved.

It's just a convenient way to ignore the years of evidence of the harms. Any bad news can be swept under the rug, labeled outdated as quickly as it happens. Well here's one that just happened, maybe this kid should have used a fancier model too? Should OpenAI pay him a large sum of money for the good he's done? https://www.cnn.com/2026/08/15/us/arjun-aravind-massachusett...

▲Izmaki an hour ago | parent [-]

I'm not sure if you're trying to conclude that because a piece of technology was misused previously in a very bad way, that it will never be able to do good things because surely it was bad previously.

Reminds me of the 2000's crazy of "if you let kids play violent video games they will become unstable, violent psychopaths when they grow up". Thank God I was able to (rather easily) convince my mom that the idea that I would also steal a car and beat somebody to death with a golf club in real life just because that's what I did in GTA on my PlayStation, was completely absurd.

▲hardbass 3 hours ago | parent | prev [-]

I wish I didn't have to keep asking this question, but do you believe in souls?

▲nvme0n1p1 2 hours ago | parent [-]

We should treat AI that recommends suicide to teenagers the same as we would treat a therapist recommending suicide to teenagers. I don't care if either/both/neither have souls.

▲bluebarbet 5 hours ago | parent | prev [-]

My (controversial) view is that a psychologist is to a friend what a prostitute is to a lover. In a truly healthy society there would be no need for psychologists and prostitutes because everyone would have a handful of good friends and at least one lover. Back in messy reality, psychologists and prostitutes are a decent fix to keep society on the rails.

So. I think I agree with you.

▲StilesCrisis 4 hours ago | parent | next [-]

Well, a psychologist is also given years of training about what healthy behavior and relationships look like. Some best friends have this skill, others absolutely do not.

▲jester997 3 hours ago | parent | prev [-]

Well, no this doesn’t make sense. If you think about intent, most psychologists probably want to help people’s mental health. Prostitutes, although they offer a service that helps some people, I’m sure their motivation is primarily financial.

▲perching_aix 5 hours ago | parent | prev [-]

> It should simply not recommend treatment if it's not sure it is correct

Oh okay, darn, guess they just forgot to make it so!

▲nvme0n1p1 5 hours ago | parent [-]

It sounds like OpenAI forgot to include "make no mistakes" in the system prompt. Rookie mistake.