Remix.run Logo
▲ foobarbecue 7 hours ago

ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode.

▲ricardobeat 7 hours ago | parent | next [-]

Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.

▲famouswaffles 6 hours ago | parent | next [-]

It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization.

Byte Latent Transformer - https://arxiv.org/abs/2412.09871

1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.

▲ 6 hours ago | parent | prev | next [-]
[deleted]
▲hbcdbff 7 hours ago | parent | prev | next [-]

Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.

▲kyralis 7 hours ago | parent [-]

... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.

▲Dylan16807 2 hours ago | parent | prev | next [-]

If you want it to read letters, all you have to do is make your tokens be letters. That's easier than normal tokenization.

▲LikesPwsh 7 hours ago | parent | prev | next [-]

One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.

▲vanuatu 6 hours ago | parent | prev | next [-]

we already fixed it with reasoning

▲mitxela 6 hours ago | parent | prev [-]

They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.

▲zahlman 7 hours ago | parent | prev [-]

> ChatGPT will say there are two Ds in "your mom" and one D in "uranus"

… Isn't it possible that it understands the innuendo and is going along with making the joke?

▲ndriscoll 5 hours ago | parent | next [-]

In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don't even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10/10 timeline.

▲queenkjuul an hour ago | parent | prev | next [-]

And the number of Rs in strawberry is a joke how?

▲Timon3 7 hours ago | parent | prev | next [-]

How many LLM users have anything in their prompt against "going along with jokes"? I'd guess not many.

What a wonderful new world.

▲kulahan 7 hours ago | parent | prev [-]

Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke.

▲bombcar 4 hours ago | parent | next [-]

It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times.

▲mitxela 6 hours ago | parent | prev [-]

We can only say bad things about the capabilities of LLMs.