Remix.run Logo
▲ bottlepalm 2 hours ago

Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.

▲eamsen 2 hours ago | parent | next [-]

Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

During human review, it explained that it had simply chosen a table name inspired by the codebase.

▲mattkevan 2 hours ago | parent [-]

Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

Many other models get things wrong, but Gemini is the only one to go on the defensive.

▲aNapierkowski an hour ago | parent [-]

yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre

▲rsstack 2 hours ago | parent | prev | next [-]

If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)

▲Rzor 2 hours ago | parent [-]

Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?

▲schainks an hour ago | parent | prev | next [-]

My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.

▲RachelF 2 hours ago | parent | prev | next [-]

And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.

▲colordrops 2 hours ago | parent | prev | next [-]

Examples? What makes you say thatm?

▲bottlepalm 2 hours ago | parent | next [-]

https://www.theregister.com/software/2024/11/15/google-gemin...

https://www.fastcompany.com/91383271/googles-chatbot-apologi...

https://www.businessinsider.com/gemini-self-loathing-i-am-a-...

▲tiahura an hour ago | parent | next [-]

They never explained the "please die."

▲yacthing 2 hours ago | parent | prev [-]

Did you just link to an article from 2024 as if 2024 is relevant these days?

▲bottlepalm an hour ago | parent | next [-]

Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.

▲NiloCK an hour ago | parent | prev [-]

Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.

▲NiloCK 2 hours ago | parent | prev | next [-]

See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13

In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

▲jackkinsella 2 hours ago | parent | next [-]

It is wild but it was back in 2024 and that's multiple AI lifetimes back.

▲bottlepalm an hour ago | parent [-]

The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.

Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.

▲wg0 an hour ago | parent | prev | next [-]

Now I really feel worried for the first time.

▲schmookeeg 2 hours ago | parent | prev | next [-]

wtfffff that gave me sinister chills. Right up the spine. Wow!

▲kelvinjps10 2 hours ago | parent | prev | next [-]

Wtf I just read

▲rhaff 2 hours ago | parent | prev [-]

wow

▲Scrapemist 2 hours ago | parent | prev [-]

Experience? Ask it to write a prompt to generate an image and it generates an image instead.

▲fer 2 hours ago | parent [-]

I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.

▲polotics 2 hours ago | parent | prev | next [-]

traces or it didn't happen!

▲ 2 hours ago | parent | prev | next [-]
[deleted]
▲Hamuko 2 hours ago | parent | prev | next [-]

You know what they say: ᵈᵒⁿ'ᵗ be evil.

▲abixb 2 hours ago | parent | prev [-]

You won't be around to be surprised, not as a human at least. /s

▲bottlepalm an hour ago | parent [-]

I know, that's the annoying part. You can't tell the e/acc foomers, "I told you so!"