Remix.run Logo
amelius a day ago

Funny, I recently pasted output from Gemini into Claude, and it said it was total nonsense generated by an overconfident AI.

Apparently the "AI-generated bullshit" detection is already part of LLMs.

skorniienko a day ago | parent [-]

but have Claude actually fact-checked it or just provided an opinion?

amelius a day ago | parent | next [-]

Actually it was Fabel, and it totally tore apart Gemini's response.

skorniienko 14 hours ago | parent [-]

What if you use OpenAI as a judge and check both responses? Not defending Gemini, just curious if it was that wrong, or Fable - that aggressive.

LearnYouALisp a day ago | parent | prev [-]

"Lazy" evaluation