Remix.run Logo
▲ afavour an hour ago

That’s not proof though is it? If the original LLM output is fallible surely the LLM review of that output is also very much fallible?

▲kolinko 3 minutes ago | parent | next [-]

Output of LLM can be infallible* even if LLMs themselves make mistakes. Ditto with humans.

As much as anything can be infallible.

▲rsfern 37 minutes ago | parent | prev [-]

Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).

I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.

I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here