Remix.run Logo
▲ dekhn 3 hours ago

I think you misunderstood. I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

I work full time on "lab in the loop" AI, so I'm pretty familiar with the need for real-world experiments. I am not proposing a fully autonomous scientist that could read an arbitrary paper and emit whether it's universally true without some verification method.

Also, to your statement: " Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation."

I'm a practicing scientist (well, ex-scientist) and it seems like most "satisfying explanations" end up being wrong or incomplete simply because they seem so satisfying.

▲raddan 3 hours ago | parent | next [-]

> I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

What's the difference? How do you find real mistakes without a model? Either you have a trusted mathematical model (in which case you already have a complete explanation) or you have to compare it against the ultimate oracle: the world. Or are you proposing something like "let's use an LLM to convert this hand-wavy English paper into a formal proof and then check it for logical fallacies?" In which case, fine, that would be useful, but that's not exactly the same thing (and also not as important) as saying that a paper advances a bad explanation. Just that the explanation is flawed in some way.

▲jltsiren 2 hours ago | parent | prev [-]

How do you determine that a mistake is serious enough that it actually invalidates the paper? How do you know that it's not possible to correct it and reach a similar conclusion with fundamentally the same approach? Especially when the fix would be complex enough to justify a follow-up paper.