Remix.run Logo
▲ jruohonen 3 hours ago

> generate a reliable list of which ones are not correct and a convincing reason why

That is not realistic, but I suppose where things are heading is that you have some indicator of the strength of evidence -- see Fig. 13 in the following insightful take:

https://news.ycombinator.com/item?id=49407226

Though, "strength" should probably be "reliability" and "validity", and I suppose those indicators are more for picking signals from the noise; i.e., what is even worth clicking and reading. That would be increasingly valuable already today due to the volume (and, yes, slop and other related stuff).

▲raddan 3 hours ago | parent | next [-]

> That is not realistic, but I suppose where things are heading ...

And to expand on this, it's not realistic because science is not armchair philosophy. You have to go out and measure the world.

Sometimes, through force of will, a person can think deeply about a problem and come up with beautiful theories that explain our measurements. Many scientists had careers like this, probably most famously, Einstein. But it's worth noting that Einstein also got a lot wrong! [1]

Even if we somehow give an LLM the ability to go out and measure things, I seriously doubt that the role of humans in science is done. There's a big difference between "an explanation" and "a good explanation." Ask any physicist. There's a surprising amount of aesthetics involved. Good theories are consistent with the evidence, but it's more than that-- there's a great deal of "taste" involved. And there's a good reason for that. For any real problem, there are effectively an infinite number of alternative hypotheses. From a "theory of science" standpoint, this should cause scientists nightmares, but it doesn't. Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation. If you spend time with scientists, especially in the "hallway track" at a conference, "taste" is a frequent topic of conversation!

[1] https://en.wikipedia.org/wiki/Einstein%27s_unsuccessful_inve...

▲dekhn 3 hours ago | parent [-]

I think you misunderstood. I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

I work full time on "lab in the loop" AI, so I'm pretty familiar with the need for real-world experiments. I am not proposing a fully autonomous scientist that could read an arbitrary paper and emit whether it's universally true without some verification method.

Also, to your statement: " Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation."

I'm a practicing scientist (well, ex-scientist) and it seems like most "satisfying explanations" end up being wrong or incomplete simply because they seem so satisfying.

▲raddan 3 hours ago | parent | next [-]

> I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

What's the difference? How do you find real mistakes without a model? Either you have a trusted mathematical model (in which case you already have a complete explanation) or you have to compare it against the ultimate oracle: the world. Or are you proposing something like "let's use an LLM to convert this hand-wavy English paper into a formal proof and then check it for logical fallacies?" In which case, fine, that would be useful, but that's not exactly the same thing (and also not as important) as saying that a paper advances a bad explanation. Just that the explanation is flawed in some way.

▲jltsiren 2 hours ago | parent | prev [-]

How do you determine that a mistake is serious enough that it actually invalidates the paper? How do you know that it's not possible to correct it and reach a similar conclusion with fundamentally the same approach? Especially when the fix would be complex enough to justify a follow-up paper.

▲dekhn 3 hours ago | parent | prev [-]

This is realistic. We already do this today: it's called "journal club". A bunch of grad students read the same paper and then criticize it. I've read papers that I thought were amazing only to. have somebody else notice a key issue in a method, or a conclusion that didn't follow, or outright omission of an important detail, in a way that could be verified by both the students and the authors of the paper.