Remix.run Logo
▲ qoez an hour ago

Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.

▲viraptor an hour ago | parent | next [-]

Have you got a link to someone pointing it out? It looks like they're still going through formal proofs so likely found the problems that way.

▲sanxiyn an hour ago | parent [-]

Yes: https://x.com/ElliotGlazer/status/2108026240582246600

▲viraptor an hour ago | parent | next [-]

So someone ran a different LLM to find an issue they'd find anyway during formalisation? That's not the same as relying on thriving community.

▲oliculipolicula an hour ago | parent [-]

The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.

This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.

▲kolinko 5 minutes ago | parent | next [-]

My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.

▲oliculipolicula a minute ago | parent [-]

No. Someone received a tip, presumably from inside OAI. He then used OAI Astra to check..

▲yorwba 26 minutes ago | parent | prev [-]

What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.

▲oliculipolicula 3 minutes ago | parent | next [-]

Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.

▲idiotsecant 6 minutes ago | parent | prev [-]

Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO.

▲afavour an hour ago | parent | prev [-]

That’s not proof though is it? If the original LLM output is fallible surely the LLM review of that output is also very much fallible?

▲kolinko 2 minutes ago | parent | next [-]

Output of LLM can be infallible* even if LLMs themselves make mistakes. Ditto with humans.

As much as anything can be infallible.

▲rsfern 37 minutes ago | parent | prev [-]

Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).

I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.

I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here

▲caaqil 28 minutes ago | parent | prev [-]

> With automated math that community as tao pointed out is at risk.

If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.

▲doc_ick 18 minutes ago | parent | next [-]

“If they can be automated, they are not necessary.” That’s a pretty interesting take as eventually everything could be automated.

▲lkey 8 minutes ago | parent | prev [-]

You've abandoned every part of yourself to the siren's song of efficiency and automation, huh?

To witness an arson and rejoice reveals an ugly kind of sadism.