| ▲ | oliculipolicula an hour ago |
| The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first. This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD
nerds weren't confident bosses pushed ahead anyway. |
|
| ▲ | yorwba an hour ago | parent | next [-] |
| What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either. |
| |
| ▲ | oliculipolicula 41 minutes ago | parent | next [-] | | Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch. | | |
| ▲ | dragonwriter 7 minutes ago | parent | next [-] | | LLMs are unpredictably complex with potentially sigbificnaly different reaults based on random seed and seemingly insignificant prompt details) in the ideal case and nondeterministic in practice, so someone finding an error with a given LLM is not strong evidence that the result was not checked with an LLM, even with the very same LLM, previously. | |
| ▲ | Hamuko 3 minutes ago | parent | prev [-] | | I can throw Opus 5.5 at my code three times for code review and get three different sets of things it considers to be issues. I imagine all of them were checked with Astra at least once, but were they checked enough times? |
| |
| ▲ | idiotsecant 44 minutes ago | parent | prev [-] | | Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO. |
|
|
| ▲ | kolinko 43 minutes ago | parent | prev [-] |
| My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models. |
| |