|
| ▲ | karmasimida 41 minutes ago | parent | next [-] |
| Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it. |
|
| ▲ | tristanj 3 hours ago | parent | prev | next [-] |
| Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user. It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes. We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries. |
| |
| ▲ | biophysboy 2 hours ago | parent [-] | | I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math |
|
|
| ▲ | enraged_camel 4 hours ago | parent | prev [-] |
| Because OpenAI says so, obviously! |