| ▲ | 20k 7 hours ago |
| It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft? Edit: OpenAI have admitted they were training on prompts at the time they made their breakthrough https://mastodon.social/@tristanbuckmaster/11723647135247030... |
|
| ▲ | ameliaquining 7 hours ago | parent | next [-] |
| If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them. If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree. |
| |
| ▲ | 20k 7 hours ago | parent | next [-] | | The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem | | |
| ▲ | orangecat 6 hours ago | parent | next [-] | | In that it isn't able to genuinely solve problems Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture. This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI. | | |
| ▲ | 20k 6 hours ago | parent [-] | | I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach | | |
| ▲ | letmevoteplease 5 hours ago | parent [-] | | You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution. |
|
| |
| ▲ | ivory54321 6 hours ago | parent | prev | next [-] | | I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants? | | |
| ▲ | Timwi 4 hours ago | parent [-] | | If it could do this while properly crediting the researchers (the “giants”) it would be a different matter. |
| |
| ▲ | lotsofpulp 6 hours ago | parent | prev [-] | | Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it? | | |
| ▲ | 20k 6 hours ago | parent [-] | | OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct | | |
| ▲ | lotsofpulp 4 hours ago | parent [-] | | Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal. | | |
| ▲ | fwip 2 hours ago | parent [-] | | Terms and conditions are almost entirely about the company doing things that would otherwise be illegal. |
|
|
|
| |
| ▲ | sdenton4 4 hours ago | parent | prev [-] | | Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers? Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible. |
|
|
| ▲ | doctoboggan 6 hours ago | parent | prev | next [-] |
| I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data? |
|
| ▲ | user43928 5 hours ago | parent | prev | next [-] |
| All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training. That's it. The rest appears to be wild speculation. |
| |
| ▲ | jsw97 4 hours ago | parent [-] | | Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms. Never ever touch those requests. If you get a side by side comparison just resend the prompt. |
|
|
| ▲ | zem 4 hours ago | parent | prev | next [-] |
| even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"? |
| |
| ▲ | flir 4 hours ago | parent [-] | | Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it. Publication, though? Slimy is right. But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on. |
|
|
| ▲ | TZubiri 4 hours ago | parent | prev | next [-] |
| >It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject. |
|
| ▲ | TZubiri 4 hours ago | parent | prev [-] |
| >It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject. |