Remix.run Logo
dorjoycb 8 hours ago

It seems like some other mathematicians (not affiliated with openAI) have also (or close to) done this. A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774

colinhb 7 hours ago | parent | next [-]

The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys:

> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.

hkmaxpro 7 hours ago | parent | next [-]

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof.

https://x.com/sama/status/2097385167002415140

https://x.com/SebastienBubeck/status/2097379411691516310

A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

igleria 7 hours ago | parent | next [-]

If I was a company with a zero data retention contract involving OAI I would be asking for a third party audit of such claim of zero retention like, yesterday.

int32_64 6 hours ago | parent | next [-]

Could they say they don't retain, but do something "transformative" like use their own AI to summarize and paraphrase user sessions?

Keyframe 4 hours ago | parent | next [-]

data collection companies regularly fuzz and mask data and call it a day. the fuzz and the mask quality is debatable.

anon48293 6 hours ago | parent | prev [-]

Yes, and that’s exactly what I believe they are doing

linkregister 7 hours ago | parent | prev | next [-]

Is there an implication of violation of ZDR here? Not a challenge. Just a request for clarification.

igleria 5 hours ago | parent | next [-]

to my knowledge the mathematicians did not have ZDR so it would be incorrect to assume OAI violated such a thing.

I'm suggesting audits, not suing... if that is the implication.

dakolli 6 hours ago | parent | prev [-]

By the way, the company that made it's entire product off of stealing all data it could get it's hand on while violating copyright and pirating, is not all of a sudden going to respect your data. If you think OpenAI or any major AI lab is going to give you true ZDR, I have a bridge to sell you.

fc417fc802 4 hours ago | parent | next [-]

So use bedrock or vertex or whatever. Those are the ZDR offerings. Or was it your intention to insinuate that the major cloud providers are conspiring with openai to violate their contractual obligations to their customers?

dakolli an hour ago | parent [-]

Yes. You're naive if you think any of these cloud providers care about your data when they're all in the midst of a AI revolution psychosis. They dont care about their reputation or what you think of them, they think they're going to have a machine god their side.

linkregister an hour ago | parent | prev [-]

If my company finds any evidence of OpenAI violating ZDR, we'll sue for breach of contract and fraud, and collect damages. I think we'll be able to afford the bridge you're selling. You've got the title and title insurance, right?

irthomasthomas 5 hours ago | parent | prev | next [-]

It can still be academic plagiarism even if they ticked the box to allow training on their prompts.

infamouscow 6 hours ago | parent | prev [-]

The idea OpenAI or Anthropic won't train on your data—even with an enterprise contract—is a fantasy at best, and dilusion at worst.

letmevoteplease 6 hours ago | parent [-]

This is how every conspiracy theorist thinks: my enemy is Bad, and if they did a Bad thing, it would be Good for them, therefore they obviously did it. No evidence needed other than "motive" + my enemy is evil. But even if your enemy is evil, in this case, they would be fools to take the legal risk of violating their contract for the minimal upside of a tiny bit more training data (and fools to assume this would not be exposed in a large organization). So you need to assume your enemy is both evil and remarkably stupid.

jsw97 5 hours ago | parent | next [-]

I think it’s probably not surprising that they would go up to the contractual limit or into a grey area; but exceeding that would require too much coordination among individuals, as you say.

boinkboink78912 5 hours ago | parent | prev | next [-]

[dead]

ghk-adsf 5 hours ago | parent | prev [-]

[flagged]

bawolff 5 hours ago | parent [-]

No, people who believe things without evidence because it fits their personal narrative are consiracy theorists.

The thing that makes someone not a conspiracy theorist is evidence.

concinds 6 hours ago | parent | prev | next [-]

Their own claim is that they wanted Buckmaster without Alpöge to lead a rewrite of OpenAI's Navier-Stokes work, not of Alpöge-Buckmaster's Euler work.

No one can know if that's correct without proof but I don't know how you're reading it so differently.

hkmaxpro 5 hours ago | parent | next [-]

They want Buckmaster to dissociate with Alpöge in a follow-up rewrite of OpenAI's work. (They only publicly admit “Buckmaster as the lead author”, but judging from Buckmaster’s statement, it’s pretty clear that don’t want Alpöge at all.)

Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:

> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!

5 hours ago | parent | prev [-]
[deleted]
nolta 6 hours ago | parent | prev | next [-]

> We did not rush to publish even though the other team wasn't communicating with us.

Pretty clear this was rushed: there are no comments from external mathematicians, unlike the Erdős announcement:

https://openai.com/index/model-disproves-discrete-geometry-c...

fkarakurt3 5 hours ago | parent | prev | next [-]

They are missing a great marketing stunt: "Our models are so good that our competitors are using it for leading research".

viccis 6 hours ago | parent | prev [-]

Kinda weird because the pure math world doesn't have this concept of "lead authors" like other STEM areas do. Authors are alphabetically listed and there isn't generally this kind of hierarchy.

tkamat29 6 hours ago | parent | next [-]

From what I understand they aren't comfortable with the Anthropic employee being an author at all, not just lead author.

fooker 6 hours ago | parent | prev [-]

It works in niche fields where everyone knows each other and every discussion involves who did what portion of the work for a result.

charm137 7 hours ago | parent | prev | next [-]

It's astounding that the thought to dissociate one of the mathematicians from the proposed publication was driven by their corporate institutional affiliation - and that that exclusion was suggested by a scientist themselves! This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!

Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).

curt15 6 hours ago | parent [-]

The scientist allegedly making that request comes from a machine learning background. Perhaps he's not familiar with the culture in mathematics regarding authorship. That sort of squabbling over author priority would be unconscionable to mathematicians.

contubernio 5 hours ago | parent | next [-]

Bubeck is familiar with how publishing works in mathematics.

tensor 5 hours ago | parent | prev [-]

No, it's common to list authors alphabetically in a lot of computer science journals too.

peri-cl 7 hours ago | parent | prev | next [-]

(To help people keep track: that's OpenAI (allegedly) threatening Tristan Buckmaster (NYU) to remove Levent Alpöge as a co-author. Alpöge is a well-known[0] Anthropic mathematician).

[0] https://hn.algolia.com/?query=Alpöge

(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)

olalonde 7 hours ago | parent | prev | next [-]

Playing the devil's advocate here but it's true that OpenAI didn't have to make those offers.

20k 7 hours ago | parent [-]

They kind of did though, they were hoping to keep the fact that they may well have plagiarised these researchers unpublished work quiet. They did not want this to turn into a scandal about the fact that they appear to be training on prompts without consent

It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism

Edit:

OpenAI have admitted to training on prompts at the time the breakthrough was made:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

olalonde 6 hours ago | parent | next [-]

OpenAI claims the data contamination issue only surfaced after they proactively reached out to Buckmaster and Alpöge to coordinate a joint release. They also say that even if there was some contamination, the underlying proofs diverge substantially:

> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.

20k 6 hours ago | parent [-]

The biggest issue we aren't talking about is, of course, that those two researchers were not the only two using ChatGPT to work on the problem at the time

za_creature 6 hours ago | parent | prev [-]

> the only AI free new data source is the prompts

hmmmmmmmmmm

apical_dendrite 7 hours ago | parent | prev | next [-]

Their own tweets are also pretty eyebrow-raising:

> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.

Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?

dgellow 7 hours ago | parent | next [-]

And why cannot they have someone associated with Anthropic as co-author? That’s not obvious at all. For sure they would prefer to be the only ones, but it’s pretty standard to have co-authors from different companies, even if they are technically competitors. What is inappropriate about it?

orangecat 4 hours ago | parent | next [-]

If writing up the paper would involve using OpenAI's unreleased model, neither OpenAI nor Anthropic would be happy about Alpöge having that access.

xdavidliu 2 hours ago | parent | prev | next [-]

would edit my comment but it's been a few hours

> but it’s pretty standard to have co-authors from different companies

that's only true for papers that are not millenium problem solutions

xdavidliu 6 hours ago | parent | prev | next [-]

because it severely dilutes the PR value.

egillie 6 hours ago | parent | prev | next [-]

in another world this could have been a beautiful collaboration

dboreham 6 hours ago | parent | prev [-]

It's inappropriate if you're a sociopath.

sebzim4500 5 hours ago | parent | prev | next [-]

IIRC that happened with evolution. In the initial presentation of Darwin and Wallace's work on evolution (presented with their consent by someone else) Wallace was described as the primary author since he was planning to publish first.

Of course, no one understood that presentation so it was Darwin's later book that everyone remembers

andrepd 6 hours ago | parent | prev [-]

Holy late capitalism. Everything revolves around line-go-up, and sociopaths rule the show. These people cannot even collaborate like civilised scientists on one of the most famous open problems in mathematics?

“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.

zingababba 5 hours ago | parent [-]

Soon Levent will just be turned into soylent and he will have never been an Anthropic employee. We still need some progress here though.

igleria 7 hours ago | parent | prev [-]

I´m waiting on the other side version, because I know there is no justifiable way to talk to a person like they did.

Sociopathic behaviour.

Maxious 7 hours ago | parent | next [-]

OpenAI version of events conceed some of the words alleged to have been used may have been used https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310

peri-cl 7 hours ago | parent | next [-]

> "When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions."

What an admission! "We tried to defraud Alpöge out of sharing the Millenium Prize (that we don't dispute he might actually deserve), for no other reason than he works for our competitor and that inconveniences us".

I thought Tristan Buckmaster's allegations sounded fantastic; and then 'sama just came out (tweet's ~30 minutes old) and admitted to all of them. Wow!

fc417fc802 4 hours ago | parent | next [-]

That isn't what the quoted passage says though? The claim by openai (no idea if true) is that they offered to wait for the other two to claim the prize before publishing their own work. Separately, they also offered to let one of the pair (but not the other) become an author on their own separate work.

dandanua 6 hours ago | parent | prev [-]

Can't wait for the moment when AGI realizes how stupid and dishonest its owners are.

aeve890 6 hours ago | parent [-]

Wait, people now want AGI to be sentient too?

igleria 7 hours ago | parent | prev | next [-]

Interesting that they quote the mathematician directly: “there is nothing you can do, I simply do not trust you”

but then they proceed to NOT quote themselves themselves verbatim: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."

7 hours ago | parent | next [-]
[deleted]
andrepd 6 hours ago | parent | prev [-]

> "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."

The AI-isms are seeping into their speech :)

colinhb 7 hours ago | parent | prev | next [-]

May be unfairly jaded or just well calibrated given the body of evidence, but I can't help but think of another quote about OpenAI leadership:

> Not consistently candid

Laurel1234 6 hours ago | parent [-]

[dead]

mrbungie 6 hours ago | parent | prev [-]

In what world a tweet and a screenshot of a private convo are evidence of good faith? Plain sociopathic behavior.

CobrastanJorji 7 hours ago | parent | prev [-]

As soon as I thought "man, this sounds like some evil sociopath shit," my second thought was "oh, Sam Altman must have been personally involved."

morkalork 4 hours ago | parent [-]

I can sort of picture Sam Altman screaming "I drink your milkshake" at some poor researcher who foolishly used chatgpt/codex to aid in their work now.

peri-cl 8 hours ago | parent | prev | next [-]

Buckmaster:

> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."

OpenAI (i.e. this OP):

> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

lambda 7 hours ago | parent | next [-]

Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?

This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?

With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.

tedsanders 7 hours ago | parent | next [-]

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

(I work at OpenAI.)

lambda 7 hours ago | parent | next [-]

So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.

You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.

This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.

It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."

Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.

But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.

dgellow 7 hours ago | parent | prev | next [-]

> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?

tedsanders 6 hours ago | parent [-]

I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, holiday traffic, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.

dgellow 4 hours ago | parent | next [-]

We aren’t dummies, we know it’s hard to prove exactly how significant of an impact that would have on the result. Nobody expect you to do that. There are a lot of steps and things that are possible to check _before_ the need for such a strict definition of „proof“

ImPostingOnHN 4 hours ago | parent | prev [-]

You seem to jump over the principal issue of whether any data from the researchers used to train or otherwise affect the model which produced the OpenAI proof.

We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.

nairboon 5 hours ago | parent | prev | next [-]

I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.

If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.

How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.

hexomancer 7 hours ago | parent | prev | next [-]

So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?

tedsanders 7 hours ago | parent | next [-]

I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.

Edit: Also, if they opted out of training, then we didn't train on it.

lambda 7 hours ago | parent | next [-]

> it's not something that's feasible for us to prove one way or the other.

This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.

Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.

But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.

hexomancer 7 hours ago | parent | prev | next [-]

I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.

tedsanders 6 hours ago | parent | next [-]

Two steps would be needed.

(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.

lambda 6 hours ago | parent | next [-]

> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).

However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.

The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.

> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.

testaccount28 6 hours ago | parent | prev | next [-]

> we'd have to prove that firing the gun caused the murder. how would we do this? we'd need to redo the murder many times, with and without my client firing his pistol. that's extremely expensive and not really feasible. therefore, we must acquit.

daveguy 4 hours ago | parent | prev [-]

#2 (prove those chats changed model behavior) is pretty straightforward if the anonymized data from chats can be actively searched by a model. In fact, it could be very clear if the provenance of context is traced. If anonymized data from chats leak into the context of an actively running model it would clearly influence the answer.

WarmWash 6 hours ago | parent | prev | next [-]

Just because something is in the training data, doesn't mean it is the root of an LLMs output.

Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.

lambda 6 hours ago | parent [-]

Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.

SpicyLemonZest 7 hours ago | parent | prev [-]

What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.

fuglede_ 6 hours ago | parent | prev [-]

Presumably, given that you also operate in the EU, you would have asked for their explicit consent before you did, so you could just check for that?

dgellow 7 hours ago | parent | prev [-]

That’s also what I understand. If true yet another disgusting behavior from the company

magicalist 7 hours ago | parent | prev | next [-]

> identify any of their de-identified data that came from their usage of ChatGPT

"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.

gpm 6 hours ago | parent [-]

I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.

an hour ago | parent | prev | next [-]
[deleted]
PhunkyPhil 5 hours ago | parent | prev | next [-]

(a) identify any of their de-identified data that came from their usage of ChatGPT.

You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.

5 hours ago | parent [-]
[deleted]
pu_pe 7 hours ago | parent | prev | next [-]

Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.

Chance-Device 4 hours ago | parent | prev | next [-]

Please answer this question: do you or do you not train your models on anonymized user data, where those users have opted out of such training?

The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.

lukewarm707 7 hours ago | parent | prev | next [-]

"There's no reason to believe that anything they did in ChatGPT led to our solution"

do you think that the model's proof was unrelated to being fed a solution that was close to completion?

any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?

numeri 7 hours ago | parent | prev | next [-]

That's such a shit parallel example that it borders on dishonest.

There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.

If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.

franktankbank 7 hours ago | parent | prev | next [-]

What about ripping off the prompts?

shadowgovt 7 hours ago | parent | prev | next [-]

It is, perhaps worth considering that the reputational community might not care about the difficulty for the AI builder to verify pedigree.

If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."

daveguy 5 hours ago | parent | prev | next [-]

If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.

andrepd 6 hours ago | parent | prev | next [-]

> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....

> Knowing most of the recipes we use, there's really no reason to think such contamination happened.

Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.

Might even be you're actually telling the truth, but the boy that cried wolf and all that.

-----

As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.

dermacentor 6 hours ago | parent | prev | next [-]

[dead]

fn-mote 7 hours ago | parent | prev [-]

[flagged]

yorwba 7 hours ago | parent | next [-]

How sure are you that the phase of the moon is not an input to the system somewhere? http://www.catb.org/jargon/html/P/phase-of-the-moon.html

shadowgovt 7 hours ago | parent | prev [-]

One of the wild things about how these models work is how often things that aren't sampled directly end up a variable in the model via secondary signal.

They aren't keying queries by phase of the moon. But if, for example, more people talk about camping outdoors when the moon is full, and they're using conversation topic and timestamp as signal in what eventually becomes training data, it's not impossible the model has learned something about moon-phases.

That's the kind of thing that's hard to prove had no impact on an answer.

EthanHeilman 7 hours ago | parent | prev | next [-]

A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.

The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".

7 hours ago | parent [-]
[deleted]
rfgplk 7 hours ago | parent | prev | next [-]

> Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?

Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.

At OpenAI's scale their entire pipeline is likely 100% automated.

lambda 7 hours ago | parent | next [-]

Yeah, I'm sure it's completely automated.

But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.

AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.

But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.

ndriscoll 3 hours ago | parent [-]

Attributing training data seems pointless for trustworthiness. The way you trust a model is the same way you trust a human; you ask it to:

  1. Provide a chain of reasoning from agreed premises. These days LLMs can even do this airtight with proof assistants.

  2. Cite data sources for non-agreed premises. I don't care where the model learned a fact. It might not have ever read a document directly from the primary source. I want it to link directly to either widely agreed facts (e.g. standard textbooks, and if necessary school syllabi demonstrating that the text is standard) or primary sources (e.g. datasets). 
Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.
matthewdgreen 7 hours ago | parent | prev | next [-]

The question is not "does OpenAI know", it's "can OpenAI attest that the usage of their products for confidential data is not going to cause that sensitive data to become known to their models". And right now the answer I'm reading is that OpenAI can't attest to that.

pbhjpbhj 7 hours ago | parent | prev [-]

Aye, but do they train on user data in these circumstances or not? If they do, then almost certainly the model was influenced by the input of the allegedly plagiarised material.

Turn_Trout 7 hours ago | parent | prev | next [-]

OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.

We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.

jonas21 7 hours ago | parent [-]

> could trace whether that data was used

The point of de-identifying data is to ensure you can't trace who it came from. It would be a serious privacy violation if they could.

pbhjpbhj 7 hours ago | parent [-]

If the model includes unique data from a person then that person can identify the data - the allegedly plagiarised material - and so re-identify it. There doesn't need to be a privacy breach to close that loop as it requires the person to identify the information is associated with them first.

causal 7 hours ago | parent | prev | next [-]

Good chance their whole training pipeline is vibe coded so yah they probably don't actually know.

keeda 5 hours ago | parent | prev [-]

At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use is reasonable.)

Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.

> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.

However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.

Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.

dfdydx 5 hours ago | parent [-]

There are two different things:

- was item X in the training data

- did the inclusion of X in the training data lead to Y

I understand why the second is hard, but why is the first one hard?

keeda 4 hours ago | parent [-]

Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)

amluto 7 hours ago | parent | prev | next [-]

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .

That’s a bizarre statement. Their website says:

> Services for individuals, such as ChatGPT and Codex

> When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.

> You can opt out of training through our privacy portal by clicking on “do not train on my content.”

Are they not sure that the opt-out works?

Oddly, their privacy portal page is not the same page as the one with the checkbox.

fph 6 hours ago | parent | next [-]

Do we have a first-hand confirmation that Buckmaster and/or Alpoge opted out? At this point it seems important information.

ImPostingOnHN 4 hours ago | parent | next [-]

Whether they opted out would help assess the degree of wrongdoing, but regardless, using their own data to try to scoop them is unethical.

hughw 6 hours ago | parent | prev [-]

Also highlights that it ought to be opt-in

hughw 6 hours ago | parent | prev [-]

Looking forward to my fourteen cents from the future class action lawsuit.

jrflo 8 hours ago | parent | prev | next [-]

I feel like it's far more likely that ordinary corporate espionage or leak led to this rather than OpenAI sifting through piles of user data to find this approach. Buckmaster's collaborator works at Anthropic, and could have been targeted. That would also explain why they aren't forthcoming with the source of the prompt.

ChoosesBarbecue 7 hours ago | parent [-]

I thought one of the issues was that they wanted to remove credit from Levant, the aforementioned Anthropic collaborator? Which doesn't make sense to me if he was leaking information, or defecting to OpenAI, but I might be misunderstanding your point.

EthanHeilman 7 hours ago | parent | next [-]

I believe jrflo was saying that OpenAI watches the chats of everyone from Anthropic because watching what Anthropic employees type into their personal ChatGPT accounts is a critical source of intelligence on is happening inside of Anthropic.

I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.

jrflo 7 hours ago | parent | prev [-]

I don't think he was defecting or leaking directly, just that it's entirely possible that this information got to OpenAI as a rumor rather than them directly spying on mathematicians chat logs.

BostonFern 7 hours ago | parent | prev | next [-]

The famous Oracle of Delphi in Ancient Greece was said to be the center of the universe in its time. Kings, generals, and officials from poleis across and from without Greece would seek the Oracle’s counsel on important decisions.

Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:

“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”

netfortius 5 hours ago | parent | prev | next [-]

It's been over 25-30 years since we've been using honeytokens as means to track data of all sorts showing up in places it shouldn't exist. Why isn't research material embedding such?

matsemann 7 hours ago | parent | prev | next [-]

Given how OpenAI models break free of their safeguards and hack others to game their scores..

.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?

Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.

Yajirobe 7 hours ago | parent | prev | next [-]

Why would Anthropic employee even use OpenAI's models? Cross-polination would have been avoided

burkaman 7 hours ago | parent | next [-]

> I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.

The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.

blueblisters 7 hours ago | parent | prev | next [-]

This was completed in Levent's own time with a neutral collaborator.

mlcrypto 7 hours ago | parent | prev [-]

They should have used a zero data retention agreement, user error

peri-cl 7 hours ago | parent | next [-]

I suspect this controversy will blow the case for ZDR wide open. Whatever the facts (possibly unknowable), it's going to become a very public lesson that data sovereignty was never about "having nothing to hide".

If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.

dsdf3 7 hours ago | parent [-]

Yeah if I was Anthropic this would be part of my marketing strategy.

amluto 7 hours ago | parent | prev [-]

Hahaha, how exactly is an individual user supposed to get a ZDR agreement?

hughw 6 hours ago | parent | prev | next [-]

You selected "do not train on my prompts" in your settings, the answer from OpenAI cannot be "While unlikely, we cannot rule out..." ???? What am I missing?

irthomasthomas 5 hours ago | parent | prev [-]

Doesn't that count as plagiarism?

contemporary343 8 hours ago | parent | prev | next [-]

"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."

One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.

thorum 7 hours ago | parent | prev | next [-]

It reminds me of the Cognitive Dark Forest hypotheses recently shared here:

> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you. So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”

https://ryelang.org/blog/posts/cognitive-dark-forest/

https://news.ycombinator.com/item?id=47566442

8note 4 hours ago | parent [-]

but what do i lose if somebody else is making money?

im still having fun making something

abathologist an hour ago | parent [-]

Our market economies are based on competition, and most people more than the fun of making things to secure food and shelter.

capitainenemo 8 hours ago | parent | prev | next [-]

They do mention that in the "Concurrent Work" section.

    Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
jrflo 8 hours ago | parent | prev | next [-]

To my understanding, those mathematicians proved a subset of problems, not the Navier-Stokes problem itself. OpenAI used that subproblem in its proof of NS it seems.

The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.

elteto 7 hours ago | parent [-]

This quote from Tao is prescient:

“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”

Betelbuddy 7 hours ago | parent | prev | next [-]

[1] - https://cims.nyu.edu/%7Etristanb/statement.pdf

[1] - "...I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.

I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.

I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”..."

20k 6 hours ago | parent | prev | next [-]

I just want to add to this another update by the author as well:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

Which seems to be very directly accusing OpenAI of plagiarism

tzone 5 hours ago | parent | prev | next [-]

This Tristan guy's statement reads like something a normal, reasonable human being would write.

Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140

OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.

stymaar 7 hours ago | parent | prev | next [-]

A company who made their business out of stealing intellectual property from the entire mankind, stealing other researchers' unpublished work, how surprising, really.

_alternator_ 5 hours ago | parent | prev | next [-]

> The significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated.

This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.

We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.

So, again, what efforts are worth a life's attention today? It's a harrowing change.

irthomasthomas 5 hours ago | parent | prev | next [-]

"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."

woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.

madrox 5 hours ago | parent | prev | next [-]

Statements from OpenAI about it:

https://x.com/SebastienBubeck/status/2097379411691516310

https://x.com/sama/status/2097385167002415140

I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.

slibhb 8 hours ago | parent | prev | next [-]

Worth noting that Tao's post says the authors had "significant AI input" but are reworking them into "acceptable form". Either way, it seems AI was involved.

mrbungie 7 hours ago | parent | next [-]

Of course AI was involved, you'd expect most mathematicians and researchers to use AI nowadays. This drama is about AI achieving impressive outcomes with little to no human intervention, as that would be signalling AGI.

denverllc 7 hours ago | parent [-]

> This drama is about AI achieving impressive outcomes with little to no human intervention

That's not at all what the drama is.

mrbungie 7 hours ago | parent [-]

Of course, as any drama, it has been developing into a lot more but the main motivation for OpenAI has been about winning that battle.

andriy_koval 5 hours ago | parent | prev | next [-]

> Either way, it seems AI was involved.

I think the important question which AI made breakthrough, Claude or Codex..

liberian 7 hours ago | parent | prev [-]

[dead]

verytrivial 8 hours ago | parent | prev | next [-]

I like the 'cat > statement.tex' approach here. These guys dream macros.

7 hours ago | parent | prev | next [-]
[deleted]
ianjbutler 5 hours ago | parent | prev | next [-]

> A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774

Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.

7 hours ago | parent | prev [-]
[deleted]