| ▲ | olmo23 6 hours ago |
| > While unlikely, we cannot rule out that de-identified data derived from [Buckmaster and Alpöge’s] usage of our products helped improve our models. Well yeah, if you use the free product they train on your data, ... I thought this was widely understood? |
|
| ▲ | greggoB 6 hours ago | parent | next [-] |
| If you read Buckmasters statement, he specifically notes that they used the paid subscriptions, iirc. |
| |
| ▲ | AlanYx 3 hours ago | parent | next [-] | | >he specifically notes that they used the paid subscriptions, iirc. Where are you seeing that? He only says "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments." I might be missing something, but he doesn't seem to confirm that he opted out, at least in the written writeup, maybe he has on social media? He also likely had early access to Astra given the timing, and I thought early access customers couldn't opt out? (Am I wrong about that?) | | | |
| ▲ | aenis 5 hours ago | parent | prev | next [-] | | The paid subscriptions have opt-out for sharing data for training purposes. I think it's on by default. | | |
| ▲ | y-curious 4 hours ago | parent [-] | | I don’t know because I use Anthropic, but I would eat my hat if this was on by default. | | |
| |
| ▲ | Tenemo 5 hours ago | parent | prev [-] | | Paying for an account doesn't opt you out by itself, right? Has he stated anywhere that he actually opted out? But if not, then I also don't understand why OpenAI's communications about this have been so vague, they could've just said that he didn't opt out, using those chats in training data follows their ToS and that's it (whether that's "fair" is a separate discussion). | | |
| ▲ | dgellow 4 hours ago | parent | next [-] | | We don’t need to guess, OpenAI pretty much indirectly they had the chats in their data set. OpenAI responses are the most suspicious part of that whole controversy, the fact they do not provide straight answers is not a sign of a good faith actor here | |
| ▲ | Topfi 5 hours ago | parent | prev [-] | | There has been no statement either way, as far as I could find beyond them only using commercially available models, though given Alpöges employer, I'd be surprised if they didn't opt out. In any case, for such work, ZDR or self-hosting seem to be an absolute must now. Unless OpenAI can show that training was permitted, this will erode the limited trust that many users have had in such toggles and may lead to further, uncomfortable inquiries. |
|
|
|
| ▲ | Topfi 6 hours ago | parent | prev | next [-] |
| They did pay [0] and substantially by the sound of things: > I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI. [0] https://cims.nyu.edu/~tristanb/statement.pdf |
|
| ▲ | Arodex 4 hours ago | parent | prev [-] |
| Then OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves. |