| ▲ | dash2 an hour ago | |||||||
He didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business. Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this. | ||||||||
| ▲ | Traster 31 minutes ago | parent | next [-] | |||||||
Parse that statement more carefully. > I was told the model did not look up user data. The naive way to read this is "Nothing you guys did influenced the way our model got to the solution". The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem". | ||||||||
| ▲ | revolvingthrow an hour ago | parent | prev | next [-] | |||||||
It would be shocking if it wasn’t trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? It’s so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?" There’s potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber? The only unlikely part is the timeline - your sessions from a week ago probably haven’t made their way into the model. It’ll just take a while longer, and will be massaged just enough so that it isn’t really your exact session word for word so you can’t sure as easily. | ||||||||
| ▲ | johnnienaked an hour ago | parent | prev | next [-] | |||||||
It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course. Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you? The whole reason they have subs is to train them on YOUR WORKFLOWS lol | ||||||||
| ||||||||
| ▲ | actionfromafar an hour ago | parent | prev [-] | |||||||
Couldn't the Enterprise have a different fine print? | ||||||||