| ▲ | torginus 5 hours ago |
| With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it. This could mean every potential serious customer would have no option but to seek alternatives to these online services. |
|
| ▲ | lynndotpy 4 hours ago | parent | next [-] |
| I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services. |
| |
| ▲ | zdragnar 4 hours ago | parent | next [-] | | Some offer zero data retention policies, but there can be weasel words. For example, on the individual pro plan, you can turn off the setting that lets them train models on your data, but they still have a section in their terms that allows them to evaluate your anonymized data for statistical and "research" purposes. You have to actually get a signed contract along with an enterprise plan that spells out exactly what they're going to use, and what settings enable what retention. https://privacy.claude.com/en/articles/10023548-how-long-do-... (see the additional info section) | |
| ▲ | ssivark 3 hours ago | parent | prev | next [-] | | What about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data. | | |
| ▲ | hazard 3 hours ago | parent | next [-] | | I worked very briefly at Baseten, and I can say that it was a perpetual annoyance (from an engineering perspective) that customers would complain about issues with their models but we couldn't actually see the inputs/outputs. I don't know about the other providers, but at Baseten they literally weren't stored anywhere. | |
| ▲ | lynndotpy 2 hours ago | parent | prev [-] | | I don't have any much exposure to the attitudes people have around them, and I haven't worked with them. So I can't really say |
| |
| ▲ | Gud 4 hours ago | parent | prev | next [-] | | No, that is not "common knowledge".
You are supposed to be able to disable that unwanted feature. | | |
| ▲ | ForHackernews an hour ago | parent [-] | | I have no inside information, but I always assume the tickboxes that "disable ____ data" from Google/Facebook/OpenAI just disconnects it from your own account, not hides it from the provider. |
| |
| ▲ | Aurornis 3 hours ago | parent | prev [-] | | > I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? The services have toggles to allow prompts to be used in the training set. There is a conspiracy theory that the toggle is a false distraction and they’re actually keeping everything, and that none of the employees involved will ever whistleblow this fact. Outside of Internet comment sections, I think most people assume these US-based companies are doing what they say. For enterprise use there are services like AWS Bedrock which have strict isolation guarantees. There are some people who still believe those guarantees are a lie, but once someone has reached that point I don’t think they trust anything that isn’t running entirely within their house. People in that category are a very small minority, but a very vocal minority. | | |
| ▲ | lynndotpy 2 hours ago | parent [-] | | The impression I have (from interacting with people IRL using OpenAI and Anthropics offerings, and how they feel about the risks involved) is just the opposite. But we probably just have different life experiences. |
|
|
|
| ▲ | Betelbuddy 3 hours ago | parent | prev | next [-] |
| Or these customers could just use AWS Bedrock...but their current CEO is an incompetent MBA unable to publicly articulate their biggest advantage, in the context of the current AI usage my companies. You have access to all the frontier models, but...your inputs are not shared with the model vendors...neither are used to train the next model. Why am I even doing the Amazon board job for them!?? |
| |
| ▲ | ballon_monkey 3 hours ago | parent | next [-] | | Bedrock is really bad. It seems like they don't host the models very well because they produce tons of bugs/errors calling the model. For example you can end up with Anthropic models not returning a stop token and you end up waiting for a timeout thinking its doing something when it isn't. | | |
| ▲ | Betelbuddy 3 hours ago | parent [-] | | Well Anthropic hosts their models at AWS, ( and at many others...) so maybe the AWS team can ask them how they do it ;-) ? |
| |
| ▲ | whatshisface 3 hours ago | parent | prev | next [-] | | Amazon is deeply invested in Anthropic and would not defame them through marketing a service whose selling point was their startup's breach of contracts. | |
| ▲ | staticautomatic 3 hours ago | parent | prev [-] | | All except Gemini which can be rather important depending on your use case. | | |
|
|
| ▲ | Aurornis 4 hours ago | parent | prev | next [-] |
| > OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work. They’re not saying that some prompts are mysteriously jumping into training data. Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers. |
| |
|
| ▲ | ronsor 5 hours ago | parent | prev [-] |
| Almost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data. |
| |
| ▲ | steveBK123 5 hours ago | parent | next [-] | | They already trained on pirated content, what makes you think they are going to honor ZDR? | | |
| ▲ | GrinningFool 4 hours ago | parent [-] | | Contractual obligations carry teeth. Scraping the internet is relatively risk-free. | | |
| |
| ▲ | Yizahi 3 hours ago | parent | prev | next [-] | | A lot substance is hinged on the exact definition of the word "data" or "user data". In the age of post-truth everyone is claiming that they keep no "user data". Except that after running it once through some transformer program it's no longer "user data", it's something entirely else and these corpos gave ZERO promises regarding such laundered/transformed data at all, ever. | |
| ▲ | nrmitchi 5 hours ago | parent | prev | next [-] | | The guarantee on this is a (contractual) “trust me bro”, and a right to try to sue a multi-trillion-dollar company who will absolutely drive you into the ground with legal red tape. If you are big enough to be able to withstand that, you’re already running (or trying to run) your own/open-weight models. | | |
| ▲ | enugu 5 hours ago | parent [-] | | Doesn't Amazon Bedrock change this, since OpenAI does not have access to the data? | | |
| ▲ | nrmitchi 5 hours ago | parent [-] | | Well that is a different thing and an entirely different provider than OpenAI/Anthropics ZDR promise. |
|
| |
| ▲ | torginus 5 hours ago | parent | prev | next [-] | | Just a thought experiment: considering training seems to be 'fair use', I wonder if they trained a tiny model to retain key info from your prompts, would mean that this would still constitute fair use, and allow them to legally claim they don't retain your data. | | |
| ▲ | ronsor 5 hours ago | parent [-] | | ZDR is shorthand for a more specified agreement of "we don't do anything other than generate your output tokens", so no. Besides, true ZDR is usually offered by third-parties with deals to host OpenAI models, such as Amazon (AWS Bedrock) and Microsoft (Azure). |
| |
| ▲ | applfanboysbgon 5 hours ago | parent | prev | next [-] | | ZDR is based on the exact same pinky-promise as training opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data. | |
| ▲ | pennomi 5 hours ago | parent | prev [-] | | Where nothing is retained at all, allegedly. |
|