Remix.run Logo
▲ eigenspace 7 hours ago

I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.

The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.

If they can do that, they'll have customers.

▲hajile 6 hours ago | parent [-]

The Navier-Stokes fiasco made me push for local/controlled models very hard. "Can't rule out" that they stole data (backed up by their backdoor offers of sharing credit).

If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.

If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).

Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.

▲sashank_1509 6 hours ago | parent | next [-]

I would also add to this, there are ways to use customer data to improve your model outside of just using it as “training-data”.

A simple loophole, use the code to create an RLVR environment where the resultant code is the end goal / max reward. Technically the customer data is never trained upon, but effectively you’re using it. Even better, use the code as a seed to generate synthetic data similar to it and use that synthetic data as rewards in an RLVR model.

Unless you can host the ChatGPT model on your own servers, which I know some enterprises are doing, I don’t think there’s any hope of protecting your data / competitive advantage from these frontier companies. Better to be paranoid, than be commodified by these companies.

▲phillmv 5 hours ago | parent | prev | next [-]

tbh to me if the AI company writes all of your code & your eng don't even review it anymore then… the AI company _controls your company_. maybe that's ok if you make widgets but less ok if you do anything in dev tooling, security, or [insert market they may suddenly decide to compete in].

▲FabHK 6 hours ago | parent | prev [-]

If I'm not mistaken, they did rule out that they stole data from the mathematicians in question.

▲airstrike 4 hours ago | parent | next [-]

IIRC they ruled out that humans knowingly stole that data, but they didn't rule out that the AI agent might have

▲user43928 22 minutes ago | parent [-]

I don't think that's correct.

> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training. The OpenAI internal model used for this result was developed through large-scale reinforcement learning on top of a previously pretrained model. Our proofs also differ significantly. In the Euler case, Alpöge and Buckmaster proved a result with external forcing, while OpenAI’s system proved a result without external forcing.

▲eigenspace 6 hours ago | parent | prev [-]

"We investigated ourselves and determined that we are not to blame"

▲chihuahua 5 hours ago | parent [-]

Surely Sam Altman would not lie to anybody.