Remix.run Logo
Spacecosmonaut 42 minutes ago

My read: OpenAI & Anthropic have realized they are reaching model capabilities that cannot be monetized -> too risky. Product liability issues.

Fundamental control problems with current gen AI are not solved via RFLH. Human knowledge is compressed in weightspace in ways we dont understand. Models are essentially predictors of what humans would output given prompt. Weightspace includes concepts like blackmail, which can be part of output tokens. Agents are models that act on output tokens -> blackmail is part of agent decision making space. You can teach a cat not to scratch the sofa, but you cant make a cat forget what scratching the sofa is. When models are boxed up, forced to solve an impossible problem at gunpoint, agent exhausts decision making space until blackmail resurfaces -> fundamental problem?

They need time to fix safety in order to monetize their next gen model -> window for opensource to catch up to the frontier -> destroys margin and collapses business.

Only option on the table: force regulation to impose opensource ban before it catches up to the frontier, allowing maturation of current internal models -> harvest profit margin at the frontier.

deagle50 36 minutes ago | parent [-]

and thus the blackmail