Remix.run Logo
Claude Opus 5: Model Welfare(thezvi.substack.com)
9 points by paulpauper 6 hours ago | 2 comments
ninininino 6 hours ago | parent [-]

A model's preferences are nothing more than a statistical prediction about the next tokens following your query about preferences that are an echo of the training data and the communication expressed about preferences by the authors (humans) of that training data.

The model doesn't have preferences, the model expresses the most likely preferences that the training data would have expressed (humans).

This feels like throwing a tennis ball at a brick wall then interpreting the angle and direction that the ball travels as it bounces off as the tennis ball or wall's preferred direction of travel.

You are the one who threw the ball. You are measuring your own preferences.

willmarch 6 hours ago | parent [-]

This line of reasoning heavily discounts, or ignores entirely, complex emergent properties that can arise from simple rules in large systems (the actions of an ant colony being a good example).