Remix.run Logo
impossiblefork 3 hours ago

Morally I agree, but since there's probably a lot of LLM text in the training data, distilling on another model will probably make your model copy the values encoded into the other model as well, even in cases where you only distill on value-neutral stuff.

By copying their programming style, you'll move the model towards that way of writing, which will move the model towards the values expressed in those documents.

I feel that Deepseek v4 got so claudified at the end that it was like Claude.