You’re complicating things.
There’s no reward for prosocial in llm rl as compared to other targets.
Humans have it since prosocial and others have evolutionary reward signals that do.