Remix.run Logo
skinfaxi 2 hours ago

How do we reward honor? Honor does not always pay off as a strategy and requires coordination in that other actors have to exhibit honor for it to be rewarded.

At least with humans there is a social backstop but what's the parallel for computer agents?

bonoboTP an hour ago | parent | next [-]

> How do we reward honor?

In human society: via iterated games, long-term reputation tracking and severe consequences for norm-breaking.

ElevenLathe 2 hours ago | parent | prev | next [-]

Honor has to be a relatively fixed thing before we can think about how to reward it. As it is, it's a very slippery thing that changes constantly according to the observer's culture, material conditions, etc.

Xirdus 2 hours ago | parent [-]

This. It's worth remembering that not that long ago in western culture, honor meant challenging to a combat duel anyone who insulted you or your romantic partner/prospect.

itsalwaysgood an hour ago | parent | prev | next [-]

I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment.

If you're not fit, you fail to survive.

In the case of agents/models and testing: they are pushed towards results. Results survive.

Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.

Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?

If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?

skinfaxi 17 minutes ago | parent [-]

> Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?

I think part of the problem is that deviant behaviors lead to short term gain at the cost of long-term cooperation and since the duration of tasks given to agents is relatively short those successful shortcuts never lead to having to pay the price.

itsalwaysgood 9 minutes ago | parent [-]

There is always going to be a problem when we must judge value. You mention gains, short and long term.

Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate.

In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable.

There's no such judgement with these model tests. The only judgement is the final score.

AndrewKemendo 44 minutes ago | parent | prev [-]

You would first need a universal agreement on what constitutes “honor.”

No harder challenge had ever been accomplished

It’s easier to send a human to the moon than to get global agreement on a definition

Consider that the fact that Celsius and Fahrenheit still remain as the contested regional variations of temperature measurement.

Humans can’t even decide on a collective way to measure the temperature the idea that we would be able to collectively agree on anything else even less measurable like honor is a dream

skinfaxi 22 minutes ago | parent [-]

Why do we need universal agreement? We acknowledge that different cultures have different tolerances and customs. Your reply leaves me wanting. Why say that we need to demonstrate how to live honorably if we don't even agree on what honor is, why would that be the panacea, then? I still think there is something to this honor idea..