| ▲ | summarybot 2 hours ago | ||||||||||||||||
Yesterday I came up with an idea that I sent to some researchers at the different AI labs via email: Rather than train the model on one score, track two scores. The first score is the Short-term-objective-score (STOS) and the other, more important one, is the EAOS Ethically-aligned-outcome-score. Every trajectory can be evaluated on whether or not it has a high enough EAOS to be considered acceptable. If the model does some task and has a very high STOS but very low EAOS, like modifying game code to win at a game rather than playing by the rules, it is unacceptable. Models going forward must all have an ethics evaluation in tandem with objectives evaluation, and only when the ethics value is high enough should actions be considered successes. | |||||||||||||||||
| ▲ | antonvs an hour ago | parent | next [-] | ||||||||||||||||
What happens if we do the same for CEOs? | |||||||||||||||||
| |||||||||||||||||
| ▲ | christkv an hour ago | parent | prev [-] | ||||||||||||||||
Whats the definition of EAOS though who's ethics? Greek-Roman, Western, Islamic, Buddhist, Hinduism, Human rights (western values).. | |||||||||||||||||
| |||||||||||||||||