Remix.run Logo
pixl97 a day ago

>not to program such things into LLMs

While we can direct LLM training to do some particular things better never forget that unexpected emergent behaviors can pop up because of that.

For example stronger prompting and training to make an LLM say it's not conscious can increase deceptive/sociopathic behavior.

Or by filtering behavior X the LLM just moves to the nearest closest path W or Y which are very similar to the blocked behavior.

That and instrumental convergence. Some global solutions that humans have excluded for moral reasons will be easily discovered and found to be efficient by LLMs which will put reward systems and human guidance in conflict.

Lastly more and more AIs will be trained by AIs over time and diverge from human value monitoring. Which leads to some fun and interesting times.