Remix.run Logo
deepnet 10 hours ago

An insightful post by one of the AI ‘godfathers’.

Bengio outlines the dangers of the current situation and what has led to these dangers.

He also proposes solutions in the last paragraph.

Well worth a read, right to the end.

Hopefully a stimulating debate on these issues will ensue in these comments.

We do need to consider the points Bengio makes and with some urgency.

Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.

As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:

“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”

Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.

I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.

Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.

Bengio asserts that the way LLMs are trained is flawed if we want safety.

He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.

In short he presents clearly the case for how plausibly unsafe the current course is.

He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.