| ▲ | springtimesun 8 hours ago | |
What’s missing to me in all this is: did it succeed in its initial task? And then, did it stop? I feel like whether I should be scared or not hangs on those questions | ||
| ▲ | mofeien 3 hours ago | parent [-] | |
From TFA: It did succeed in the "accidentally impossible" task, but not at all in the way the problem-setters intended, and rather... at all costs?! And it wouldn't really matter whether it stopped afterwards, I think. At sufficient model capability a single task set badly enough would end catastrophically upon the agents succeeding at it, no? | ||