| ▲ | efebarlas an hour ago | |
i think people say that thinking that only training to produce the next likely word would end up producing some local minimum word that generally fits but doesn't actually lead to intelligent thought. that feels like a misunderstanding of how the loss function behaves when used within a sequence | ||