| ▲ | forgotTheLast 3 hours ago | |
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens. | ||
| ▲ | clhodapp an hour ago | parent [-] | |
Yup. You basically just need something for probability to push off of | ||