Remix.run Logo
8note an hour ago

250k because thats what the smaller models supported and theres better training data in that part of the window

there have been some papers suggesting that the useful context is even smaller, and stays fixed as you change the context window size.

as a more general case though, i think the possibilities for what youd need to include in training to have many different paths of text be well represented enough over the long window means the later tokens will almost always be a lot more random than early ones?

there's a sheer amount of bits problem.