Remix.run Logo
anon373839 a day ago

I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.