> future models can be trained to be even more aware of external knowledge
Then you need longer contexts, which is proving to a much more stubborn problem than general knowledge compression.