So would it be a 42 trillion parameter model, because that's how many tokens there are in the training data?
Does that mean, it's not compressed anymore?