it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.