| ▲ | Iolaum 5 hours ago | |
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell. | ||
| ▲ | ismailmaj 3 hours ago | parent [-] | |
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data. | ||