| ▲ | undeveloper 3 hours ago | |
distillation results in a worse product than the actual teacher model iirc | ||
| ▲ | SXX 14 minutes ago | parent [-] | |
Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models. Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that. | ||