| ▲ | numpad0 4 hours ago | |
Those labs publicly said during GPT-3/4 era that the optimal epoch count, or dataset repetition count, for foundation model training, is one. So it's a forward 1-pass compression. But it's a black box! Nobody knows whats going on inside! It's all transformative! Sure... | ||