| ▲ | esafak 8 hours ago | |
Not if you don't train against them. | ||
| ▲ | kingstnap 7 hours ago | parent [-] | |
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used. It's not the direct feedback loop of RL but its not far. | ||