| ▲ | arkmm 3 hours ago | |
Very cool! Can you say a little bit about the size of the DPO training examples and how long training took? | ||
| ▲ | simedw 3 hours ago | parent [-] | |
For DPO I only had around 700 preference examples, so not much data at all. That took about 12 minutes to train on a single GPU. Pretraining was obviously a a lot slower, the 125M model took roughly half a day. | ||