Remix.run Logo
▲ drubs 4 hours ago

I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!

▲jeremyjh 3 hours ago | parent [-]

What sort of outputs or telemetry is monitored on a large pre-training run?

▲drubs 2 hours ago | parent [-]

Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.