Remix.run Logo
▲ demibabs a day ago

Is this an example of the bitter lesson? Could diffusion models not be this good with equal scaling/compute as LLMs?

▲moojacob a day ago | parent [-]

Well the LLMs are much bigger than diffusion models so I think that's the bitter lesson. You could scale compute for diffusion models though.

▲demibabs a day ago | parent [-]

Well no, the bitter lesson isn’t the scaling laws themselves. It’s that approaches which can take advantage of scaling laws will ultimately beat ones that can’t.

▲moojacob a day ago | parent | next [-]

You learn something new every day! thanks

▲E-Reverance a day ago | parent | prev [-]

Scale with compute*

not explicit to reference scaling laws of training