Remix.run Logo
mark_l_watson 5 hours ago

This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.

jonas_scholz 5 hours ago | parent | next [-]

I really hope they dont stop at the small models though! The bigger ones that dont fit on a single GPU are more interesting I think

mips_avatar 2 hours ago | parent | prev [-]

Problem is right now the biggest GPU boxes they have is single rtx pro 6000s.