Remix.run Logo
▲ r_lee a day ago

if it's agentic stuff, they likely aren't hammering a core constantly and they will maybe sit idle quite often between model requests, so it makes sense. I do wonder how much memory they allocate to each one though.

it's just very efficient use of shared cores that is required to make these kinds of workloads cost efficient

▲eru 20 hours ago | parent [-]

If they sit mostly idle, you can swap out a lot of the memory to SSD, I guess.