| ▲ | r_lee a day ago | |
if it's agentic stuff, they likely aren't hammering a core constantly and they will maybe sit idle quite often between model requests, so it makes sense. I do wonder how much memory they allocate to each one though. it's just very efficient use of shared cores that is required to make these kinds of workloads cost efficient | ||
| ▲ | eru 20 hours ago | parent [-] | |
If they sit mostly idle, you can swap out a lot of the memory to SSD, I guess. | ||