| ▲ | hypfer 2 hours ago | |
I'll be the one to ask the obvious question: What does this mean for compute workloads? Specifically, LLM inference. Does it mean anything at all, or is this purely a games-thing? | ||
| ▲ | skew-aberration 2 hours ago | parent [-] | |
I doubt it makes much of a difference, and you can always manually manage what data lives in the GPU when if you 100% have to overcommit. Games have a much larger and more diverse set of objects in the VRAM, and their usage is less predictable, so manual scheduling of the memory is infeasible typically. | ||