Remix.run Logo
What happens when a GPU writes memory(blog.doubleword.ai)
47 points by ibobev 2 days ago | 1 comments
mailonce 4 hours ago | parent [-]

I've been running local image models on an older laptop recently, and memory behavior surprised me more than raw inference time. One experiment briefly pushed private memory past 18 GB before I changed the allocation behavior. After terminating the worker/process, it dropped dramatically. It made me realize how different "the model fits in memory" is from "the whole inference pipeline behaves well in memory."