| ▲ | anerli an hour ago | |
Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions. However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory. | ||