| ▲ | smallerize 2 hours ago | |
If you're dropping thousands on API tokens, you're going to be slowed down at least 10x trying to do everything on a single MBP. | ||
| ▲ | dannyw an hour ago | parent [-] | |
But you could grab a 5090, and paired with some DRAM for MoE offloading of bigger models, and be a happy camper with 1.8TB/s of memory bandwidth. Or just use Luna honestly. Worth considering if you’re ok with hosted APIs. | ||