Remix.run Logo
dbbk 3 hours ago

Well if you're spending thousands on API tokens already, you could just drop the same amount on a 128GB MacBook Pro and that's a one time cost.

smallerize 2 hours ago | parent | next [-]

If you're dropping thousands on API tokens, you're going to be slowed down at least 10x trying to do everything on a single MBP.

dannyw an hour ago | parent [-]

But you could grab a 5090, and paired with some DRAM for MoE offloading of bigger models, and be a happy camper with 1.8TB/s of memory bandwidth.

Or just use Luna honestly. Worth considering if you’re ok with hosted APIs.

Gigachad an hour ago | parent | prev | next [-]

The models people are spending thousands on require more on the range of 600-800gb memory.

128gb hardly runs deepseek v4 flash which is almost free via api pricing.

neuroticnews25 2 hours ago | parent | prev [-]

Don't forget about energy usage, you'll probably never break even vs same model on openrouter.

jurgenburgen 2 hours ago | parent [-]

If you can’t do it cheaper on your own hardware it does make you wonder how much of the cost of inference those large LLM providers are eating? Datacenter hardware isn’t magic.

flaunf221 an hour ago | parent | next [-]

Your personal hardware probably isn't running useful tasks 24/7. If you spend 60% of your 8h work day on full on agentic work, then your hardware is paying off for itself only 20% of available time.

petu 2 hours ago | parent | prev | next [-]

Datacenter hardware can batch at large scale, probably over 90% more energy efficient per token than a MacBook.

Der_Einzige 29 minutes ago | parent | prev [-]

Datacenter hardware might as well be magic compared to consumer. "Oh the F35 isn't magic compared to my M16 bro!"