Remix.run Logo
cyanydeez 14 hours ago

mmm, the chinese models are also working on local GPUs at consumer grades. so theyre not just drainig cloud moats.

vrm 14 hours ago | parent [-]

good luck running a 2.4T model on any local hardware. it’s not gonna happen. the arrow is to specialized hardware at least for the smartest models

Flere-Imsaho 7 hours ago | parent | next [-]

Yes but someone who has access to that kind of hardware can distill down to a smaller model that is specialised for a specific task. I don't need my local model to be an oracle for everything, I want a coding AI, one that knows medicine, another that recognises objects in my security camera, etc.

matheusmoreira 13 hours ago | parent | prev | next [-]

I have hope it'll happen one day, even if not now.

nekusar 13 hours ago | parent [-]

Already is possible. On a machine with 32GB ram, and NO gpu. Just need a large SSD or NVME. Streams from disk to memory.

https://github.com/JustVugg/colibri

khuey 10 hours ago | parent [-]

Painfully slow tok/s though.

NekkoDroid 3 hours ago | parent | next [-]

Isn't it more s/tok?

nekusar 2 hours ago | parent | prev [-]

I didn't say it was fast!

But its also a 800B sized model running on a ram constrained system with no GPU.

Techniques, GPUs, more ram, and faster disks can always speed it up. But the point being is they run on low end machines now. Its now an optimization problem, not a possibility assessment.

cyanydeez 3 hours ago | parent | prev [-]

sir, I'm not running a multi billion dollar code base; I just want my nose wiped and a clean fork of whatever repo might be the target of supply chain attacks, and a few nicissities.

I don't need 2.4T to do that; I'm doing it with 35B or 27B. If they get me a model in ~80B with a A5B or A7B, that will be the end point.

It's bizarre people, by themselves, believe all these parameters are getting them much more.

Lets be serious: if we as a civilization really wanted the advancements promised, we'd find the 1000 best scientists and give them free access to these models while the rest of us get personal GPUs for specific use cases.

But instead, we have to endeour this penis measuring contest for the infinite bikeshedding of the universe.