Remix.run Logo
walrus01 4 hours ago

3.8-flash-next quantized in a "large" Q4 that just fits in 128GB RAM even more so, in how close it can get to state of the art in a number of benchmarks. Or a large Q8 version of it that fits in under 190GB. Competing against things that are closed weights/opaque information about the model and might very well be 600B+ in size.

nicman23 3 hours ago | parent [-]

it "fits" in 64 ram with mmap. granted it runs at 15 tk/s with a 9070xt but it runs

walrus01 3 hours ago | parent [-]

Right, I meant "fits" in the sense of I can load the whole thing into some combination of system RAM and GPU at llama-server launch.

15 tk/s isn't useless if you can give it big tasks to do overnight, or like ask it to do something and check back 3-4 hours later.