Remix.run Logo
flutetornado an hour ago

GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation.

Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens.

Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.

cmrdporcupine an hour ago | parent [-]

What are you getting for prefill?

flutetornado 21 minutes ago | parent [-]

450 with PTQ_01 and 900 with the other PQ2_0.