| ▲ | flutetornado an hour ago | |||||||
GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation. Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens. Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth. | ||||||||
| ▲ | cmrdporcupine an hour ago | parent [-] | |||||||
What are you getting for prefill? | ||||||||
| ||||||||