JetSpec Enables Up to 9.64x Lossless LLM Inference Speedup with Up to 1000TPS

	▲	JetSpec Enables Up to 9.64x Lossless LLM Inference Speedup with Up to 1000TPS(haoailab.com)
		4 points by snyhlxde 11 hours ago \| 1 comments

	▲	snyhlxde 11 hours ago \| parent [-]
		We find speculative decoding can push LLM generation latency to extreme by co-optimizing drafting cost and drafting quality with causal parallel tree drafting. JetSpec reaches up to 9.64× end-to-end speedup on MATH-500 and 4.58× on open-ended chat while keeping lossless. With CUDA graph and kernel optimizations, JetSpec further translates to around 1000 TPS on a single B200 GPU.