| ▲ | New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode | |
| 4 points by medicis123 9 hours ago | ||
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website. WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive. First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers. MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST DeepSeek V4 Flash C1 1,518.91 21.15 21.15 DeepSeek V4 Flash C4 1,533.15 55.99 14.00 Gemma 4 26B A4B C1 4,579.73 30.22 30.22 Gemma 4 26B A4B C4 4,702.16 63.75 15.94 Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42 Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71 Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait DeepSeek V4 Flash 4,154.34 49.30 16s Gemma 4 26B A4B 18 4,781.44 64.67 6s Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/ | ||