| ▲ | coder543 5 hours ago | |
As you allude, the prompt processing speeds are a killer improvement of the Spark which even 2 Strix Halo boxes would not match. Prompt processing is literally 3x to 4x higher on GPT-OSS-120B once you are a little bit into your context window, and it is similarly much faster for image generation or any other AI task. Plus the Nvidia ecosystem, as others have mentioned. One discussion with benchmarks: https://www.reddit.com/r/LocalLLaMA/comments/1oonomc/comment... If all you care about is token generation with a tiny context window, then they are very close, but that’s basically the only time. I studied this problem extensively before deciding what to buy, and I wish Strix Halo had been the better option. | ||