| ▲ | verdverm an hour ago |
| do you have a HF link? HF search is not uncovering it for me (or is it somewhere else) |
|
| ▲ | girvo an hour ago | parent | next [-] |
| https://github.com/spark-arena/eugr-recipes/blob/main/recipe... This one! I'd recommend pointing your agent at it (after installing sparkrun), and asking it to research the absolute latest in TP=1 Flash-Next - mine grabbed particular vLLM nightlies and mods to improve performance, and it was well worth it. |
| |
| ▲ | verdverm an hour ago | parent [-] | | I have a quirky vLLM on k8s on 2x OEM sparks setup with 9 models available to me. I'm not keen to run nightly vLLM, too many issues with it in the past. Going the qwen-next path means displacing things I use daily :/ I have a watchful eye on the diffusion ~ Jev/Kev PR https://github.com/vllm-project/vllm/pull/57250 | | |
| ▲ | girvo an hour ago | parent [-] | | For what it's worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don't think you'll need the nightly for that at all, just use the recipe) I'm so tempted to buy a second one... | | |
| ▲ | verdverm an hour ago | parent [-] | | prices have gone up quite a bit... I'm running embedding, reranking, and policy tuned models too, and a Jev/Kev when that's landed. Flash Next is not a substitute for those I have OpenCode/Fireworks to access big models |
|
|
|
|
| ▲ | verdverm an hour ago | parent | prev [-] |
| looks like this is likely it https://github.com/spark-arena/eugr-recipes https://github.com/eugr/spark-vllm-docker |