| ▲ | girvo 3 hours ago |
| Not quite: not all of this needs to be in VRAM It has a set of n-gram tables which you can stream from system RAM or even NVMe That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two? |
|
| ▲ | jonsoft 2 hours ago | parent | next [-] |
| It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache): https://github.com/christopherowen/spark-ds41f |
| |
| ▲ | girvo 2 hours ago | parent [-] | | Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand. I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware | | |
|
|
| ▲ | rsolva 3 hours ago | parent | prev [-] |
| I have access to two and will explore this the coming weeks. |
| |
| ▲ | girvo 2 hours ago | parent [-] | | Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun |
|