| ▲ | anoncow 5 hours ago | |
3000 tokens per sec on 32 mb Ram? | ||
| ▲ | fc417fc802 5 hours ago | parent [-] | |
fast != practical You can get lots of tokens per second on the CPU if the entire network fits in L1 cache. Unfortunately the sub 64 kiB model segment isn't looking so hot. But actually ... 3000? Did GP misplace one or two zeros there? | ||