| ▲ | throwa356262 3 hours ago | |
The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd. In practice, you will be able to run models a bit bigger than 35B. | ||
| ▲ | cyanydeez 2 hours ago | parent [-] | |
https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried. 38GB of vram resident. more tk/s, more prefill. | ||