| ▲ | acd 2 hours ago | |
I have written a llm compression quantization library called glq because of the high ram prices. Glq uses qtip Trellis quantization which are efficient at low bpw 2-4 bits. As a gamer and ai developer I want to squeeze more out of the same hardware. Glq runs with vLLM. Open source https://github.com/cnygaard/glq Pypi glq https://pypi.org/project/glq/ | ||