Remix.run Logo
acd 2 hours ago

I have written a llm compression quantization library called glq because of the high ram prices. Glq uses qtip Trellis quantization which are efficient at low bpw 2-4 bits. As a gamer and ai developer I want to squeeze more out of the same hardware. Glq runs with vLLM.

Open source https://github.com/cnygaard/glq

Pypi glq https://pypi.org/project/glq/