| ▲ | verdverm 2 hours ago | |
The tool I've been using, llm-compressor, can quant models that do not fit in memory (use the sequential pipeline) https://github.com/vllm-project/llm-compressor my setup to help you on your way: https://github.com/verdverm/quantr Though it seems these will not be needed as Poolside has published quants & dflash with their models. | ||