| ▲ | yalok 7 hours ago | |
sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference. And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done... 0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | ||