Remix.run Logo
▲ Sub-1-Bit LLM Compression via Latent Factorization(github.com)
61 points by brainless 5 hours ago | 8 comments
▲augment_me 2 hours ago | parent | next [-]

Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.

I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

▲big-chungus4 2 hours ago | parent | prev | next [-]

Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory

▲tcdent an hour ago | parent | next [-]

I don't think it's trying to be a useful implementation, but the significance they do provide is that they are able to improve on the relative loss at lower quants.

So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.

Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).

▲GaggiX 2 hours ago | parent | prev [-]

I recently found this 1.58-bit model for ASR and it's surprising good (and very fast), that being said it's not a LLM.

https://huggingface.co/moondream/parakeet-redux

▲badatnames 2 hours ago | parent | prev | next [-]

Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means

▲nbutton762 2 hours ago | parent | prev | next [-]

Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!

▲bArray an hour ago | parent | prev | next [-]

Has anybody tested this? Are there any available computed models to test?

▲nico 2 hours ago | parent | prev [-]

Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?