Sub-1-Bit LLM Compression via Latent Factorization

(github.com)

39 points | by brainless 3 hours ago

3 comments

  • big-chungus4 15 minutes ago
    Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory
  • badatnames 19 minutes ago
    Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means
  • nico 17 minutes ago
    Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?