Tech News
← Home  ·  All topics

Neurips 2025

1 GoKawiil brief on this topic

Researchers release LittleBit code for compressing LLMs below 1 bit per weight

A research team has published the implementation of LittleBit, a method accepted at NeurIPS 2025, which compresses large language model weights into a sub-1-bit range, down to 0.1 bits per weight, by factorizing weight matrices into binarized low-rank latent factors with learned scaling. A follow-up method, LittleBit-2 (accepted at ICML 2026), adds an optional initialization technique called Joint Iterative Quantization that aligns latent factors geometrically before training, without changing the model's inference-time structure. The codebase supports models including OPT, Llama, Phi-4, Qwen2.5, QwQ, Gemma 2/3, and Qwen3.