Researchers release LittleBit code for compressing LLMs below 1 bit per weight
A research team has published the implementation of LittleBit, a method accepted at NeurIPS 2025, which compresses large language model weights into a sub-1-bit range, down to 0.1 bits per weight, by factorizing weight matrices into binarized low-rank latent factors with learned scaling. A follow-up method, LittleBit-2 (accepted at ICML 2026), adds an optional initialization technique called Joint Iterative Quantization that aligns latent factors geometrically before training, without changing the model's inference-time structure. The codebase supports models including OPT, Llama, Phi-4, Qwen2.5, QwQ, Gemma 2/3, and Qwen3.
GoKawiil's interpretation of the reporting above, not reported fact.
Pushing model compression below one bit per weight could let large language models run with drastically reduced memory and compute requirements, which may matter for deploying big models on constrained hardware. Because LittleBit-2's improvement is confined to initialization and adds no inference overhead, it suggests compression gains can be layered onto existing deployed architectures without redesigning them. Broad support across popular model families indicates the researchers are aiming for practical adoption rather than a narrow proof of concept.
- LittleBit compresses LLM weights to as low as 0.1 bits per weight via latent factorization and binarization.
- LittleBit-2 introduces an opt-in Joint-ITQ initialization that improves latent geometry alignment without added inference cost.
- The open-source codebase supports a wide range of models, including Llama, Qwen, Gemma, and Phi-4 families.
Source: github.com, 2026-10-08
Published there as: “Sub-1-Bit LLM Compression via Latent Factorization”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.