The world's first commercially viable end-to-end 1-bit LLMs
Caltech origins ยท Apache 2.0 ยท April 2026
01 โ The problem
10Bโ700B params
GBs of VRAM required
Phones, robots
Vehicles, secure systems
1-bit quantization ยท ~1 GB footprint ยท 10ร intelligence density
02 โ The family
03 โ Architecture
Every layer is strictly 1-bit โ no mixed-precision rescue weights. Grouped scaling (g128) per layer.
04 โ The metric
โlog(avg error rate) / model size (GB)
10ร+ improvement
05 โ Speed
4โ5ร more energy efficient than 16-bit equivalents
06 โ Benchmarks
Bonsai 8B matches or approaches leading 8B-class models across standard evaluation suites:
07 โ What becomes possible
Persistent agents running hours without cloud calls
Sub-100ms latency for physical-world decisions
Full capability in remote or bandwidth-constrained environments
Compliant assistants that never send data off-device
08 โ Roadmap
8B, 4B, 1.7B models open-sourced under Apache 2.0
Sub-1-bit regimes, hybrid architectures, multimodal and agent-native capabilities
Cloud-scale efficiency gains, open-source ecosystem expansion
09 โ Considerations
Noted limitations: Edge cases in long-context reasoning need further evaluation. Hardware-optimized 1-bit kernels still maturing across full device landscape. Standard benchmarks may not fully reflect real-world application performance โ domain-specific testing recommended.