Adinkra Labs / Research / Bonsai

PrismML Bonsai

The world's first commercially viable end-to-end 1-bit LLMs

Caltech origins ยท Apache 2.0 ยท April 2026

Intelligence trapped in the cloud

Cloud models

10Bโ€“700B params
GBs of VRAM required

โœ• Gap

Edge devices

Phones, robots
Vehicles, secure systems

Bonsai bridges the gap

1-bit quantization ยท ~1 GB footprint ยท 10ร— intelligence density

Three models, one architecture

Bonsai 1.7B
0.24
GB
Ultra-light edge
Bonsai 4B
0.57
GB
Balanced performance

True 1-bit end-to-end

Every layer is strictly 1-bit โ€” no mixed-precision rescue weights. Grouped scaling (g128) per layer.

Input tokens
Embeddingsโœ“ 1-bit
Attentionโœ“ 1-bit
MLPโœ“ 1-bit
LM Headโœ“ 1-bit
Output tokens

Intelligence density

โˆ’log(avg error rate) / model size (GB)

1.06
Bonsai 8B
per GB
0.10
Full-precision 8B
per GB

10ร—+ improvement

Inference performance

Hardware
Runtime
Speed
RTX 4090
llama.cpp
440 t/s
Apple M4 Pro
MLX
136 t/s
iPhone 17 Pro
MLX
~44 t/s

4โ€“5ร— more energy efficient than 16-bit equivalents

Competitive at a fraction of the size

Bonsai 8B matches or approaches leading 8B-class models across standard evaluation suites:

IFEval
Instruction following
GSM8K
Math reasoning
HumanEval+
Code generation
BFCL
Function calling
MuSR
Multi-step reasoning
MMLU-Redux
Broad knowledge

New product categories

๐Ÿค–

On-device agents

Persistent agents running hours without cloud calls

โšก

Real-time robotics

Sub-100ms latency for physical-world decisions

๐ŸŒ

Offline intelligence

Full capability in remote or bandwidth-constrained environments

๐Ÿ”’

Private copilots

Compliant assistants that never send data off-device

What's next

Now โ€” Bonsai family launched

8B, 4B, 1.7B models open-sourced under Apache 2.0

Near-term

Sub-1-bit regimes, hybrid architectures, multimodal and agent-native capabilities

Medium-term

Cloud-scale efficiency gains, open-source ecosystem expansion

Noted limitations: Edge cases in long-context reasoning need further evaluation. Hardware-optimized 1-bit kernels still maturing across full device landscape. Standard benchmarks may not fully reflect real-world application performance โ€” domain-specific testing recommended.