Open weights
Models
Three ternary models, from four billion parameters down to sixty megabytes. All of them run on the CPU in front of you.
01 · Size on disk · was 8.90 GB
4B · 2.66 GB · ~1.94 bits/weight
2.66 GB
Cloe 1.2
Qwen3-4B, ternarized after training. From 8.9 GB to 2.66 GB, still answering.
- Post-training ternarization, no retraining
- 8.90 GB → 2.66 GB packed
- ~1.94 bits per weight, measured
- Rotation, ternarization, error compensation
02 · Size on disk · was 3,037 MB
752M · 635 MB · 1.58 bits/weight
635 MB
Cloe 1.1
Qwen3.5-0.8B converted to three states: 4.8 times smaller on disk, faster to decode.
- Ternary conversion with quantization-aware training
- 3,037 MB → 635 MB on disk
- 1.5821 bits per weight, measured
- +24% decode throughput vs the latent baseline
03 · Size on disk
70M · 60 MB
60 MB
S1.0
Seventy million parameters, trained from scratch in three states. Sixty megabytes.
- Ternary from the first step
- 60 MB on disk
- Speech models at 5 and 11 MB alongside it
- Report to follow
Scaling trajectory
Models converted after training keep most of what they knew. A model trained small from the first step keeps all of it, so every byte it sheds is pure gain. That is the line we are on, and it does not stop at ten.
↑ Intelligence per byte, relative to the original
→ Times smaller than the original
- The original model · 1×
- Cloe 1.2 · converted · 3.4× smaller · 87% kept
- Cloe 1.1 · converted · 4.8× smaller · 77% kept
- OneBit-T 3B · trained small · 4.3× smaller · 103% kept
- Projected · trained small · 10× smaller
Intelligence per byte = capability kept × size saved. Dashed parts are extrapolated.