OneBit

Open weights

Models

Three ternary models, from four billion parameters down to sixty megabytes. All of them run on the CPU in front of you.

01 · Size on disk · was 8.90 GB

4B · 2.66 GB · ~1.94 bits/weight

2.66 GB

Cloe 1.2

Qwen3-4B, ternarized after training. From 8.9 GB to 2.66 GB, still answering.

  • Post-training ternarization, no retraining
  • 8.90 GB → 2.66 GB packed
  • ~1.94 bits per weight, measured
  • Rotation, ternarization, error compensation

02 · Size on disk · was 3,037 MB

752M · 635 MB · 1.58 bits/weight

635 MB

Cloe 1.1

Qwen3.5-0.8B converted to three states: 4.8 times smaller on disk, faster to decode.

  • Ternary conversion with quantization-aware training
  • 3,037 MB → 635 MB on disk
  • 1.5821 bits per weight, measured
  • +24% decode throughput vs the latent baseline

03 · Size on disk

70M · 60 MB

60 MB

S1.0

Seventy million parameters, trained from scratch in three states. Sixty megabytes.

  • Ternary from the first step
  • 60 MB on disk
  • Speech models at 5 and 11 MB alongside it
  • Report to follow

Scaling trajectory

Models converted after training keep most of what they knew. A model trained small from the first step keeps all of it, so every byte it sheds is pure gain. That is the line we are on, and it does not stop at ten.

↑ Intelligence per byte, relative to the original

1×5×10×Converted after trainingTrained small from the startThe original modelCloe 1.2Cloe 1.1OneBit-T 3BProjected
1×10×

→ Times smaller than the original

  • The original model · 1×
  • Cloe 1.2 · converted · 3.4× smaller · 87% kept
  • Cloe 1.1 · converted · 4.8× smaller · 77% kept
  • OneBit-T 3B · trained small · 4.3× smaller · 103% kept
  • Projected · trained small · 10× smaller

Intelligence per byte = capability kept × size saved. Dashed parts are extrapolated.