OneBit

What we publish

Research

Ternary models, learned quantization boundaries, and what it takes to keep reasoning intact at 1.58 bits.

Papers

01 · CLOE V1.3

Scaling Post-Training Ternarization to Qwen3-8B

The same pipeline scaled from Qwen3-4B to Qwen3-8B. The larger model keeps more of what it knew, 78.5% against 69.6%, so size itself makes a model more robust to aggressive conversion.

Technical research report · September 2026 · Malik, Devan, Mehra
Read
02 · CLOE V2.0

Post-Training Ternarization of Qwen3-4B

Qwen3-4B pushed to 1.641 bits per weight with no retraining. Compression is demonstrated; faster execution is still an open problem.

Technical research report · August 2026 · Malik, Devan, Mehra
Read
03 · CLOE V1.2

Post-Training Ternarization of Qwen3 Language Models

Rotation, ternarization and error compensation on Qwen3-4B, end to end: capability, effective bits per weight, storage and inference, measured.

Technical report · August 2026 · Malik, Devan, Mehra
Read
04 · CLOE V1.1

Capability-Stratified Degradation in Ternary Language Models

What survives when a pretrained 752M model is pushed to three states: not a uniformly weaker model, but a stratified one.

Technical paper · August 2026 · Malik, Devan, Mehra
Read
05 · CLOE V1.0

Cloe: Hybrid Surgical Distillation for 1.58-Bit Edge Quantization of Large Language Models

Attention stays at 16 bits, the MLPs go ternary, and a teacher restores what the cut removed.

Paper · 2026 · Malik, Vishnuprasad, Devan, Mehra
Read
06 · ITBO

A Dynamic Ternary Architecture for Latency-Critical Edge Agents

Learned quantization boundaries let a ternary network decide, per layer, where a weight becomes zero.

Paper · February 2026 · Malik, Mehra, Vishnuprasad, Pundir, Tyagi, Arutkeerthi
Read
07 · Overview

Edge AI using Ultra Low Bit LLMs

Why the next generation of models will run on the chips that already exist, and what that changes.

White paper · January 2026 · Malik, Mehra, Vishnuprasad, Pundir, Tyagi
Read

The idea

From 65,536 shades to three

Shades per cell

65,536

Space it takes

100%

A language model is a web of billions of connections. What it knows is the shape of that web: which connections exist, which way each one pushes, and roughly how hard. Today every connection is stored as a high-precision number, and the precision is the expensive part. It is why the model fills a data center, why it needs a GPU, and why you reach it through a wire.

The precision is almost entirely unused. The intelligence is in the pattern, not in the decimals. A connection needs to say push, pull or stay out of it, and very little more. Written that way, the same web takes about a tenth of the space, and the arithmetic collapses to addition, which every processor made in the last decade does quickly.

That is what OneBit builds: models trained from the first step to live in that form, rather than large models squeezed into it afterwards. The intelligence is the same. The places it can live are not.

What OneBit does to a model

Original, FP16OneBit, ternary
Weights on disk8.90 GB2.66 GB
Bits per weight16~1.94
Capability kept100%87%, target 100%
HardwareGPUany CPU
Inferencein a data centeron the device
Cost per tokenmeterednone
Datacloudlocal

Qwen3-4B in FP16 against Cloe 1.2, its ternary conversion. Numbers from the paper.