Precision as a design choice

QAT: shrinking the weight footprint

Quantization-aware training lets the model practise at lower precision instead of discovering the damage only after compression.

16 bitBF16 weight block
memory footprint 100%with QAT · steadier signal
16 bitlive model