Keywords :
learnable bit-widths; numeric format optimization; quantization-aware training; Bit-Width; Design architecture; Extensive design; Fixed points; Floating points; Learnable bit-width; Numeric format optimization; Optimisations; Quantisation; Quantization-aware training; Control and Systems Engineering; Signal Processing; Hardware and Architecture; Computer Networks and Communications; Electrical and Electronic Engineering
Abstract :
[en] Quantization-aware training (QAT) has emerged as a key strategy for enabling efficient deep learning inference on resource-constrained platforms. Yet, most existing approaches rely on static, manually selected numeric formats—fixed-point or floating-point—and fixed bit-widths, limiting their adaptability and often requiring extensive design effort or architecture search. In this work, we introduce a novel QAT framework that breaks this rigidity by jointly learning, during training, both the numeric representation format and the associated bit-widths in an end-to-end differentiable manner. At the core of our method lies a unified parameterization that is capable of emulating both fixed- and floating-point arithmetic, paired with a bit-aware loss function that penalizes excessive precision in a hardware-aligned fashion. We demonstrate that our approach achieves state-of-the-art trade-offs between accuracy and compression on MNIST, CIFAR-10, and CIFAR-100, reducing average bit-widths to as low as 1.4 with minimal accuracy loss. Furthermore, FPGA implementation using Xilinx FINN confirms over (Formula presented.) LUT and (Formula presented.) BRAM savings. This is the first QAT method to unify numeric format learning with differentiable precision control, enabling highly deployable, precision-adaptive deep neural networks.
Scopus citations®
without self-citations
0