Article (Scientific journals)
Differentiable Selection of Bit-Width and Numeric Format for FPGA-Efficient Deep Networks
Dellel, Kawthar; Trabes, Emanuel; Zayed, Aymen et al.
2025In Electronics, 14 (18), p. 3715
Peer reviewed
 

Files


Full Text
electronics-14-03715-1.pdf
Author postprint (923.88 kB)
Download

All documents in ORBi UMONS are protected by a user license.

Send to



Details



Keywords :
learnable bit-widths; numeric format optimization; quantization-aware training; Bit-Width; Design architecture; Extensive design; Fixed points; Floating points; Learnable bit-width; Numeric format optimization; Optimisations; Quantisation; Quantization-aware training; Control and Systems Engineering; Signal Processing; Hardware and Architecture; Computer Networks and Communications; Electrical and Electronic Engineering
Abstract :
[en] Quantization-aware training (QAT) has emerged as a key strategy for enabling efficient deep learning inference on resource-constrained platforms. Yet, most existing approaches rely on static, manually selected numeric formats—fixed-point or floating-point—and fixed bit-widths, limiting their adaptability and often requiring extensive design effort or architecture search. In this work, we introduce a novel QAT framework that breaks this rigidity by jointly learning, during training, both the numeric representation format and the associated bit-widths in an end-to-end differentiable manner. At the core of our method lies a unified parameterization that is capable of emulating both fixed- and floating-point arithmetic, paired with a bit-aware loss function that penalizes excessive precision in a hardware-aligned fashion. We demonstrate that our approach achieves state-of-the-art trade-offs between accuracy and compression on MNIST, CIFAR-10, and CIFAR-100, reducing average bit-widths to as low as 1.4 with minimal accuracy loss. Furthermore, FPGA implementation using Xilinx FINN confirms over (Formula presented.) LUT and (Formula presented.) BRAM savings. This is the first QAT method to unify numeric format learning with differentiable precision control, enabling highly deployable, precision-adaptive deep neural networks.
Disciplines :
Electrical & electronics engineering
Author, co-author :
Dellel, Kawthar ;  Université de Mons - UMONS > Faculté Polytechnique > Service d'Electronique et Microélectronique ; Laboratory of Electronics and Microelectronics, University of Monastir, Monastir, Tunisia
Trabes, Emanuel  ;  Université de Mons - UMONS > Faculté Polytechnique > Service d'Electronique et Microélectronique ; Departamento de Electronica, Universidad Nacional de San Luis, San Luis, Argentina
Zayed, Aymen ;  Université de Mons - UMONS > Faculté Polytechnique > Service d'Electronique et Microélectronique ; National Engineering School of Sousse, University of Sousse, Sousse, Tunisia
Faiedh, Hassene;  Higher Institute of Applied Science and Technology, University of Sousse, Sousse, Tunisia
Valderrama, Carlos Alberto  ;  Université de Mons - UMONS > Faculté Polytechnique > Service d'Electronique et Microélectronique
Language :
English
Title :
Differentiable Selection of Bit-Width and Numeric Format for FPGA-Efficient Deep Networks
Publication date :
September 2025
Journal title :
Electronics
ISSN :
2079-9292
eISSN :
2079-9292
Publisher :
Multidisciplinary Digital Publishing Institute (MDPI)
Volume :
14
Issue :
18
Pages :
3715
Peer reviewed :
Peer reviewed
Research unit :
Electronics and Microelectronics
Research institute :
Numediart
Funders :
European Union’s Horizon 2020 research and innovation programme
Funding text :
This project has received funding from the European Union\u2019s Horizon 2020 research and innovation programme under the Marie Sklodowska Curie grant agreement No 101034383.
Available on ORBi UMONS :
since 23 March 2026

Statistics


Number of views
48 (4 by UMONS)
Number of downloads
158 (1 by UMONS)

Scopus citations®
 
0
Scopus citations®
without self-citations
0
OpenCitations
 
0
OpenAlex citations
 
0

Bibliography


Similar publications



Contact ORBi UMONS