
Muhammad Fahad
Electrical and Computer EngineeringMaster's research student in the Ko Lab, supervised by Prof. Seok-Bum Ko. He received his BS in Electrical Engineering from the Pakistan Institute of Engineering and Applied Sciences, and has over a year of experience as an FPGA design engineer at Renzym, Islamabad. His research interests include computer architecture, deep learning accelerators and FPGA/ASIC implementation of compute-intensive applications.
Research
Muhammad's work centres on efficient digital hardware for artificial intelligence: computer arithmetic, mixed- and variable-precision computing, approximate computing, FPGA and ASIC accelerator design, and hardware-software co-design. The recurring theme is reducing the cost of computation without giving up the accuracy the target application actually needs.
Current project
M.Sc. research, Ko Lab, supervised by Prof. Seok-Bum Ko.
A reconfigurable multi-precision posit multiplier with overflow-aware variable-precision approximation. A runtime-reconfigurable posit multiplier for energy-efficient AI hardware. A shared radix-4 Booth datapath performs four Posit(8,0), two Posit(16,1) or one Posit(32,2) multiplication at runtime, while Overflow-Aware Precision Control (OAPC) uses mantissa information to prevent a residual precision loss that can occur after normalization. The design is studied from arithmetic, application and hardware angles: the RTL is verified in combinational and pipelined form, evaluated on mixed-precision CIFAR-10 inference, and synthesized in a TSMC 65 nm flow for power, performance and area.
- One shared multiplier covers all three posit widths at runtime, avoiding a separate arithmetic datapath for each precision mode
- OAPC raises the bit-exact result rate by about 7 percentage points, eliminates every observed error of 3 ulps or more, and halves the maximum observed error from 4 ulps to 2
- That accuracy costs 35.1 µm² of mean additional area, about 0.13% of the multiplier or roughly 24 NAND2 gates, with no delay overhead and no reduction in the 531.9 MHz maximum frequency
- At maximum frequency the design occupies 28,396.1 µm², draws 7.139 mW, and delivers 2.13 billion Posit(8,0) products per second at 3.355 pJ per operation and 74.9 GOPS/mm²
- Where full-width exact rounding can be traded away, a compile-time TRUNC_LSB of 12 cuts area a further 16.7% and power 12.2% while keeping runtime selection across all three widths
Publications
FPGA based artificial neural network accelerator
E. Fatima, M. Fahad, H. Abrar, H. Haroon-ur-Rashid, H. Waris. 2024 26th International Multi-Topic Conference (INMIC), 1-6, 2024.
- Implements a LeNet-5 convolutional network as an FPGA accelerator using Vivado HLS on a Xilinx ZCU102, comparing floating-point against fixed-point formats and applying hardware-oriented optimizations to cut latency and resource use
- Pipelining with array partitioning reduced accelerator delay by 51.4% to 12.9 ms, for only about a 1% increase in BRAM and flip-flop use
- Loop unrolling cut a further 21.56% to 9.57 ms, and a second unrolling step reached 9.04 ms, another 5.54% faster
- A 16-bit fixed-point implementation held 97.78% top-1 accuracy over 10,000 MNIST test images, only 0.94 percentage points below the 98.72% FP32 reference
- Taken through the full deployment flow with C/RTL co-simulation, IP generation, AXI-based communication and on-board testing, measuring 10.2 ms per image; the gap to the 9.04 ms synthesized figure was attributed mainly to AXI-Lite interface overhead