Ko Lab

Research

Hamish designs the arithmetic that deep learning runs on: multipliers and compute units that trade a little exactness for large savings in energy and area, and hardware that makes privacy-preserving training practical.

Publications

Decoder reduction approximation scheme for Booth multipliers

M. H. Haider, H. Zhang, S.-B. Ko. IEEE Transactions on Computers, 73(3), 735-746, 2024.

  • Approximate Booth multipliers had fallen behind truncation-based approximate logarithmic multipliers; this scheme closes the gap
  • Uses only N/4 Booth decoders instead of the traditional N/2, at negligible error rates
  • The 16-bit BD16.4 design cuts normalised mean error deviation by 96.5% and power-area product by 69.6% against a state-of-the-art approximate logarithmic multiplier
Block diagram of the proposed BD16.4 approximate Booth multiplier using four decoders.

Booth encoding-based energy-efficient multipliers for deep learning

M. H. Haider, S.-B. Ko. IEEE Transactions on Circuits and Systems II: Express Briefs, 70(6), 2241, 2023.

  • A re-encoding scheme that combines Booth encoding with power-of-two quantization to shrink network weights
  • Model size down 30.77% for a CNN and 49.86% for a linear network, with minimal accuracy loss
  • Inference energy down 50.6% for the CNN and 90.1% for the linear network
Datapath of the power-of-two multiplier showing encoders, decoders and accumulator.

Memory-efficient differential privacy accelerator

M. H. Haider, N. Kim, H. Zhang, J. Arias-Garcia, H. J. Lee, S.-B. Ko. IEEE Asia Pacific Conference on Circuits and Systems (APCCAS), 2025.

  • Generates the Gaussian noise in line rather than moving it through memory, which is where the overhead usually goes
  • Controlled clock phase mismatches induce metastability in flip-flop arrays, and that entropy becomes the noise source
  • Paired with an approximate computation unit, implemented on FPGA and modelled as a PyTorch extension
Accelerator diagram with a grid of compute units, the proposed noise generator and buffers.

Power-efficient and reconfigurable compute unit for multi-precision AI inference

M. H. Haider, H. Zhang, S.-B. Ko. IEEE International Symposium on Circuits and Systems (ISCAS), 2026.

  • Fitting several fixed-precision multipliers and using one at a time is wasteful: the largest dominates the critical path and caps the clock
  • R4RC16 and R4RC32 instead reconfigure at run time between a low-power 8-bit mode and a default 16- or 32-bit mode
  • In low-power mode, up to 7.6 times the energy efficiency of state-of-the-art approximate logarithmic multipliers and 13.8 times that of approximate Booth designs
Gate-level diagram of the 4-2 compressor used in the reconfigurable multiplier.

Reconfigurable multi-precision multipliers for CNN acceleration

M. H. Haider, H. J. Lee, S.-B. Ko. International Journal of Contents, 21(4), 2025.

  • Targets edge computing, IoT devices and mobile platforms, where energy efficiency and throughput both matter
Block diagram of the 32-bit reconfigurable Booth multiplier split into two 8x32 units.

FFT-based deep learning for combustion instability prediction

S. Valarezo-Plaza, A. Erazo, M. H. Haider, J. Bae, P. Canteenwalla, S. Yun, S.-B. Ko. International Journal of Hydrogen Energy, 2026.

  • Two LSTM models compared: one on time-series pressure and heat release rate, one on frequency-domain features from an FFT
  • Measured across power levels of 15-30 kW, hydrogen content from 0% to 80%, air flow of 400-600 slpm and three downstream blockage ratios
Schematic of the combustor rig with fuel injector, V-gutter and the three downstream blockage ratios.

An introduction to AI for clinicians

S. B. Lee, A. B. Carter, M. H. Haider, S.-B. Ko. Interactive Journal of Medical Research, 2026.

  • A tutorial written for practising clinicians on what AI is already doing in medicine and what is coming
Diagram contrasting a labelled data table with unsupervised grouping of circles and squares.