Ko Lab

Research

Sayed works on making large models cheap enough to train, fine-tune and run on constrained hardware, and on the high-throughput error-correcting decoders used in satellite communication.

Publications

LYRA: low-frequency rank adaptation via factored DCT coefficients

S. Muhsin, S.-B. Ko. IEEE Signal Processing Letters, vol. 33, 2026.

  • Parameterises each weight update with a small set of low-frequency 2D discrete cosine transform coefficients, chosen separately along each axis
  • That separable structure allows a factored forward pass of three small matrix multiplications, avoiding the dense inverse transform FourierFT needs at every step
  • On GLUE and SuperGLUE with BERT-base and RoBERTa-base it matches FourierFT at identical parameter and optimizer-memory budgets, with the lowest peak GPU memory of every method tested
Grid of low-frequency discrete cosine transform coefficient maps used to parameterise the weight update.

Optimized transformer models: pruning and quantization for the edge

M. H. Haider, S. Valarezo-Plaza, S. Muhsin, H. Zhang, S.-B. Ko. IEEE International Symposium on Circuits and Systems (ISCAS), 2024.

Diagram comparing a twelve-layer BERT model with a compressed six-layer version.
  • Pruning and quantization were designed for CNNs and do not transfer to transformers automatically, because the computation patterns differ
  • A comparative analysis of what does carry over, aimed at real-world edge deployment
  • Large gains in compression ratio while transformer accuracy holds
Flowchart of the compression pipeline from BERT base through knowledge distillation to pruning and quantization.

A high-throughput QC-LDPC decoder for near-earth satellite communication

K. M. S. Nampoothiri, K. N. Menon, S. Muhsin, S. Manoj, S. P, D. K. J, D. P. P. National Institute of Technology Calicut.

  • Decoder architecture for the (8176,7154) quasi-cyclic LDPC code recommended by the CCSDS for near-earth applications
  • Shift-register memory circuits and a pipe-stage forwarding mechanism avoid memory conflict, which allows the core processing unit to be heavily pipelined
  • 2.65 Gbps at ten iterations on a Xilinx XCVU9P FPGA, clocked at 253 MHz
Block diagram of the LDPC decoder showing input interface, bit node storage, check node units and output interface.