Zhaohui Xu

PhD Student · ShanghaiTech University · AI Infrastructure

Portrait of Zhaohui Xu

I am Zhaohui Xu (徐兆辉), a PhD student in Computer Science and Technology at ShanghaiTech University, jointly supervised by Prof. Yajun Ha and Prof. Li Jiang.

My research focuses on deployable low-precision inference. I work on low-bit quantization, model compression, and efficient deployment for AI inference systems, with an emphasis on kernel-runtime co-optimization and heterogeneous software-hardware co-design. My published work includes CCF-A papers at ASPLOS, HPCA, and DAC.

Research Focus

Low-bit quantization and model compression · Efficient inference and deployment · Kernel-runtime and heterogeneous-system co-optimization

Education

  • 2025—2028 | ShanghaiTech University | PhD in Computer Science and Technology
    Joint training with the Qizhi Institute; Prof. Yajun Ha (ShanghaiTech University); Prof. Li Jiang (Shanghai Jiao Tong University)
  • 2022—2025 | ShanghaiTech University | Master in Computer Science and Technology
    Joint training with the Institute of Computing Technology, Chinese Academy of Sciences; Prof. Yajun Ha (ShanghaiTech University); Researcher Ying Wang (Institute of Computing Technology, Chinese Academy of Sciences)
  • 2018—2022 | China Three Gorges University | Bachelor in Computer Science and Technology

Selected Publications

COMET: Towards Practical W4A4KV4 LLMs Serving

Authors: Lian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
ASPLOS 2025 · CCF A · DOI

Research role: Research and implementation of the FMPQ fine-grained mixed-precision quantization method.

Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM

Authors: Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
HPCA 2025 · CCF A · DOI

Research role: Efficient hot/cold-neuron prediction using MLP pruning.

Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network Acceleration

Authors: Lian Liu, Zhaohui Xu, Yintao He, Ying Wang, Huawei Li, Xiaowei Li, Yinhe Han
DAC 2024 · CCF A · DOI

Research role: Hardware-friendly dynamic mixed-precision quantization and the design and implementation of the dynamic quantization accelerator.