Resume

AI Infrastructure | Low-Precision Inference, Model Compression, and Efficient Deployment

Research Focus

I work on low-bit quantization, model compression, and efficient deployment for AI inference systems, with an emphasis on kernel-runtime co-optimization and heterogeneous software-hardware co-design.

Education

  • 2025—2028 | ShanghaiTech University | PhD in Computer Science and Technology
    Joint training with the Qizhi Institute; jointly supervised by Prof. Yajun Ha and Prof. Li Jiang.
  • 2022—2025 | ShanghaiTech University | Master in Computer Science and Technology
    Joint training with the Institute of Computing Technology, Chinese Academy of Sciences; jointly supervised by Prof. Yajun Ha and Researcher Ying Wang.
  • 2018—2022 | China Three Gorges University | Bachelor in Computer Science and Technology

Selected Publications

COMET: Towards Practical W4A4KV4 LLMs Serving

Lian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
ASPLOS 2025 · CCF A
Research role: Research and implementation of the FMPQ fine-grained mixed-precision quantization method.

Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM

Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
HPCA 2025 · CCF A
Research role: Efficient hot/cold-neuron prediction using MLP pruning.

Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network Acceleration

Lian Liu, Zhaohui Xu, Yintao He, Ying Wang, Huawei Li, Xiaowei Li, Yinhe Han
DAC 2024 · CCF A
Research role: Hardware-friendly dynamic mixed-precision quantization and the design and implementation of the dynamic quantization accelerator.

Industry Experience

Baidu Kunlunxin | Deep Learning Framework Development Engineer Intern | 2024-10 - 2025-01

Technical focus: PyTorch compatibility engineering, custom-operator integration, and performance-tooling extension for domestic XPU accelerators.
Integrated and tested high-performance custom operators for PyTorch/XPU, adapted the NVTX backend for XPU, and implemented XSight System profiling visualization.

Huawei Technologies | AI Engineer Intern | 2024-06 - 2024-09

Technical focus: FlashAttention operator tuning, workload adaptation, and fused implementation for Ascend NPUs.
Tuned and adapted Ascend C FlashAttention and PromptFlashAttention fused operators, and evaluated accuracy and performance for speculative-inference workloads.

Awards & Recognition

  • 2026 · CANN Ecosystem Contributor Certificate — Core Developer
  • 2025 · Open Hardware 2025 Winner
  • 2025 · CCF Sys2025 Graph Computing System Design Competition Grand Prize (team member)