Resume
AI Infrastructure | Low-Precision Inference, Model Compression, and Efficient Deployment
Research Focus
I work on low-bit quantization, model compression, and efficient deployment for AI inference systems, with an emphasis on kernel-runtime co-optimization and heterogeneous software-hardware co-design.
Education
-
2025—2028 | ShanghaiTech University | PhD in Computer Science and Technology
Joint training with the Qizhi Institute; jointly supervised by Prof. Yajun Ha and Prof. Li Jiang. -
2022—2025 | ShanghaiTech University | Master in Computer Science and Technology
Joint training with the Institute of Computing Technology, Chinese Academy of Sciences; jointly supervised by Prof. Yajun Ha and Researcher Ying Wang. - 2018—2022 | China Three Gorges University | Bachelor in Computer Science and Technology
Selected Publications
COMET: Towards Practical W4A4KV4 LLMs Serving
Lian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
ASPLOS 2025 · CCF A
Research role: Research and implementation of the FMPQ fine-grained mixed-precision quantization method.
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, Ying Wang
HPCA 2025 · CCF A
Research role: Efficient hot/cold-neuron prediction using MLP pruning.
Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network Acceleration
Lian Liu, Zhaohui Xu, Yintao He, Ying Wang, Huawei Li, Xiaowei Li, Yinhe Han
DAC 2024 · CCF A
Research role: Hardware-friendly dynamic mixed-precision quantization and the design and implementation of the dynamic quantization accelerator.
Industry Experience
Baidu Kunlunxin | Deep Learning Framework Development Engineer Intern | 2024-10 - 2025-01
Technical focus: PyTorch compatibility engineering, custom-operator integration, and performance-tooling extension for domestic XPU accelerators.
Integrated and tested high-performance custom operators for PyTorch/XPU, adapted the NVTX backend for XPU, and implemented XSight System profiling visualization.
Huawei Technologies | AI Engineer Intern | 2024-06 - 2024-09
Technical focus: FlashAttention operator tuning, workload adaptation, and fused implementation for Ascend NPUs.
Tuned and adapted Ascend C FlashAttention and PromptFlashAttention fused operators, and evaluated accuracy and performance for speculative-inference workloads.
Awards & Recognition
- 2026 · CANN Ecosystem Contributor Certificate — Core Developer
- 2025 · Open Hardware 2025 Winner
- 2025 · CCF Sys2025 Graph Computing System Design Competition Grand Prize (team member)