AI Serving
Delta-MoE
Rethinks transfer granularity for CPU-GPU MoE serving to reduce data movement and improve efficient model execution.
Ph.D. Student, HKUST ECE
Computer Architecture · ML Systems · Hardware-Software Co-Design
I am a second-year Ph.D. student at The Hong Kong University of Science and Technology (HKUST), advised by Prof. Yuan Xie. My research focuses on efficient AI systems and computer architecture, especially CPU-GPU serving for LLM and MoE workloads, memory-centric acceleration, and hardware-software co-design for irregular AI computation.
Before joining HKUST, I received my M.S. and B.S. in Microelectronics from Fudan University.
I build efficient AI systems and infrastructure, spanning AI serving, memory-centric computing, and hardware-software co-design.
A few recent directions that connect my work across AI serving systems, memory-centric acceleration, and architecture design.
AI Serving
Rethinks transfer granularity for CPU-GPU MoE serving to reduce data movement and improve efficient model execution.
Compute-in-Memory
Explores ReRAM-coupled digital CIM with a fully-fused multi-head attention dataflow for FlashAttention-style workloads.
Edge AI Acceleration
Builds digital stochastic computing-in-memory support with accurate OR accumulation for energy-efficient edge AI models.
[Paper] SwiftCIM was accepted to ESSERC 2026.
[Paper] Delta-MoE was accepted to ICCAD 2026.
[Paper] DSCIM on Digital Stochastic Computing in Memory was accepted to DATE 2026.
[Paper] McPAL on Multi-Chiplet HBM-PIM architecture for LLMs was accepted to DAC 2025.
[Paper] DIRC-RAG on edge RAG acceleration was accepted to ISLPED 2025.
[Paper] FullSparse on sparse-aware GEMM acceleration was published at CF 2024.
Personal website launched.
Ph.D. in Electronic and Computer Engineering
The Hong Kong University of Science and Technology (HKUST), 2024 - Present
Advisor: Prof. Yuan Xie
M.S. in Microelectronics
Fudan University (复旦大学), 2021 - 2024
B.S. in Microelectronics
Fudan University (复旦大学), 2017 - 2021
*: Equal contributions
SwiftCIM: A 55nm 23.2µJ/Token L-0.5 ReRAM Coupled Digital CIM Accelerator with Fully-Fused Multi-Head Attention Dataflow for FlashAttention
ESSERC, 2026 Accepted
Delta-MoE: Rethinking Transfer Granularity for CPU-GPU MoE Serving
ICCAD, 2026 Accepted
Scalable Sparse Transformer Accelerator with In-Memory Butterfly Zero Skipper and Local Attention Reusable Engine for Irregular-Pruned NN
IEEE JETCAS, 2026 Accepted
DSCIM: Digital Stochastic Computing in Memory Featuring Accurate OR Accumulation for Edge AI Models
DATE, 2026 Accepted
McPAL: Scaling Unstructured Sparse Inference with Multi-Chiplet HBM-PIM Architecture for LLMs
DAC, 2025 Accepted
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
ISLPED, 2025 Accepted
FullSparse: A Sparse-Aware GEMM Accelerator with Online Sparsity Prediction
CF, 2024
TPNoC: An Efficient Topology Reconfigurable NoC Generator
GLSVLSI, 2023
NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area Models
ISCAS, 2022
Hobbies: Running and hiking.