Jiangnan Yu

Ph.D. Student, HKUST ECE

Jiangnan Yu

Computer Architecture · ML Systems · Hardware-Software Co-Design

I am a second-year Ph.D. student at The Hong Kong University of Science and Technology (HKUST), advised by Prof. Yuan Xie. My research focuses on efficient AI systems and computer architecture, especially CPU-GPU serving for LLM and MoE workloads, memory-centric acceleration, and hardware-software co-design for irregular AI computation.

Before joining HKUST, I received my M.S. and B.S. in Microelectronics from Fudan University.

  • FocusAI serving & memory-centric systems
  • MethodHardware-software co-design
  • CurrentPh.D. @ HKUST ECE
  • AdvisorProf. Yuan Xie

Efficient AI Systems

I build efficient AI systems and infrastructure, spanning AI serving, memory-centric computing, and hardware-software co-design.

Selected Projects

A few recent directions that connect my work across AI serving systems, memory-centric acceleration, and architecture design.

AI Serving

Delta-MoE

Rethinks transfer granularity for CPU-GPU MoE serving to reduce data movement and improve efficient model execution.

ICCAD 2026 · Accepted

Compute-in-Memory

SwiftCIM

Explores ReRAM-coupled digital CIM with a fully-fused multi-head attention dataflow for FlashAttention-style workloads.

ESSERC 2026 · Accepted

Edge AI Acceleration

DSCIM

Builds digital stochastic computing-in-memory support with accurate OR accumulation for energy-efficient edge AI models.

DATE 2026 · Accepted

News

Education

Selected Publications

*: Equal contributions

  1. SwiftCIM: A 55nm 23.2µJ/Token L-0.5 ReRAM Coupled Digital CIM Accelerator with Fully-Fused Multi-Head Attention Dataflow for FlashAttention

    Jiangnan Yu, et al.

    ESSERC, 2026 Accepted

  2. Delta-MoE: Rethinking Transfer Granularity for CPU-GPU MoE Serving

    Jiangnan Yu

    ICCAD, 2026 Accepted

  3. Scalable Sparse Transformer Accelerator with In-Memory Butterfly Zero Skipper and Local Attention Reusable Engine for Irregular-Pruned NN

    Jiangnan Yu, et al.

    IEEE JETCAS, 2026 Accepted

  4. DSCIM: Digital Stochastic Computing in Memory Featuring Accurate OR Accumulation for Edge AI Models

    Jiangnan Yu*, et al.

    DATE, 2026 Accepted

  5. McPAL: Scaling Unstructured Sparse Inference with Multi-Chiplet HBM-PIM Architecture for LLMs

    Shiwei Liu, Jiangnan Yu, et al.

    DAC, 2025 Accepted

  6. DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation

    Jiangnan Yu*, et al.

    ISLPED, 2025 Accepted

  7. FullSparse: A Sparse-Aware GEMM Accelerator with Online Sparsity Prediction

    Jiangnan Yu, Yang Fan, et al.

    CF, 2024

  8. TPNoC: An Efficient Topology Reconfigurable NoC Generator

    Jiangnan Yu, et al.

    GLSVLSI, 2023

  9. NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area Models

    X. Yi, Jiangnan Yu, et al.

    ISCAS, 2022

Misc

Hobbies: Running and hiking.