GAUDI LabGeneral-Purpose Architectures with Unleashed Design Innovations
Menu

GAUDI LAB / YONSEI UNIVERSITY

Research

From the processor core to the complete computing system.


Our research topics include, but are not limited to, the following areas:

CPU Microarchitecture

CPU pipeline with prefetcher and branch predictor, six-wide decode and allocation, reconfigurable IQ cluster, twelve-wide issue, PRF scoreboard, L1-D cache and retirement structures
Description

The Central Processing Unit (CPU) is the foundation of modern computing systems, executing applications and system software while coordinating memory, accelerators, and I/O devices. We investigate microarchitectural and multi-core techniques that improve performance, energy efficiency, scalability, and adaptability across general-purpose, datacenter, and on-device workloads.

Topics
  • Dynamic instruction scheduling, issue-queue design, and scalable out-of-order execution
  • Register renaming, early resource reclamation, and efficient instruction commit
  • Reconfigurable, composable, and heterogeneous multi-core architectures
  • Simultaneous multithreading (SMT), speculation, branch prediction, and redundant-load elimination
  • Energy-efficient processor design, functional-unit gating, and architectural simulation

Representative Publications


GPU & Accelerators

GPU streaming multiprocessor hierarchy and detailed register-bank, scheduler and operand-collector datapaths
Description

Workloads in graphics, AI, and data analytics require highly parallel and specialized execution. We study energy-efficient GPGPU, NPU, and on-chip accelerator architectures, focusing on the microarchitectural and memory-system bottlenecks that limit performance, programmability, and quality of service.

Topics
  • GPU register-file bandwidth, hierarchical register storage, and operand collection
  • Warp scheduling, inter-warp value reuse, and Tensor Core execution
  • GPU virtual memory, concurrent page-table walks, UVM, and page-fault handling
  • Adaptive-precision NPUs and multi-DNN acceleration with balanced QoS and throughput
  • CPU–GPU/NPU cooperation, accelerator offloading, and heterogeneous simulation frameworks

Representative Publications


Memory & CXL

CXL Type-2 device controller with PHY, DCOH, CXL.io, user AFUs, memory controllers and DDR channels
Description

Emerging coherent interconnects such as Compute Express Link (CXL) connect CPUs with memory expanders, accelerators, SmartNICs, and computational storage. We investigate CXL devices and fabrics from microarchitecture to system software, aiming to enable efficient memory expansion, coherent offloading, resource sharing, and dependable heterogeneous computing.

Topics
  • Characterization and optimization of genuine CXL memory systems and devices
  • CXL Type-2 accelerators and Type-3 memory-expansion architectures
  • Coherence protocols, memory-side directories, and scalable CXL memory fabrics
  • CXL-enabled end-host networking, accelerator offloading, and host–device cooperation
  • Reliability and security mechanisms for disaggregated memory, including RowHammer defense

Representative Publications


System Orchestration

Four interconnected cores and LLC/CHA/SF regions with memory controllers, GPU, SmartNIC, SmartSSD, DRAM, PIM and CXL devices
Description

Modern systems combine CPUs, GPUs, NPUs, memory, caches, networks, and storage devices whose shared resources interact in complex ways. We design system architectures and orchestration mechanisms that coordinate these components, reduce interference and data movement, and deliver predictable performance across datacenter and on-device platforms.

Topics
  • Microarchitecture-aware LLC management and Data Direct I/O for emerging I/O devices
  • CPU–accelerator coordination and offloading of datacenter and system-level tasks
  • CPU–NPU contention management and QoS for ultra-low-latency on-device AI SoCs
  • SmartNIC–host cooperative computing and hardware-assisted load balancing
  • Resource scheduling, performance characterization, and simulation for heterogeneous systems and chiplets

Representative Publications


Near-Data Processing

Computational SSD with a highlighted Compute Engine, DRAM, controller internals, FTL functions and four NAND Flash banks connected to the flash interface
Description

Data-centric applications are increasingly limited by moving data between processors, memory, storage, and networks. We bring computation closer to data through Processing-In/Near-Memory (PIM/PNM), In-Storage Processing (ISP), and In-Network Computing (INC), co-designing architectures, execution models, and system software to reduce movement overheads.

Topics
  • PIM/PNM architectures, dataflow execution models, and multi-PIM task scheduling
  • CXL-based near-memory acceleration and simulation platforms for LLM workloads
  • Computational SSD architectures and predicate pushdown for data-intensive applications
  • SSD latency prediction, scalable page caching, and storage-aware system optimization
  • Near-network processing and cooperative execution across SmartNICs, storage, and hosts

Representative Publications


Collaboration

**Future Architecture and System Technology for Scalable Computing (FAST)**  University of Illinois Urbana-Champaign (UIUC), IL, United States

Future Architecture and System Technology for Scalable Computing (FAST)
University of Illinois Urbana-Champaign (UIUC), IL, United States

**Embedded Systems and Computer Architecture Lab (eSCaL)**  Yonsei University, Seoul, Republic of Korea

Embedded Systems and Computer Architecture Lab (eSCaL)
Yonsei University, Seoul, Republic of Korea