Modern high-performance computing (HPC) systems operate at massive scales, comprising thousands of nodes equipped with high-end CPUs and GPUs to support complex workloads such as large language model training, quantum simulation, and high-resolution scientific simulations. As these systems continue to scale, two major challenges identified by the U.S. Department of Energy (DOE) become increasingly critical: managing the growing volume of data and ensuring robust error resilience. My research addresses both challenges by developing flexible, efficient, and broadly applicable software solutions. On the data-efficiency side, I design ultra-fast GPU-based compression frameworks, such as cuSZp, that achieve high compression ratios while preserving data fidelity for diverse applications. On the reliability side, I develop low-overhead fault-tolerance techniques that enable effective detection of complex faults with minimal performance impact. Together, these designs provide scalable software solutions that improve data efficiency and reliability in next-generation HPC and AI systems.
Yafan Huang is an Assistant Professor in the Department of Computer Science at the University of Maryland, College Park. His research focuses on high-performance computing (HPC), with particular interests in data compression, fault tolerance, parallel computing, and compiler optimizations. He develops efficient and resilient computing systems to address the growing computational and data challenges in scientific computing and AI applications. Previously, he was a visiting graduate student at Argonne National Laboratory. He received his Ph.D. in Computer Science from the University of Iowa. He is a recipient of the 2025 ACM-IEEE CS George Michael Memorial HPC Fellowship and was named a 2026 MLCommons ML and Systems Rising Star. His research has also received multiple best paper and best student paper recognitions at leading HPC conferences.

