Contents
Education
-
Since 2025 | Cornell University | Ph.D.
Ph.D. student, Computer Science
-
2019 – 2025 | Seoul National University | B.S.
Computer Science & Engineering, minor in Linguistics
Research Experience
-
Since 2025 | Cornell University | Advisors: Prof. Mohamed S. Abdelfattah and Prof. Jae-sun Seo
Built a neural-network/analog-hardware co-design framework for hardware-aware evaluation (ICCAD 2026).
Developed GPU kernels with robust end-to-end inference evaluation for RaZeR.
Optimized GPU kernels to accelerate experimentation and evaluation for HBQ (MICRO 2026).
-
2023 – 2025 | SNU Architecture and Code Optimization Lab | Advisor: Prof. Jae W. Lee
Built the Any-Precision LLM quantization pipeline, cutting runtime by over four orders of magnitude (ICML 2024).
Built the low-bit inference system underlying DecDEC (OSDI 2025).
-
2023 | SNU Human-Computer Interaction Lab | Advisor: Prof. Jinwook Seo
Contributed systems optimizations to UMATO (TVCG 2025) and ZADU (VIS 2023).
Publications
-
2026 | HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference
Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed S. Abdelfattah, and Jae-sun Seo
IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)
-
2026 | Algorithm-Hardware Co-Design of a Robust Large-Scale Analog MAC for Neural Network Acceleration using Single-Cut Chopping
Maysara Hamada, Youngrae Kim, Sreetama Sakar, Chee-An Yu, Jake Hyun, Huixuan Yin, Hsiang-Chun Cheng, Yifan He, Jian Meng, Jae-sun Seo, C.-C. Jay Kuo, Peter A. Beerel, and Shuo-Wei Chen
IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)

-
2025 | DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization Code
Yeonhong Park*, Jake Hyun*, Hojoon Kim, Jae W. Lee (* Equal Contribution)
Symposium on Operating Systems Design and Implementation (OSDI '25)
We present a novel inference scheme for low-bit quantized LLMs that dynamically mitigates quantization errors on a per-token basis. By leveraging CPU memory to store residuals and selectively fetching critical channels over PCIe in real time, our approach significantly enhances LLM performance with only minimal overheads in memory usage and latency.

-
2025 | UMATO: Bridging Local and Global Structures for Reliable Visual Analytics with Dimensionality Reduction Code
Hyeon Jeon, Kwon Ko, Soohyun Lee, Jake Hyun, Taehyun Yang, Gyehun Go, Jaemin Jo, Jinwook Seo
IEEE Transactions on Visualization and Computer Graphics (TVCG '25)
We provide deeper insights into UMATO, a novel dimensionality reduction technique that achieves state-of-the-art performance in terms of accuracy, scalability, and stability, preserving both local and global structures of the data.

-
2024 | Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs Code
Yeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim, Jae W. Lee
International Conference on Machine Learning (ICML '24) - Oral Presentation (top 1.5%)
Any-precision LLM enables the creation of variable bitrate models, significantly reducing deployment costs for multiple Large Language Models (LLMs) through lightweight post-training quantization and optimized software.

-
2023 | ZADU: A Python Library for Evaluating the Reliability of Dimensionality Reduction Embeddings Code
Hyeon Jeon, Aeri Cho, Jinhwa Jang, Soohyun Lee, Jake Hyun, Hyung-Kwon Ko, Jaemin Jo, Jinwook Seo
IEEE Visualization Conference (VIS '23)
ZADU is a Python library that offers efficient and comprehensive evaluation of dimensionality reduction (DR) embeddings through optimized distortion measures.
Preprints
-
2026 | RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
Yuzong Chen*, Xilai Dai*, Jake Hyun*, Chi-Chih Chang, Wonsuk Jang, Yuheng Wu, Thierry Tambe, Jae-sun Seo, and Mohamed S. Abdelfattah (* Equal Contribution)
Preprint
Awards & Achievements
-
2024 | Accelerator Programming Winter School, CUDA competition | 1st
place, team of two
Organized by SNU THUNDER Research Group & Manycoresoft
1st place by performance, final project on model inference throughput optimization using CUDA C++.
-
2022 | Korean AI Competition | 1st place, undergrad div.,
team of four, prize: $8,000
Organized by Korea Ministry of Science and ICT, National Information Society Agency
Developed a speech-to-text model for the Korean language & its dialects.
Awarded by Korean Minister of Science and Technology. -
2020 | SNUH Medical AI Challenge | 4th place, team of
11
Organized by Seoul National University Hospital
Developed an intraoperative hypotension predictor from arterial pressure waveforms.
-
2020 | Digital Health Hackathon | 1st place, team of
three, prize:
$2,500
Organized by Samsung Advanced Institute for Health Sciences & Technology, Digital Healthcare Partners
Created a drug treatment decision model for a rare cancer utilizing a two-model ensemble approach.
-
2017 | Korean Olympiad in Informatics, project division | Silver (3rd
place)
Organized by Korea Ministry of Science and ICT
Created an RL-based AI agent for the games of Othello and Omok.
Open Source Contributions
-
2024 | flash1dkmeans | GitHub
Library implementation of the novel log-time $k$-means algorithms proposed in my thesis, used in Any-Precision LLM for quantization.
-
2024 | SNU CSE Thesis in LaTeX Format | GitHub
LaTeX source of my undergraduate thesis, shared for SNU CSE students to reference the formatting.Feel free to adapt the structure—please do not reuse the thesis content itself.
-
2023 | Steadiness & Cohesiveness | GitHub
Metrics for evaluating the reliability of dimensionality reduction embeddings, used in ZADU to provide a comprehensive evaluation of DR techniques.
CS 박사 유학 지원 수기 / My CS PhD Application Journey (Korean)
CS 박사 유학 준비 과정에서의 제 경험과 생각을 정리한 수기를 적었습니다.
유학을 준비하시는 분들께 도움이 되기를 바랍니다. 아래 링크에서 확인하실 수 있습니다.