Jake Hyun

Jake Hyun

Hi, I'm Jake Hyun🇰🇷 현재익 [hjʌn dʑɛik], a Ph.D. student in Computer Science at Cornell University, co-advised by Prof. Mohamed S. Abdelfattah and Prof. Jae-sun Seo as part of the Computer Systems Laboratory. Before beginning my Ph.D. at Cornell, I earned a B.S. in Computer Science and Engineering and a minor in Linguistics from Seoul National University in 2025.

My research focuses on hardware–software co-design and efficient LLM inference, spanning low-bit quantization and agentic inference systems.

At Cornell, I develop frameworks for neural-network and hardware co-design and build GPU kernels for low-bit inference. Previously, at SNU, I worked on the quantization pipeline for Any-Precision LLM and the inference system underlying DecDEC.

Contents

Education

Research Experience

  • Since 2025 | Cornell University | Advisors: Prof. Mohamed S. Abdelfattah and Prof. Jae-sun Seo

    Built a neural-network/analog-hardware co-design framework for hardware-aware evaluation (ICCAD 2026).

    Developed GPU kernels with robust end-to-end inference evaluation for RaZeR.

    Optimized GPU kernels to accelerate experimentation and evaluation for HBQ (MICRO 2026).

  • 2023 – 2025 | SNU Architecture and Code Optimization Lab | Advisor: Prof. Jae W. Lee

    Built the Any-Precision LLM quantization pipeline, cutting runtime by over four orders of magnitude (ICML 2024).

    Built the low-bit inference system underlying DecDEC (OSDI 2025).

  • 2023 | SNU Human-Computer Interaction Lab | Advisor: Prof. Jinwook Seo

    Contributed systems optimizations to UMATO (TVCG 2025) and ZADU (VIS 2023).

Publications

  • 2026 | HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference

    Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed S. Abdelfattah, and Jae-sun Seo

    IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

    HBQ overview: relative accuracy and efficiency trade-off between weight-only, block, and hierarchical block quantization
  • 2026 | Algorithm-Hardware Co-Design of a Robust Large-Scale Analog MAC for Neural Network Acceleration using Single-Cut Chopping

    Maysara Hamada, Youngrae Kim, Sreetama Sakar, Chee-An Yu, Jake Hyun, Huixuan Yin, Hsiang-Chun Cheng, Yifan He, Jian Meng, Jae-sun Seo, C.-C. Jay Kuo, Peter A. Beerel, and Shuo-Wei Chen

    IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)

    Single-cut chopping: proposed concept and efficient implementation with a shared output chopper
  • 2025 | DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization Code

    Yeonhong Park*, Jake Hyun*, Hojoon Kim, Jae W. Lee (* Equal Contribution)

    Symposium on Operating Systems Design and Implementation (OSDI '25)

    We present a novel inference scheme for low-bit quantized LLMs that dynamically mitigates quantization errors on a per-token basis. By leveraging CPU memory to store residuals and selectively fetching critical channels over PCIe in real time, our approach significantly enhances LLM performance with only minimal overheads in memory usage and latency.

    OSDI25
  • 2025 | UMATO: Bridging Local and Global Structures for Reliable Visual Analytics with Dimensionality Reduction Code

    Hyeon Jeon, Kwon Ko, Soohyun Lee, Jake Hyun, Taehyun Yang, Gyehun Go, Jaemin Jo, Jinwook Seo

    IEEE Transactions on Visualization and Computer Graphics (TVCG '25)

    We provide deeper insights into UMATO, a novel dimensionality reduction technique that achieves state-of-the-art performance in terms of accuracy, scalability, and stability, preserving both local and global structures of the data.

    UMATO
  • 2024 | Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs Code

    Yeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim, Jae W. Lee

    International Conference on Machine Learning (ICML '24) - Oral Presentation (top 1.5%)

    Any-precision LLM enables the creation of variable bitrate models, significantly reducing deployment costs for multiple Large Language Models (LLMs) through lightweight post-training quantization and optimized software.

    Any-Precision LLM
  • 2023 | ZADU: A Python Library for Evaluating the Reliability of Dimensionality Reduction Embeddings Code

    Hyeon Jeon, Aeri Cho, Jinhwa Jang, Soohyun Lee, Jake Hyun, Hyung-Kwon Ko, Jaemin Jo, Jinwook Seo

    IEEE Visualization Conference (VIS '23)

    ZADU is a Python library that offers efficient and comprehensive evaluation of dimensionality reduction (DR) embeddings through optimized distortion measures.

    ZADU: UMAP projection, CheckViz, and reliability map

Preprints

Awards & Achievements

  • 2024 | Accelerator Programming Winter School, CUDA competition | 1st place, team of two

    Organized by SNU THUNDER Research Group & Manycoresoft

    1st place by performance, final project on model inference throughput optimization using CUDA C++.

  • 2022 | Korean AI Competition | 1st place, undergrad div., team of four, prize: $8,000

    Organized by Korea Ministry of Science and ICT, National Information Society Agency

    Developed a speech-to-text model for the Korean language & its dialects.
    Awarded by Korean Minister of Science and Technology.

  • 2020 | SNUH Medical AI Challenge | 4th place, team of 11

    Organized by Seoul National University Hospital

    Developed an intraoperative hypotension predictor from arterial pressure waveforms.

  • 2020 | Digital Health Hackathon | 1st place, team of three, prize: $2,500

    Organized by Samsung Advanced Institute for Health Sciences & Technology, Digital Healthcare Partners

    Created a drug treatment decision model for a rare cancer utilizing a two-model ensemble approach.

  • 2017 | Korean Olympiad in Informatics, project division | Silver (3rd place)

    Organized by Korea Ministry of Science and ICT

    Created an RL-based AI agent for the games of Othello and Omok.

Open Source Contributions

  • 2024 | flash1dkmeans | GitHub

    Library implementation of the novel log-time $k$-means algorithms proposed in my thesis, used in Any-Precision LLM for quantization.

  • 2024 | SNU CSE Thesis in LaTeX Format | GitHub

    LaTeX source of my undergraduate thesis, shared for SNU CSE students to reference the formatting.
    Feel free to adapt the structure—please do not reuse the thesis content itself.

  • 2023 | Steadiness & Cohesiveness | GitHub

    Metrics for evaluating the reliability of dimensionality reduction embeddings, used in ZADU to provide a comprehensive evaluation of DR techniques.

CS 박사 유학 지원 수기 / My CS PhD Application Journey (Korean)

CS 박사 유학 준비 과정에서의 제 경험과 생각을 정리한 수기를 적었습니다.
유학을 준비하시는 분들께 도움이 되기를 바랍니다. 아래 링크에서 확인하실 수 있습니다.

수기 보러가기