NVIDIA Research Research Intern · Taipei, Taiwan (Remote)
Jan 2026 – Present
Mentors: Ryo Hachiuma and Min-Hung Chen
I am a Ph.D. student at KAIST AI and a research intern at NVIDIA. My research focuses on spatial and 4D understanding for vision–language models — equipping them with fine-grained motion perception (4DP-QA) and agentic spatial reasoning (SpatialClaw). This direction builds on my work on Track Any Point (point tracking) and its applications to 3D and 4D video understanding and reconstruction.
I have had the opportunity to intern at NVIDIA and Adobe, where I was honored to work with inspiring mentors, including Orazio Gallo, Abhishek Badki, Hang Su, and Jindong Jiang at NVIDIA, as well as Joon-Young Lee and Gabriel Huang at Adobe. I am currently working with Ryo Hachiuma and Min-Hung Chen at NVIDIA Research in Taiwan.
* denotes equal contribution. Selected work
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Preprint, 2026
Work done during my internship at NVIDIA.
4DP-QA: Scalable QA for 4D Perception in Vision Language Models
CVPR, 2026
Work done during my internship at NVIDIA.
MV-TAP: Tracking Any Point in Multi-View Videos
CVPR, 2026
AnthroTAP: Learning Point Tracking with Real-World Motion
CVPR, 2026
Pose-dIVE: Pose-Diversified Augmentation for Person Re-Identification
CVPR Findings, 2026
Seurat: From Moving Points to Depth
CVPR, 2025 · Highlight (3.0% acceptance rate)
Work done during my internship at Adobe. Selected as a Qualcomm Innovation Fellowship 2025 finalist.
Exploring Temporally-Aware Features for Point Tracking
CVPR, 2025
DiffFace: Diffusion-based Face Swapping with Facial Guidance
Pattern Recognition, 2025
Multi-Granularity Video Object Segmentation
AAAI, 2025
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
NeurIPS, 2024
Local All-Pair Correspondence for Point Tracking
ECCV, 2024
FlowTrack: Revisiting Optical Flow for Long-Range Dense Tracking
CVPR, 2024
Work done during my internship at Adobe.
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
CVPR, 2024 · Highlight (2.8% acceptance rate)
Unifying Feature and Cost Aggregation with Transformers for Dense Correspondence
ICLR, 2024
DäRF: Boosting Radiance Fields from Sparse Inputs with Monocular Depth Adaptation
NeurIPS, 2023
LANIT: Language-Driven Image-to-Image Translation for Unlabeled Data
CVPR, 2023
MIDMs: Matching Interleaved Diffusion Models for Exemplar-based Image Translation
AAAI, 2023
CATs++: Boosting Cost Aggregation with Convolutions and Transformers
TPAMI, 2023
Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence
NeurIPS, 2022
Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation
ECCV, 2022
CATs: Cost Aggregation Transformers for Visual Correspondence
NeurIPS, 2021
NVIDIA Research Research Intern · Taipei, Taiwan (Remote)
Jan 2026 – Present
Mentors: Ryo Hachiuma and Min-Hung Chen
NVIDIA Research Research Intern · Santa Clara, CA, USA
Jun 2025 – Nov 2025
Mentors: Orazio Gallo, Abhishek Badki, Hang Su, and Jindong Jiang
Adobe Research Research Scientist Intern · San Jose, CA, USA
Jun 2024 – Sep 2024
Mentors: Gabriel Huang and Joon-Young Lee
Adobe Research Research Scientist Intern · San Jose, CA, USA
Jun 2023 – Sep 2023
Mentors: Joon-Young Lee and Gabriel Huang
KAIST Integrated M.S./Ph.D. in Artificial Intelligence · Seoul, Korea
2024 – 2027 (exp.)
Korea University Integrated M.S./Ph.D. in Computer Science and Engineering · Seoul, Korea
2022 – 2024
Transferred to KAIST with supervisor (degree incomplete).
Yonsei University B.S. in Computer Science · Seoul, Korea
2018 – 2022