NAVER
Researcher
Current team: NAVER Cloud · Vision Team
September 2019 – PresentSeoul, Korea
I build video understanding systems, from representation learning and Video LLM training to production APIs and interactive applications. My work spans video analysis, multimodal retrieval, summarization, and agentic video editing.
Media Intelligence & Sinossi
Developed Sinossi, a video metadata engine powering automatic chapterization and action analysis for NAVER ShoppingLive. The engine became the foundation of Media Intelligence (formerly MAIU), NAVER Cloud Platform’s B2B video understanding service. Sinossi received a TOP 3 recognition at NAVER’s N INNOVATION AWARD.
Video Understanding Agent
Building on Sinossi through follow-up research and development on a video understanding agent combining multimodal search, speech and scene summaries, and MCP tools, with natural-language video editing grounded in retrieved video segments.
Video LLM Training & Efficiency
Developed Video LLM training and evaluation pipelines for scene description and summarization, alongside work on long-video understanding and highlight prediction. Developed training-free spatio-temporal token merging for efficient Video LLM inference and integrated it into serving-engine plugins, achieving roughly 2× serving throughput.
Through Video-Oasis, I investigate how to reliably evaluate video understanding by distinguishing visual and temporal understanding from shortcuts based on language priors.
Video Representation Learning & Summarization
Researched temporal and frequency-based data augmentation for video representation learning, alongside masked-autoencoder-based unsupervised video summarization. Created and maintain the team’s shared codebase for large-scale self-supervised pretraining, downstream evaluation, and optimized inference.
Real-Time Talking Head Generation
Built real-time audio-driven talking-head generation and a streaming API.
Baseball Defense Highlight Detection
Built a baseball defense-highlight detection service, including annotation tools and model development.
Person Re-Identification & Tracking
Improved person re-identification and tracking performance for AutoCam and worked on pose-invariant representations and the DanceReID dataset for dance videos. Developed READ, an image-to-video re-identification method that uses reciprocal attention between query images and video sequences to improve identity matching.