MINHO SHIM

Welcome. This is the Research Home of Shim Minho.

Welcome to my website. This is a place where I share my research interests. Have a great time.

About Me

Hi, I’m Minho Shim, a researcher on the Vision Team at NAVER Cloud. My research focuses on video understanding, spanning representation learning, multimodal models, and reliable evaluation. I’m particularly interested in how models understand visual and temporal information, and how we can tell whether they truly do.

Recently, I’ve been exploring how computer vision should account for human subjectivity—the different perspectives, preferences, and interpretations people bring to the same video. Beyond understanding people’s intentions, I aim to extend video understanding toward a deeper understanding of the people who create, watch, and interpret videos.

I connect these research questions to practical systems for video search, summarization, and agentic editing.

Experience

Researcher

Current team: NAVER Cloud · Vision Team

September 2019 – PresentSeoul, Korea

I build video understanding systems, from representation learning and Video LLM training to production APIs and interactive applications. My work spans video analysis, multimodal retrieval, summarization, and agentic video editing.

Media Intelligence & Sinossi

Developed Sinossi, a video metadata engine powering automatic chapterization and action analysis for NAVER ShoppingLive. The engine became the foundation of Media Intelligence (formerly MAIU), NAVER Cloud Platform’s B2B video understanding service. Sinossi received a TOP 3 recognition at NAVER’s N INNOVATION AWARD.

Video Understanding Agent

Building on Sinossi through follow-up research and development on a video understanding agent combining multimodal search, speech and scene summaries, and MCP tools, with natural-language video editing grounded in retrieved video segments.

Video LLM Training & Efficiency

Developed Video LLM training and evaluation pipelines for scene description and summarization, alongside work on long-video understanding and highlight prediction. Developed training-free spatio-temporal token merging for efficient Video LLM inference and integrated it into serving-engine plugins, achieving roughly 2× serving throughput.

Through Video-Oasis, I investigate how to reliably evaluate video understanding by distinguishing visual and temporal understanding from shortcuts based on language priors.

Video Representation Learning & Summarization

Researched temporal and frequency-based data augmentation for video representation learning, alongside masked-autoencoder-based unsupervised video summarization. Created and maintain the team’s shared codebase for large-scale self-supervised pretraining, downstream evaluation, and optimized inference.

Real-Time Talking Head Generation

Built real-time audio-driven talking-head generation and a streaming API.

Baseball Defense Highlight Detection

Built a baseball defense-highlight detection service, including annotation tools and model development.

Person Re-Identification & Tracking

Improved person re-identification and tracking performance for AutoCam and worked on pose-invariant representations and the DanceReID dataset for dance videos. Developed READ, an image-to-video re-identification method that uses reciprocal attention between query images and video sequences to improve identity matching.

Yonsei University

Research Assistant · March 2017 – August 2019

Research Intern · December 2014 – February 2017

Seoul, Korea

Conducted research on video understanding and automatic summarization with Prof. Seon Joo Kim during my master’s studies. In 2018, built a large-scale baseball video database with temporal segment annotations generated semi-automatically from play-by-play text, alongside methods for analyzing baseball videos. The database supports action recognition, temporal localization, text–video alignment, and highlight generation.

Teaching Experience

Teaching Assistant, Artificial Intelligence · SK Telecom · June 2018
8 hours of lectures and labs covering style transfer, autoencoders, and GANs.

Teaching Assistant, Deep Learning · Samsung Electronics · August 2017
20 hours of lectures and labs covering reinforcement learning, GANs, RNNs, image captioning, and video understanding.

Teaching Assistant, Computer Graphics · Yonsei University · Spring 2017
Undergraduate course.

Projects

Video Highlight Generation · Samsung Future Technology Foundation · 2017–2019
Developed a large-scale video database and methods for analyzing its videos.

Video Story Understanding for the Video Turing Test · Ministry of Science and ICT, South Korea · 2017–2019
Researched automatic video summarization for video question answering.

Fake Face Detection · IITP · 2019
1st place in the second-year AI R&D Challenge Fake Face Detecting Competition, developing a module to distinguish images produced by generative models from other images.

Teaching Machines to Dribble using Unsupervised Reinforcement Learning · Yonsei University · 2017
1st place in the Advanced Computer Graphics course project competition; awarded an NVIDIA GTX 1070 GPU.

Color Correction Algorithm for Transparent Displays · LG Display · 2015
Investigated the CIE-L*a*b* color space of transparent displays and compared it with the standard sRGB color space.

GAIA, Web Application for Reducing Food Waste · SWKorea (formerly AppCenter), K-Hackathon · 2014
Reached the semifinals.

Department Webpage Reconstruction · Yonsei University · 2014
Rebuilt the Department of Computer Science website.

Mocoga Games

Software Intern

December 2013 – March 2014Seoul, Korea

Improved EvryPlay, a game publishing platform.

Publications

  1. Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
    Geo Ahn, Inwoong Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, and Jinwoo Choi
    In European Conference on Computer Vision (ECCV), 2026
  2. Video-Oasis: Rethinking Evaluation of Video Understanding
    Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Minho Shim, and Yukyung Choi
    In European Conference on Computer Vision (ECCV), 2026
  3. Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
    Su Ho Han*, Jeongseok Hyun*, Pilhyeon Lee, Minho Shim, Dongyoon Wee, and Seon Joo Kim
    In International Conference on Learning Representations (ICLR), 2026
  4. Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
    Jeongseok Hyun, Sukjun Hwang, Su Ho Han, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Joon-Young Lee, Seon Joo Kim, and Minho Shim
    In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  5. Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
    WonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim, Inwoong Lee, Dongyoon Wee, Minho Shim, and Jae-Pil Heo
    In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  6. Classification Matters: Improving Video Action Detection with Class-Specific Attention
    Jinsung Lee, Taeoh Kim, Inwoong Lee, Minho Shim, Dongyoon Wee, Minsu Cho, and Suha Kwak
    In European Conference on Computer Vision (ECCV), 2024
  7. Towards Multi-Domain Learning for Generalizable Video Anomaly Detection
    MyeongAh Cho, Taeoh Kim, Minho Shim, Dongyoon Wee, and Sangyoun Lee
    In Advances in Neural Information Processing Systems (NeurIPS), 2024
  8. Decomposed Cross-Modal Distillation for RGB-based Temporal Action Detection
    Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, and Hyeran Byun
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
  9. Exploring Temporally Dynamic Data Augmentation for Video Recognition
    Taeoh Kim, Jinhyung Kim, Minho Shim, Sangdoo Yun, Myunggu Kang, Dongyoon Wee, and Sangyoun Lee
    In International Conference on Learning Representations (ICLR), 2023
  10. Frequency Selective Augmentation for Video Representation Learning
    Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, and Junmo Kim
    In AAAI Conference on Artificial Intelligence (AAAI), 2023
  11. Masked Autoencoder for Unsupervised Video Summarization
    Minho Shim, Taeoh Kim, Jinhyung Kim, and Dongyoon Wee
    arXiv, 2023
  12. READ: Reciprocal Attention Discriminator for Image-to-Video Re-Identification
    Minho Shim, Hsuan-I Ho, Jinhyung Kim, and Dongyoon Wee
    In European Conference on Computer Vision (ECCV), 2020
  13. Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking
    Hsuan-I Ho, Minho Shim, and Dongyoon Wee
    In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020
  14. Teaching Machines to Understand Baseball Games: Large-Scale Baseball Video Database for Multiple Video Understanding Tasks
    Minho Shim*, Young Hwi Kim*, Kyungmin Kim*, and Seon Joo Kim
    In European Conference on Computer Vision (ECCV), 2018

* equal contribution, equal corresponding