Lei Qinqian

I am a Ph.D. student majoring in Electrical and Computer Engineering(ECE) at National University of Singapore (NUS), supervised by Prof. Robby T. Tan. Before that, I received my Bachelor's Degree from Beijing Institute of Technology, majoring in Automation.

My research lies at the intersection of computer vision, multimodal foundation models, and embodied AI. I develop vision-language models that understand human-object interactions, open-vocabulary visual concepts, and multimodal reasoning. More recently, I have been extending these ideas toward embodied agents, investigating how multimodal foundation models can support reward learning, world modeling, and long-horizon decision making under sparse supervision. My long-term goal is to build foundation models that can reason about actions, interactions, and their consequences in open-world environments.

Email: qinqian.lei@u.nus.edu

Google Scholar  /  Github  /  Linkdin

profile photo
News

[Jul. 2026] Honored to be recognized as an Outstanding Reviewer for ECCV 2026!

[Jun. 2026] Our work "Token-Based Affordance Grounding with Large Vision-Language Models" is accepted by ECCV 2026!

[Jun. 2026] Our work "SHOE: Semantic HOI Open-Vocabulary Evaluation Metric" is accepted by the CVPR 2026 GRAIL-V workshop and receives the Best Paper Award!

[Feb. 2026] Our work "CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods" is accepted by CVPR 2026!

[Jun. 2025] Our work "HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation" is accepted by ICCV 2025!

[Sep. 2024] Our work "Efficient Zero-Shot HOI Detection: Enhancing VLM Adaptation with Innovative Prompt Learning" is accepted by NeurIPS 2024!

[Dec. 2023] Our work "Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOI" is accepted by AAAI 2024!

Publications
Token-Based Affordance Grounding with Large Vision-Language Models
Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang
European Conference on Computer Vision (ECCV), 2026
PDF  /  Code

TokAG is a zero-shot affordance grounding framework that leverages latent semantic-spatial signals within pretrained LVLMs, eliminating the need for affordance-specific supervision.

CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
Qinqian Lei, Bo Wang, Robby T. Tan
Computer Vision and Pattern Recognition (CVPR), 2026
PDF  /  Code

CrossHOI-Bench is the first work introducing a unified HOI evaluation, revealing complementary pros and cons of VLMs and HOI models.

HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
Qinqian Lei, Bo Wang, Robby T. Tan
International Conference on Computer Vision (ICCV), 2025
PDF  /  Code

HOLa boosts zero-shot HOI detection with low-rank VLM feature adaptation, achieving better action distinction and generalization ability.

EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
Qinqian Lei, Bo Wang, Robby T. Tan
Neural Information Processing Systems (NeurIPS), 2024
PDF  /  Code

We introduce a novel Efficient Zero-shot HOI Detection (EZ-HOI) method to enhance VLM adaptation through prompt learning with a small computational cost.

Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOI
Qinqian Lei, Bo Wang, Robby T. Tan
Association for the Advancement of Artical Intelligence (AAAI), 2024
PDF  /  Code

We propose novel label-uncertain augmentations to enhance query diversity, facilitating effective learning of Bongard-HOI by learning from those new augmented queries.

Honors & Awards

[2026 Jul. ] Outstanding Reviewer, ECCV 2026

[2026 Jun. ] Best Paper Award, CVPR 2026 GRAIL-V Workshop

[2022 Jun. ] Excellent Graduate of Beijing

[2019, 2021 Sep. ] National Scholarship of China



The template of this page is from Jon Barron.