|
Lei Qinqian
I am a Ph.D. student majoring in Electrical and Computer Engineering(ECE) at National University of Singapore (NUS), supervised by Prof. Robby T. Tan. Before that, I received my Bachelor's Degree from Beijing Institute of Technology, majoring in Automation.
My research lies at the intersection of computer vision, multimodal foundation models, and embodied AI. I develop vision-language models that understand human-object interactions, open-vocabulary visual concepts, and multimodal reasoning. More recently, I have been extending these ideas toward embodied agents, investigating how multimodal foundation models can support reward learning, world modeling, and long-horizon decision making under sparse supervision. My long-term goal is to build foundation models that can reason about actions, interactions, and their consequences in open-world environments.
Email: qinqian.lei@u.nus.edu
Google Scholar  / 
Github  / 
Linkdin
|
|
|
[Jul. 2026] Honored to be recognized as an Outstanding Reviewer for ECCV 2026!
[Jun. 2026] Our work "Token-Based Affordance Grounding with Large Vision-Language Models" is accepted by ECCV 2026!
[Jun. 2026] Our work "SHOE: Semantic HOI Open-Vocabulary Evaluation Metric" is accepted by the CVPR 2026 GRAIL-V workshop and receives the Best Paper Award!
[Feb. 2026] Our work "CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods" is accepted by CVPR 2026!
[Jun. 2025] Our work "HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation" is accepted by ICCV 2025!
[Sep. 2024] Our work "Efficient Zero-Shot HOI Detection: Enhancing VLM Adaptation with Innovative Prompt Learning" is accepted by NeurIPS 2024!
[Dec. 2023] Our work "Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOI" is accepted by AAAI 2024!
|
|
|
Token-Based Affordance Grounding with Large Vision-Language Models
Seung Il Lee,
Qinqian Lei,
Daguang Xu,
Dong Yang,
Robby T. Tan,
Yixin Chen,
Bo Wang
European Conference on Computer Vision (ECCV), 2026
PDF  / 
Code
TokAG is a zero-shot affordance grounding framework that leverages latent semantic-spatial signals within pretrained LVLMs, eliminating the need for affordance-specific supervision.
|
|
|
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
Qinqian Lei,
Bo Wang,
Robby T. Tan
Computer Vision and Pattern Recognition (CVPR), 2026
PDF  / 
Code
CrossHOI-Bench is the first work introducing a unified HOI evaluation, revealing complementary pros and cons of VLMs and HOI models.
|
|
|
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
Qinqian Lei,
Bo Wang,
Robby T. Tan
International Conference on Computer Vision (ICCV), 2025
PDF  / 
Code
HOLa boosts zero-shot HOI detection with low-rank VLM feature adaptation, achieving better action distinction and generalization ability.
|
|
|
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
Qinqian Lei,
Bo Wang,
Robby T. Tan
Neural Information Processing Systems (NeurIPS), 2024
PDF  / 
Code
We introduce a novel Efficient Zero-shot HOI Detection (EZ-HOI) method to enhance VLM adaptation through prompt learning with a small computational cost.
|
|
|
Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOI
Qinqian Lei,
Bo Wang,
Robby T. Tan
Association for the Advancement of Artical Intelligence (AAAI), 2024
PDF  / 
Code
We propose novel label-uncertain augmentations to enhance query diversity, facilitating effective learning of Bongard-HOI by learning from those new augmented queries.
|
|
[2026 Jul. ] Outstanding Reviewer, ECCV 2026
[2026 Jun. ] Best Paper Award, CVPR 2026 GRAIL-V Workshop
[2022 Jun. ] Excellent Graduate of Beijing
[2019, 2021 Sep. ] National Scholarship of China
|
The template of this page is from Jon Barron.
|
|