|
Recent Publications
* Indicates Equal Contribution
|
|
VoCo-LLaMA: Towards Vision Compression with Large Language Models
Xubing Ye,
Yukang Gan,
Xiaoke Huang,
Yixiao Ge,
Yansong Tang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
[arXiv]
[Code]
Compresses hundreds of video tokens into one via attention distillation for efficient long-video understanding.
|
|
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
Xubing Ye,
Yukang Gan,
Yixiao Ge,
Xiao-ping Zhang,
Yansong Tang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
[arXiv]
Adaptively prunes visual tokens inside LLM decoder layers with minimal performance loss.
|
|
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team
Technical Report, 2026
[arXiv]
[Code]
[HF]
An open-source deep research agent advancing model, context, and interactive scaling.
|
|
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
WeLM Team
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
[arXiv]
[Code]
[HF]
A distillation-free document conversion model for extracting text, tables, and formulas.
|
|
Language-Aware Vision Transformer for Referring Segmentation
Xubing Ye*,
Zhao Yang*,
Jiaqi Wang*,
Yansong Tang,
Kai Chen,
Hengshuang Zhao,
Philip H.S. Torr
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI, IF=20.8), 2024
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
[IEEE]
[Conference Version]
[Code]
A universal referring image and video segmentation framework with language-aware visual encoding.
|
|
Selected Honors and Awards
|
|
|
Shanghai, China · Work with Yuan Liu, Le Tian, and Xiao Zhou.
|
|
Beijing, China · Work with Baihan Shu, Zeyu Cui, and Cheng Lin.
|
|
|
|
TPAMI 2026 · TIP 2026 · CVPR 2025 · JVCIR 2024, 2025
|
© Xubing Ye | Last updated: Mar. 17, 2024 | Website Template
|