Xubing Ye (叶栩冰)

Hi! I'm Xubing Ye. I'm currently an algorithm engineer at ByteDance Seed. I obtained my master's degree from Shenzhen International Graduate School at Tsinghua University in 2026, supervised by Prof. Yansong Tang. I obtained my bachelor's degree from the School of Software Engineering at Tongji University in 2023.

My current research focuses on data engineering for agentic LLMs and MLLMs.

Email: yxb_tongji@163.com  /  Github  /  Scholar

profile photo
Recent Publications

* Indicates Equal Contribution

dise VoCo-LLaMA: Towards Vision Compression with Large Language Models
Xubing Ye, Yukang Gan, Xiaoke Huang, Yixiao Ge, Yansong Tang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
[arXiv] [Code]

Compresses hundreds of video tokens into one via attention distillation for efficient long-video understanding.

dise ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
Xubing Ye, Yukang Gan, Yixiao Ge, Xiao-ping Zhang, Yansong Tang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
[arXiv]

Adaptively prunes visual tokens inside LLM decoder layers with minimal performance loss.

MiroThinker benchmark comparison MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team
Technical Report, 2026
[arXiv] [Code] [HF]

An open-source deep research agent advancing model, context, and interactive scaling.

POINTS-Reader training pipeline POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
WeLM Team
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
[arXiv] [Code] [HF]

A distillation-free document conversion model for extracting text, tables, and formulas.

dise Language-Aware Vision Transformer for Referring Segmentation
Xubing Ye*, Zhao Yang*, Jiaqi Wang*, Yansong Tang, Kai Chen, Hengshuang Zhao, Philip H.S. Torr
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI, IF=20.8), 2024
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
[IEEE] [Conference Version] [Code]

A universal referring image and video segmentation framework with language-aware visual encoding.

Selected Honors and Awards
Nanhu Elite Scholarship · Tsinghua University · 清华大学综合优秀奖学金(校级一等) 2025
Zhaoyi Scholarship · Tsinghua University · 清华大学综合优秀奖学金(校级一等) 2024
First Prize Scholarship · Tongji University · 同济大学综合优秀奖学金(校级一等) 2023
Internship Experience
Shanda · MiroMind Team Oct. 2025 – Jan. 2026
Shanghai, China · Work with Song Bai and Jifeng Dai.
Tencent · WXG · WeLM Team Mar. 2025 – Oct. 2025
Shanghai, China · Work with Yuan Liu, Le Tian, and Xiao Zhou.
ByteDance · SEED · Application Dec. 2024 – Mar. 2025
Beijing, China · Work with Baihan Shu, Zeyu Cui, and Cheng Lin.
Tencent · PCG ARC Lab · Applied Research Center Feb. 2024 – Dec. 2024
Shenzhen, China · Work with Yixiao Ge and Ying Shan.
Academic Services
Reviewer
TPAMI 2026 · TIP 2026 · CVPR 2025 · JVCIR 2024, 2025

© Xubing Ye | Last updated: Mar. 17, 2024 | Website Template