website picture

Boyuan Sun (孙博远) is currently a 4th year Ph.D. candidate at Nankai University, supervised by Prof. Qibin Hou and Prof. Ming-Ming Cheng. He received his bachelor's degree from the School of Computer Science and Technology at Xidian University in 2021. And now he is taking the Master-Ph.D. combined program in Nankai University. His research interests include Computer Vision and Multimodal Large Language Model, particularly focusing on multi-modal visual perception, vison-language model, semi-supervised learning, etc.

Internship Experience

Ubiquant
Aug 2026 – Present
Research Intern

Working on foundation models for agentic science, focusing on AI for Science and building Science Agent systems.

Agentic Science AI for Science Science Agent
ByteDance
Apr 2026 – Aug 2026
Research Intern

Working on streaming video understanding and omni models. Focusing on agentic RL and on-policy training for streaming video.

Streaming Video Understanding Omni Model
Shanghai AI Laboratory
Dec 2025 – Apr 2026
Research Intern

Working on Science Discovery Agent System, focusing on MLE agentic model design, Data analyze agent, and agentic model training.

Science Discovery Agent MLE Data Analyze Agent
Tongyi Lab, Alibaba
Sep 2024 – Dec 2025
Research Intern

Focused on video understanding techniques, especially on training-free token compression, omni model architecture, and fine-grained object understanding for large multimodal models.

Token Compression Video Understanding Video-LLM

Selected Publications

* Equal contribution. # Corresponding author.

SAIL: Scientific Agentic Intelligence via a Science-Aware Loop

Boyuan Sun, Bryan Dai, Che Liu, Chi Liu, Derek Li, Hongming Piao, Mengzhuo Chen, Xidong Wang, Yan Shu, Yinda Chen, Ziyang Zeng

Technical Report 2026

A scientific agent model that uses failure diagnosis to guide training-task construction for literature research, scientific coding, and multi-step research workflows.

IQuest-Q1: Advancing Foundation Capabilities for Agentic CLI Systems

IQuest Team

Open-Weight Model 2026

A 320B mixture-of-experts model with 15B active parameters for agentic coding, reasoning, and multi-step tool use.

DeltaStreamer: Towards Human-Like Agentic Streaming Video Understanding

Boyuan Sun, Deshui Miao, Wenzhao Gao, Shaoyong Jia, Shaohui Jiao, Chi Liu, Yan Shu, Derek Li, Bryan Dai, Shengsheng Qian, Qibin Hou

Preprint 2026

An agentic streaming video framework that records visual changes in structured memory and uses query-driven planning and targeted replay to recover relevant evidence.

arxiv

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Lei Bai, Zongsheng Cao, Yang Chen, Zhiyao Cui, Shangheng Du, Yue Fan, Shiyang Feng, Zijie Guo, Haonan He, Liang He, Xiaohan He, Shuyue Hu, Yusong Hu, Songtao Huang, Yichen Jiang, Hao Li, Xin Li, Dahua Lin, Weihao Lin, Fenghua Ling, Dongrui Liu, Zhuo Liu, Wenjie Lou, Runmin Ma, Chunjiang Mu, Haoyang Peng, Tianshuo Peng, Jinxin Shi, Luohe Shi, Boyuan Sun, Zelin Tan, Shengji Tang, Yan Teng, Qianyi Wang, Xiaosong Wang, Yiming Wu, Yi Xie, Xiangchao Yan, Jingqi Ye, Peng Ye, Fangchen Yu, Jiakang Yuan, Bihao Zhan, Bo Zhang, Chen Zhang, Shufei Zhang, Shuaiyu Zhang, Wenlong Zhang, Yiqun Zhang, Junpeng Zhao, Zhijie Zhong, Bowen Zhou, Yuhao Zhou

Technical Report

A long-horizon scientific agent that combines domain-specialist training and multi-teacher on-policy distillation to support extended research workflows.

arxiv

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao, Shiyang Feng, Zichen Liang, Boyuan Sun, Tianshuo Peng, Yifan Zhou, Xin Li, Jie Zhou, Liang He, Bo Zhang, Lei Bai

Technical Report

A machine learning discovery framework that combines search, experimental feedback, and reusable experience to iteratively improve algorithms.

arxiv

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

Boyuan Sun, Bo-Wen Yin, Yuan-Ming Li, Xihan Wei, Qibin Hou

IEEE Computer Vision and Pattern Recognition 2026 (CVPR 2026)

Training-time mask supervision aligns natural-language references with visual objects for fine-grained video understanding without visual prompts at inference.

arxiv

GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics

Modi Jin, Yiming Zhang, Boyuan Sun, Dingwen Zhang, Ming-Ming Cheng, Qibin Hou

IEEE Computer Vision and Pattern Recognition 2026 (CVPR 2026 Highlight)

A visual geolocation model trained with expert reasoning data and geographic similarity and consistency rewards to infer fine-grained locations.

arxiv

Depth Anything at Any Condition

Boyuan Sun*, Modi Jin*, Bowen Yin, Qibin Hou#

Submitted to TPAMI

Consistency learning and spatial relation constraints adapt depth foundation models to adverse imaging conditions while preserving general-scene capabilities.

arxiv

LLaVA-Scissor: token Compression with Semantic Connected Components for Video LLMs

Boyuan Sun*, Jiaxing Zhao*, Xihan Wei, Qibin Hou#

Submitted to TMM

Training-free video token compression uses semantic connected components to reduce spatial and temporal redundancy while prioritizing semantic coverage.

arxiv

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Qize Yang*, Shimin Yao*, Weixuan Chen*, Shenghao Fu, Detao Bai, Jiaxing Zhao, Boyuan Sun, Bowen Yin, Xihan Wei, Jingren Zhou

Technical Report

A multimodal reasoning model that organizes audiovisual context before reasoning and uses reinforcement learning to encourage evidence-grounded answers.

arxiv

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Jiaxing Zhao*, Qize Yang*, Yixing Peng*, Detao Bai* Shimin Yao*, Boyuan Sun, Xiang Chen, Shenghao Fu, Weixuan Chen, Xihan Wei, Liefeng Bo#

Technical Report

A human-centric vision-speech model that combines face, body, and interaction representations with audio for richer video understanding.

arxiv

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Boyuan Sun*, Jiaxing Zhao*, Xiang Chen, Xihan Wei, Qibin Hou#

Arxiv preprint 2025

Instruction-driven fusion of complementary visual projectors adapts video representations to the static and temporal information required by each question.

arxiv

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

Jiaxing Zhao*, Boyuan Sun*, Xiang Chen, Xihan Wei

Association for the Advancement of Artificial Intelligence 2026 (AAAI 2026)

A facial dynamics dataset and instruction-tuning framework support video-based expression understanding, with evaluation of event content and temporal order.

arxiv

AODRaw: Towards RAW Object Detection in Diverse Conditions

Zhong-Yu Li, Xin Jin, Boyuan Sun, Chun-Le Guo, Ming-Ming Cheng#

IEEE Computer Vision and Pattern Recognition 2025 (CVPR 2025 Highlight)

A RAW-image detection dataset and benchmark paired with knowledge-distilled RAW pretraining address diverse lighting and weather conditions.

arxiv

CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic Segmentation

Boyuan Sun, Yuqi Yang, Weifeng Yuan, Le Zhang, Ming-Ming Cheng, Qibin Hou#

IEEE Computer Vision and Pattern Recognition 2024 (CVPR 2024)

Pixel and region label propagation through feature correlations makes better use of unlabeled images for semi-supervised semantic segmentation.

arxiv

CamoFormer: Masked Separable Attention for Camouflaged Object Detection

Bowen Yin, Xuying Zhang, Qibin Hou, Bo-Yuan Sun, Deng-Ping Fan, & Luc Van Gool.

Arxiv preprint 2022

Masked separable attention distinguishes camouflaged objects from their surroundings and progressively refines segmentation through a top-down decoder.

Find me on WeChat with the ID sby123bb, or scan my QR code:

QR code