Yu Chen

AI researcher & builder, working on multimodal agents.

Incoming PhD Student · University of Chinese Academy of Sciences B.S. · Southeast University
Portrait of Yu Chen at golden hour, overlooking a river

I’m an incoming PhD student focusing on multimodal agent post-training, with the goal of enabling agents to perceive the world and make complex decisions more like humans.

My perspective is shaped by both research training and hands-on industry experience. I believe impactful research should not only push the boundaries of knowledge, but also solve real problems with scalable, simple, and robust methods.

Education

Ph.D.
University of Chinese Academy of Sciences (UCAS) · Beijing
2026 — Present
B.S. in Artificial Intelligence
Southeast University · Nanjing
2022 — 2026

Research

2025 — Present

Video Understanding & Multimodal Agent Post-Training

Video understanding, long audio–video reasoning, and multimodal agent post-training — teaching agents to hold hours of context and act on it.

2024 — 2025

Multimodal Representation Learning

Fine-grained vision–language representation and pixel-text alignment, down to any granularity.

2023 — 2024

EEG & Brain–Computer Interfaces

EEG-based BCI and mental-health modeling — reading signal directly off the brain.

Publications

2026
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
arXiv
Yu Chen*, Caorui Li*, Ziyu Xiong, Yidong Wang, Mingqi Gao, Shuman Liu, Biao Liu, Chunfeng Yang, Anxiang Zeng, Haibo Zhang, Chaofan Chen
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
ICLR
Caorui Li*, Yu Chen*, Yiyan Ji*, Jin Xu, Zhenyu Cui, Shihao Li, Yuanxing Zhang, Wentao Wang, Zhenghao Song, Dingling Zhang, et al.
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
ICML
Yicheng Xiao*, Yu Chen*, Haoxuan Ma*, Jiale Hong, Caorui Li, Lingxiang Wu, Haiyun Guo, Jinqiao Wang
Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
arXiv
Mingqi Gao, Hongyuan Dong, Yifei Chen, Zhisheng Zhong, Zheng Ruan, Wenjin Hou, Yu Chen, Han Hu, Yansong Tang
2025
FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network
NeurIPS
Shuo Xu, Yu Chen, Shuxia Lin, Xin Geng, Xu Yang
STGE-Former: Spatial–Temporal Graph-Enhanced Transformer for EEG-Based Major Depressive Disorder Detection
ICASSP
Yu Chen, Chunfeng Yang
MLLM-CBench: A Comprehensive Benchmark for Continual Instruction Tuning of Multimodal LLMs with Chain-of-Thought Reasoning Analysis
arXiv
Haiyun Guo, ZhiYan Hou, Yu Chen, Jinghan Sun, Yadu Zhou, Yuzhe Guo, Shujing Zhu, Kuan Zhu, Jinqiao Wang
A Multi-scale Representation Learning Framework for Long-Term Time Series Forecasting
arXiv
Boshi Gao, Qingjian Ni, Fanbo Ju, Yu Chen, Ziqi Zhao

* Equal contribution

Experience

Shopee · LLM Team, Pre-training

Large Language Model Algorithm Intern
Multimodal agentic post-training within the pre-training LLM team.
Sep 2025 — Apr 2026

Pinch · YC-backed AI startup

Machine Learning Engineer Intern
Voice cloning and text-to-speech for AI simultaneous interpretation — preserving speaker identity while keeping translation fast and accurate.
Jan 2025 — Jun 2025

Huawei · ICT

AI Engineer Intern
Multimodal large-model training and inference optimization — accelerating and migrating workloads from GPUs to Ascend NPUs.
Jun 2024 — Sep 2024