I’m an incoming PhD student focusing on multimodal agent post-training, with the goal of enabling agents to perceive the world and make complex decisions more like humans.
My perspective is shaped by both research training and hands-on industry experience. I believe impactful research should not only push the boundaries of knowledge, but also solve real problems with scalable, simple, and robust methods.
Education
Research
Video Understanding & Multimodal Agent Post-Training
Video understanding, long audio–video reasoning, and multimodal agent post-training — teaching agents to hold hours of context and act on it.
Multimodal Representation Learning
Fine-grained vision–language representation and pixel-text alignment, down to any granularity.
EEG & Brain–Computer Interfaces
EEG-based BCI and mental-health modeling — reading signal directly off the brain.
Publications
* Equal contribution
Experience

Shopee · LLM Team, Pre-training

Pinch · YC-backed AI startup
