About
Hi there!
I am a Master's student in Electrical Engineering at Stanford University. I graduated with my Bachelor's degree in Computer Science from Tongji University in June 2026.
I conducted research at UC San Diego MLPC Lab, advised by Prof. Zhuowen Tu. I was also fortunate to work with Prof. Jiaqi Wang during my internship at JD Explore Academy, part of JD.com. My research interests include agent harnesses and 3D/4D spatial intelligence.
News
- One paper has been accepted to NeurIPS 2026. See you in Sydney!
- I joined JD Explore Academy as a research intern
- One paper has been accepted to CVPR 2026. See you in Denver!
- One paper has been accepted to ICASSP 2026. See you in Barcelona!
- I joined UC San Diego MLPC Lab as a research intern
- One paper has been accepted to ICCV 2025. See you in Hawaii!
Publications
* equal contribution · click a figure to pause it
Preprint
Paper
Project
Code
NeurIPS 2026
Towards Realistic Conversational Multimodal Instruction Following
Paper
Project
Code
Demos
Things you can try in the browser
Experience
-
JD Explore Academy
May 2026 – Oct 2026Research Intern
-
UC San Diego
Jul 2025 – Dec 2025Research Intern
-
Zhejiang University
Dec 2024 – Mar 2025Research Intern
-
Tongji University
Nov 2024 – Aug 2025Research Intern
Before research · I miss those happy days
ASCE Concrete Canoe Competition
Hull designer · May 2022 – Apr 2024 · Hosted by ASCE in Sacramento, CA
Honors
- National Scholarship · top 0.2% nationwide, the highest scholarship in China
- Interdisciplinary Contest in Modeling (COMAP), Finalist Prize · top 1.8% worldwide
- Outstanding Undergraduate Thesis Award · much to my surprise; graduated just fine anyway
- ASCE Concrete Canoe Competition, 2nd Place in California Section, 2024 · almost beat UC Berkeley, lol
- School Sports Meet, Silver Medal in 4×100 m Relay and Bronze Medal in 4×400 m Relay · I used to be fast
- Up-and-Coming Dessert Baker · unanimously and highly acclaimed by friends
PixARMesh: Auto-Regressive Mesh-Native Single-View Scene Reconstruction
FOLK: Fast Open-Vocabulary 3D Instance Segmentation via Label-guided Knowledge Distillation
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction