Xiangxuan Ren任相璇
A little curiosity goes a little way.Ph.D. candidate at Institute of AI, Shanghai Jiao Tong University
Perception reveals structure across space and time.
Geometry, semantics, and identity connect individual observations across space and time.
Spatial anchors and a unified vocabulary let language describe visual locations and reason over them.
A spatial axis extends reasoning beyond semantics, allowing semantics and localization to evolve together.
Spatial structure connects what is visible, what must be inferred, and how an agent can act.
Imagined futures become useful when models can reason over them to guide their next action.
About me

I am a Ph.D. candidate in Computer Science and Technology at the Institute of AI, Shanghai Jiao Tong University, advised by Prof. Chao Ma. I began my doctoral studies in April 2023 and expect to graduate in June 2027.
My research centers on multimodal large language models, with a focus on how they build representations of the visual world and reason over them. I study how intelligence emerges through increasingly capable ways of representing, reasoning about, and interacting with the world, towards more coherent and generalizable world understanding.
I received my master’s degree from Shanghai Jiao Tong University’s Artificial Intelligence and Robotics Center (September 2020–March 2023), advised by Prof. Weidong Zhang. Before that, I earned my bachelor’s degree in Automation from the University of Electronic Science and Technology of China (September 2016–June 2020).
News
- : AgentWorld, our paper on vision-language-action models, was accepted to NeurIPS 2026.
- : Two papers accepted to ECCV 2026: PixelPilot on scalable vision-language-action learning and FocusGS on sparse-view 3D reconstruction.
- : Our occupancy world model OccTrans was accepted to ICME 2026 (Oral).
- : GETok, our unified spatial vocabulary for multimodal grounding, was accepted to CVPR 2026; VLA-World to CVPR 2026 Findings.
- : LiteFusion is online: enhancing camera-based 3D detection with LiDAR geometry and minimal adaptation.
- : OccGen on generative occupancy prediction and VEON on open-vocabulary 3D perception were accepted to ECCV 2024.
- : SparseOcc and Un-Track were accepted to CVPR 2024, advancing sparse 3D perception and unified multimodal tracking.
Selected Publications
Full publication listInternships
–
Research Intern at Huawei, Noah’s Ark Lab.
Focused on 3D perception and multimodal large language models.
Honors & Awards
- Outstanding Graduate, Shanghai Jiao Tong University, .
- COSCO SHIPPING Scholarship (First Class), Shanghai Jiao Tong University (ranked 1st in the department), .
- Outstanding Student, University of Electronic Science and Technology of China, .
- Top 100 Outstanding Youth Volunteers, University of Electronic Science and Technology of China, .
- Outstanding Student Scholarship, University of Electronic Science and Technology of China (top 10%), .
- Outstanding Youth League Student Leader, University of Electronic Science and Technology of China, .
Academic Service
Reviewer for CVPR, ICLR, ICCV, NeurIPS, ECCV, AAAI, IEEE TMM, IEEE TIV, IEEE TVT.
Partners & Collaborators
I have collaborated closely with Dr. Zhongdao Wang · Dr. Guoqing Wang · Dr. Pin Tang · Dr. Zongwei Wu · Dr. Jilai Zheng.
I am sincerely grateful to my advisor and collaborators for their generous guidance and support, encouragement, trust, and companionship through the challenges of research. Their perspectives and support have shaped my growth both as a researcher and as a person. I cherish the experience of learning from one another and growing together.
