Zhuojiang Cai

Research Assistant, PhD Student
Room: 274
Office Hours: by Appointment
E-Mail: cai.zhuojiang@tum.de
Please visit my personal homepage, LinkedIn, and Github for more details.
Open Research Topics
🤖 Motivated Master's and undergraduate students who are interested in these topics are welcome to reach out. I am always happy to discuss research ideas and collaborate on projects targeting top-tier computer vision and machine learning conferences.
1. Feed-Forward Generative 3D Scene Reconstruction
Towards accurate and scalable monocular 3D scene reconstruction with feed-forward generative models, focusing on geometry, camera pose estimation, and large-scale scene understanding.
2. Multi-View Human Foundation Models
Building geometry-aware foundation models for multi-view human understanding, including pose, gaze, interaction, and 3D reasoning.
3. Unified 3D Gaze and Attention Understanding
Learning a unified representation for 3D gaze estimation, point-of-gaze (PoG), gaze target prediction, and attention reasoning.
4. Foundation Models for Joint Gaze and Scene Understanding
Leveraging vision-language models and visual foundation models for joint reasoning over human attention, semantics, and 3D scenes.
Career Development
2018–2022: B.S. in Information and Computing Science, École Centrale de Pékin, Beihang University, China
2022–2025: M.Eng. in Electronic Information, École Centrale de Pékin, Beihang University, China
2025–present: Ph.D. Student, Technical University of Munich, Germany
Research Interests
Human & Scene Understanding
3D Reconstruction & Generation
Gaze Estimation
Natural Interaction