Kaiyuan Deng

Kaiyuan Deng

M.Eng. Student at UESTC | Embodied Intelligence | Multimodal Perception

I am Kaiyuan Deng, an M.Eng. student in Computer Technology at the University of Electronic Science and Technology of China (UESTC). My research focuses on embodied intelligence and multimodal perception, with particular interests in robotic manipulation and multimodal large language models.

My work spans egocentric visual reasoning, vision-language-action learning, 3D hand-trajectory prediction, high-value data selection, and embodied robotic systems. I enjoy turning research ideas into reliable systems through real-robot deployment, 3D perception, and simulation.

Email: mantou.cloud@gmail.com

Education

University of Electronic Science and Technology of China (UESTC)
Sep. 2024 - Jun. 2027 (expected)
  • Professional M.Eng. in Computer Technology, School of Computer Science and Engineering
  • Research: embodied intelligence and multimodal perception
  • Graduate Academic Scholarship
University of Electronic Science and Technology of China (UESTC)
Sep. 2020 - Jun. 2024
  • B.Eng. in Data Science and Big Data Technology, School of Computer Science and Engineering
  • GPA: 3.98/4.00
  • First-Class Academic Scholarship for two consecutive years

Experience

华为2012实验室
自动驾驶算法实习生
2026.05-Present

    Publications

    Journal Papers

    Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning

    Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning

    Shenshen Li, Kaiyuan Deng, Lei Wang, Hao Yang, Chong Peng, Peng Yan, Fumin Shen, Heng Tao Shen, Xing Xu

    IEEE Transactions on Multimedia (under review) · TMM, under review

    multimodal reasoning; high-value data selection; reinforcement learning; causal discrepancy

    Conference Papers

    Ego3S: Select, Strengthen, and Synchronize for Efficient Egocentric Reasoning

    Ego3S: Select, Strengthen, and Synchronize for Efficient Egocentric Reasoning

    Shenshen Li, Kaiyuan Deng, Ruohuai Xie, Xing Xu, Heng Tao Shen, Yazhou Yao, Fumin Shen

    International Conference on Machine Learning · ICML 2026

    egocentric reasoning; multimodal large language models; reinforcement learning; data selection

    MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training

    MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training

    Zhenhan Yin, Xuanhan Wang, Jiahao Jiang, Kaiyuan Deng, Pengqi Chen, Shuangle Li, Chong Liu, Xing Xu, Jingkuan Song, Lianli Gao, Heng Tao Shen

    IEEE/CVF Conference on Computer Vision and Pattern Recognition, Findings · CVPR 2026 Findings

    vision-language-action; mutual imitation; robot learning; cross-embodiment learning

    Visual Causal Intervention for 3D Egocentric Multimodal Hand Trajectory Forecasting

    Visual Causal Intervention for 3D Egocentric Multimodal Hand Trajectory Forecasting

    Kaiyuan Deng, Xun Jiang, Zheng Wang, Fumin Shen, Xing Xu

    IEEE International Conference on Multimedia and Expo · ICME 2026

    3D hand trajectory forecasting; visual causal intervention; egocentric vision; multimodal learning

    Awards