JS Jun Sun

CV

Jun Sun · luciojunsun@gmail.com · github.com/LucioSunj · WeChat qw2667180048 · Download PDF

Education

BSc in Information and Computing Science

Xi'an Jiaotong-Liverpool University · Suzhou, China

  • GPA 3.94 / 4.00  ·  Major GPA 4.00 / 4.00 (Rank 1)
  • IELTS 7.5 (Listening 8.0, Reading 9.0)
2022.09 – 2026.07

Publications

TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation

J. Sun, B. Yang, J. Zhang, N. Ma, C. Wu, S. Zhang, Y. Huang, Q. Wang, S. Liang, Y. Chen

IROS 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems

A training-free retrofit that gives frozen VLA policies history awareness by caching intermediate-layer prefix K/V states — no parameter change, no longer inputs. +4.0 points on LIBERO-Long with π0.5, and up to +14.2% full-task success on a real Realman RM-65B.

2026

MiTPose: Multi-Granularity Guided Vision Transformer for Human Pose Estimation

Y. Wu, Q. Gao, Y. Liu, J. Sun, Z. Li, Y. Jin, Y. Yue, X. Zhu

INDIN 2025 IEEE International Conference on Industrial Informatics

2025

Research experience

World Action Models and Reinforcement Learning for Robot Manipulation Current

Research Assistant · School of Robotics and Automation, Nanjing University · Advisor: Prof. Shangke Lyu

  • Research on world action models and reinforcement learning for robot manipulation: how a policy represents future state, and how it allocates that prediction over a long horizon.
2026.03 – Present

TempoFit: Training-Free Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation

Research Assistant · Advisor: Prof. Yaran Chen

Memoryless VLA policies struggle over long horizons under occlusion, state aliasing and subtle post-action changes — which calls for efficient temporal reasoning without retraining or architectural surgery.

  • Temporal memory design: proposed TempoFit, a training-free retrofit that equips frozen VLA policies with history awareness by caching intermediate-layer prefix K/V states, leaving parameters and input length untouched.
  • Retrieval & fusion: developed K-to-K retrieval with a frame-gap temporal bias and norm-preserving residual loading, suppressing stale context and minimising distribution shift under frozen weights.
  • Benchmarks: LIBERO-Long success from 92.6% to 96.6% on π0.5 and 90.8% to 94.4% on QwenGR00T, with consistent gains on CALVIN.
  • Real-robot validation: validated on a Realman RM-65B across three multi-stage tasks, raising full-task success by up to 14.2% with minor latency and memory overhead.
2025.11 – 2026.03

Multimodal Embodied Intelligence with the Octo Framework and 3D Depth Perception

Research Assistant · Advisor: Prof. Yaran Chen

Language-guided manipulation is limited by RGB-only geometric ambiguity, occlusion and sim-to-real shift; robust VLA needs explicit 3D grounding for stable grasping decisions.

  • Depth-aware fusion: added a monocular-depth stream for multi-scale geometric features, aligned to Octo embeddings and injected via gated cross-attention with FiLM modulation — without modifying the core Octo VLA architecture.
  • Dataset construction: extended Bridge by generating monocular depth from RGB sequences into an RGB–Depth–Language corpus; collected teleoperated RGB-D real-robot data for fine-tuning and sim-to-real alignment.
  • Training & evaluation: behaviour cloning plus RL fine-tuning lifted unseen-object grasping success from ~77% to ~89%; robustness verified under adversarial occlusion and depth noise.
  • Contribution: benchmarked against OpenVLA and other baselines to isolate depth gains; co-authored technical reports and led weekly reviews on architecture and pipeline reliability.
2025.01 – 2025.06

Indoor 3D Reconstruction and Novel View Synthesis with UDF-Guided 3DGS

Research Assistant · Advisor: Prof. Yong Yue

In low-texture indoor scenes, appearance-driven 3D reconstruction suffers from geometric ambiguity and unstable optimisation, limiting surface fidelity and novel-view quality.

  • Modelling: built a NeRF-style UDF (unsigned distance + normals) coupled to a 3DGS renderer, optimised with joint feature–geometry losses and normal/depth-consistency regularisation for low-texture stability.
  • Optimisation & pruning: devised multi-stage UDF-guided optimisation — coarse geometry first, then refining Gaussian positions, scales and opacities — with voxel-based partitioning for efficient clone/prune, concentrating primitives near the zero-level set.
  • Evaluation: processed ScanNet++ with COLMAP calibration and point-cloud initialisation; evaluated with Chamfer Distance and PSNR, outperforming strong baselines.
2024.05 – 2024.12

Where I've worked

Algorithm R&D Intern

Westlake Robotics · Hangzhou, China

Deformable-object manipulation is hard because object state changes continuously, observations are partial and occluded, and small perception errors compound into unstable long-horizon action.

  • Task design & evaluation: designed real-robot protocols for folding, unfolding and stacking, with quantitative metrics for perception–action decision quality.
  • Real-robot data & training: led real-robot data collection and processing; fine-tuned on physical robot data and ran large-scale pretraining across tasks and embodiments to improve transferability.
  • End-to-end pipeline: built the annotation → conversion → training → online inference pipeline; improved deployability with distributed training and inference optimisation.
  • System integration: delivered a dexterous-hand towel-folding demo and led dual-arm long-horizon garment folding — the first team in China (after Figure) to demonstrate real-robot dexterous towel folding on a physical platform.
2025.06 – 2025.10

Awards & scholarships

  • 2025/26University Academic Excellence Award, XJTLU (Top 5%)
  • 2024/25University Academic Excellence Award, XJTLU (Top 5%)
  • 2023/24University Academic Excellence Award, XJTLU (Top 5%)

Skills

  • Programming & MLPython, Java, PyTorch, JAX
  • Embodied AI & RoboticsVision-Language-Action models (Octo, OpenVLA, π0), world action models, reinforcement learning
  • MathematicsCalculus, linear algebra, probability theory
  • LanguagesChinese (native), English (fluent, IELTS 7.5), Spanish (fluent)