Ximeng Sun

👋 Hello, I am an Applied Research Scientist on the AMD GenAI team, where I work on multimodal foundation models. My research spans multimodal large language models (VLMs / MLLMs), data-centric AI, multimodal reasoning, and efficient learning.

At AMD, I work hands-on across the full foundation-model stack — large-scale data curation and data engines, pretraining, supervised fine-tuning, multimodal reasoning, and rigorous evaluation — with distributed training on AMD Instinct GPUs. I am a main contributor to InstellaVL-1B, the first fully-open multimodal LLM trained end-to-end on AMD GPUs, released with open weights, code, and a technical blog. I also contribute across the broader AMD Instella family of open models, including the Instella open language models and the Instella-T2I text-to-image model, spanning multimodal understanding, generation, and reasoning.

I earned my Ph.D. in Computer Science from Boston University (2019–2024), advised by Prof. Kate Saenko, with research internships at Meta AI, Google Cloud, and IBM Research. Earlier, I received my M.S. from the University of Michigan, Ann Arbor and my B.Eng. from Beijing University of Posts and Telecommunications.

I'm always open to collaboration! If you're interested in an internship with me, or would like to reach out for academic collaboration, feel free to get in touch. 📬


💼 Experience

  • Applied Research Scientist, AMD GenAI — Multimodal Foundation Models. San Jose, CA. Jun 2024 – Present
  • Ph.D. / Research Fellow, Boston University. Advised by Prof. Kate Saenko. 2019 – 2024
  • Research Scientist Intern, Meta AI. Host: Xide Xia. 2022
  • Research Intern, Google Cloud. Host: Clayton Mellina. 2021
  • Research Intern, IBM Research. Mentor: Rameswar Panda; Manager: Rogerio Feris. 2019, 2020

🚀 Selected Open Model Releases

  • InstellaVL-1B — the first fully-open multimodal LLM trained end-to-end on AMD GPUs (main contributor). Released with weights, code, and a technical blog.

    model / code / blog

  • Instella — a family of fully-open language models with stellar performance, trained on AMD GPUs (contributor).

    paper / code / blog

  • Instella-T2I — an open text-to-image model built on a compact 1D discrete latent space (contributor).

    paper / code / model


📚 Publications

Author of 20+ papers with 2,100+ citations. A curated list is below; see my Google Scholar for the full list.

📄 Conference & Journal Publications

  • Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, Zicheng Liu. "Latent Visual Reasoning". ICLR 2026.
  • paper

  • Jingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun, Ze Wang, Xiaodong Yu, Jiebo Luo, Zicheng Liu, Emad Barsoum. "VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking". CVPR 2026.
  • paper

  • Shijia Yang, Yunong Liu, Bohan Zhai, Ximeng Sun, Zicheng Liu, Emad Barsoum, Manling Li, Chenfeng Xu. "CaptionQA: Is Your Caption as Useful as the Image Itself?". CVPR 2026.
  • paper

  • Xingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu, Ze Wang, Ximeng Sun, Jialian Wu, Alan Yuille, Emad Barsoum, Zicheng Liu. "XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models". ICLR 2026.
  • paper

  • Yujian Liu, Ze Wang, Hao Chen, Ximeng Sun, Xiaodong Yu, Jialian Wu, Jiang Liu, Emad Barsoum, Zicheng Liu, Shiyu Chang. "Learning from Online Videos at Inference Time for Computer-Use Agents". TMLR 2026.
  • paper

  • Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Xiaodong Yu, Hao Chen, Jiebo Luo, Zicheng Liu, Emad Barsoum. "Unleashing Hour-Scale Video Training for Long Video-Language Understanding". NeurIPS 2025 (Spotlight).
  • paper

  • Yufan Zhuang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Jingbo Shang, Zicheng Liu, Emad Barsoum. "Self-Taught Agentic Long Context Understanding". ACL 2025.
  • paper

  • Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, Emad Barsoum. "Agent Laboratory: Using LLM Agents as Research Assistants". EMNLP 2025 (Findings).
  • paper

  • Hao Chen, Ze Wang, Xiang Li, Ximeng Sun, Fangyi Chen, Jiang Liu, Jindong Wang, Bhiksha Raj, Zicheng Liu, Emad Barsoum. "SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer". CVPR 2025.
  • paper

  • Reuben Tan, Ximeng Sun, Ping Hu, Jui-hsien Wang, Hanieh Deilamsalehy, Bryan A. Plummer, Bryan Russell, Kate Saenko. "Koala: Key Frame-Conditioned Long Video-LLM". CVPR 2024.
  • paper

  • Ping Hu, Ximeng Sun, Stan Sclaroff, Kate Saenko. "DualCoOp++: Fast and Effective Adaptation to Multi-Label Recognition with Limited Annotations". TPAMI 2024.
  • paper

  • Ximeng Sun, Rameswar Panda, Chun-Fu Chen, Naigang Wang, Bowen Pan, Aude Oliva, Rogerio Feris, Kate Saenko. "All at Once Network Quantization via Collaborative Knowledge Transfer". WACV 2024.
  • Ximeng Sun, Pengchuan Zhang, Peizhao Zhang, Hardik Shah, Kate Saenko, Xide Xia. "DIME-FM: DIstilling Multimodal and Efficient Foundation Models". ICCV 2023.
  • paper / code

  • Ximeng Sun, Ping Hu, Kate Saenko. "DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations". NeurIPS 2022.
  • paper / code

  • Ximeng Sun, Rameswar Panda, Chun-Fu Chen, Aude Oliva, Rogerio Feris, Kate Saenko. "Dynamic Network Quantization for Efficient Video Inference". ICCV 2021.
  • paper / code

  • Rameswar Panda, Chun-Fu Chen, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, Rogerio Feris. "AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition". ICCV 2021.
  • paper / code

  • Ximeng Sun, Rameswar Panda, Rogerio Feris, Kate Saenko. "AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning". NeurIPS 2020.
  • paper / code

  • Ximeng Sun, Huijuan Xu, Kate Saenko. "TwoStreamVAN: Improving Motion Modeling in Video Generation". WACV 2020.
  • paper / code

  • Xingchao Peng, Zijun Huang, Ximeng Sun, Kate Saenko. "Domain Agnostic Learning with Disentangled Representations". ICML 2019.
  • paper / code

  • Ximeng Sun, Ryan Szeto, Jason Corso. "A Temporally-Aware Interpolation Network for Video Frame Inpainting". ACCV 2018 (extended in TPAMI 2019).
  • paper / demo / code

📝 Preprints

  • Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang, Emad Barsoum, Zicheng Liu, Sandeep Chinchali, Ufuk Topcu. "VEGAS: Human-Aligned Video Caption Evaluation via Gaze". 2026.
  • paper

  • Yihao Liang, Ze Wang, Hao Chen, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Emad Barsoum, Zicheng Liu, Niraj K. Jha. "CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models". 2026.
  • paper

  • Chao Huang, Zeliang Zhang, Jiang Liu, Ximeng Sun, Jialian Wu, Xiaodong Yu, Ze Wang, Chenliang Xu, Emad Barsoum, Zicheng Liu. "Directional Reasoning Injection for Fine-Tuning MLLMs". 2025.
  • paper

  • Yuxiang Guo, Jiang Liu, Ze Wang, Hao Chen, Ximeng Sun, Yang Zhao, Jialian Wu, Xiaodong Yu, Zicheng Liu, Emad Barsoum. "ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning". 2025.
  • paper

  • Ze Wang, Hao Chen, Benran Hu, Jiang Liu, Ximeng Sun, Jialian Wu, Yusheng Su, Xiaodong Yu, Emad Barsoum, Zicheng Liu. "Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation". 2025.
  • paper

  • Xingrui Wang, Jiang Liu, Ze Wang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Yusheng Su, Alan Yuille, Zicheng Liu, Emad Barsoum. "KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation". 2025.
  • paper

  • Aimon Rahman, Jiang Liu, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Yusheng Su, Vishal M. Patel, Zicheng Liu, Emad Barsoum. "MOVi: Training-free Text-conditioned Multi-Object Video Generation". 2025.
  • paper

🔖 Patents

  • Rameswar Panda, Ximeng Sun, Richard Chen, Rogerio Schmidt Feris, Ekaterina Saenko. "Dynamic network quantization for efficient video inference". US Patent App. 17/566,782.

🤝 Collaborators

At AMD, I closely collaborate with Zicheng Liu, Jiang Liu, Jialian Wu, Ze Wang, Xiaodong Yu, Hao Chen, Yusheng Su, Shijia Yang, Bangzheng Li, and Shenghui Chen.

I am also grateful to my Ph.D. advisor Kate Saenko and to earlier mentors and collaborators, including Rogerio Feris, Rameswar Panda, Xide Xia, Pengchuan Zhang, Peizhao Zhang, Clayton Mellina, Xiao Bian, and Kihyuk Sohn.


✉️ Contact

Email: sunxmgreatwork [AT] gmail [dot] com