Email: jiangyi0425 [at] gmail.com · jiangyi.enjoy [at] bytedance.com

Biography

I am a Research Lead at ByteDance Seed, where I work on generative foundation models.

I received my master's degree from the Department of Computer Science and Engineering at Zhejiang University.

Our work on Visual Autoregressive Modeling (VAR) received a NeurIPS 2024 Best Paper Award.

Research Interests

Visual foundation models, generative pretraining, and large language models.

Unified multimodal generation and understanding for open-world interaction.

Large-scale multimodal pretraining and alignment.

Invited Talks

"Elucidating the Design Space of Visual Autoregressive Models and Image Tokenizers", Tutorial: Autoregressive Models Beyond Language, NeurIPS, 2025.

"Towards Autoregressive Modeling for Scalable and Versatile Visual Generation", Workshop: What Makes a Good Video: Next Practices in Video Generation and Evaluation, NeurIPS, 2025.

"Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction", invited talk at the BAAI Conference, 2024.

"Spark from Large Language Models: Pretraining, Open-World, Generalized Vision Models", invited talk at IDEA, 2024.

Highlights

  • Visual AutoRegressive modeling: new visual generation Framework elevates Autoregressive models beyond diffusion, indicate scaling law in image generation.

  • Waver: next-generation foundation model for unified image & video generation, built on rectified flow Transformers and engineered for industry-grade performance.

  • Unitok: unified tokenizer for visual generation & understanding, can be seamlessly integrated into MLLMs to unlock visual generation & understanding capability.

  • ByteTrack ranks 1th of the most influential papers in ECCV 2022. Code is available on github with 5.1k stars.

  • Sparse R-CNN accepted by CVPR'21. Sparse R-CNN is integrated into several famous frameworks(Detectron2, MMDetection, PaddlePaddle).

Selected Publications [Google Scholar]

(* Equal contribution; † Project lead; ‡ Corresponding author)

Conference Papers, Journal Articles, and Preprints

Honors and Awards

Competitions

  • Winner of the CVPR 2022 Large-Scale Video Object Segmentation Challenge: Video Instance Segmentation

  • Runner-up in CVPR 2021 FGVC8 iNaturalist Challenge

  • Runner-up in ICCV 2019 WIDER Face and Person Challenge: Face Detection

  • Kaggle Competitions Master, 2018

Professional Activities