Email: jiangyi0425 [at] gmail.com · jiangyi.enjoy [at] bytedance.com

Biography

I am a Research Lead at ByteDance Seed, where I work on generative foundation models.

I received my master's degree from the Department of Computer Science and Engineering at Zhejiang University.

Our work on Visual Autoregressive Modeling (VAR) received a NeurIPS 2024 Best Paper Award.

Research Interests

Visual foundation models, generative pretraining, and large language models.

Unified multimodal generation and understanding for open-world interaction.

Large-scale multimodal pretraining and alignment.

Invited Talks

"Elucidating the Design Space of Visual Autoregressive Models and Image Tokenizers", Tutorial: Autoregressive Models Beyond Language, NeurIPS, 2025.

"Towards Autoregressive Modeling for Scalable and Versatile Visual Generation", Workshop: What Makes a Good Video: Next Practices in Video Generation and Evaluation, NeurIPS, 2025.

"Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction", invited talk at the BAAI Conference, 2024.

"Spark from Large Language Models: Pretraining, Open-World, Generalized Vision Models", invited talk at IDEA, 2024.

Highlights

  • Visual Autoregressive Modeling (VAR): an image generation framework based on next-scale prediction that demonstrates scaling laws and outperforms diffusion transformers on the ImageNet benchmarks evaluated in the paper.

  • Waver: a foundation model for unified image and video generation, supporting text-to-image, text-to-video, and image-to-video generation.

  • Liquid: a scalable autoregressive multimodal model with a shared vocabulary for images and text, enabling visual understanding and generation.

  • UniTok: a unified tokenizer for visual generation and understanding that integrates into multimodal large language models (MLLMs) to enable visual generation while preserving understanding capabilities.

  • ByteTrack ranked 1st among ECCV 2022 papers in Paper Digest's influence ranking. Code is available on GitHub.

  • Sparse R-CNN was accepted at CVPR 2021 and has implementations in widely used frameworks, including Detectron2, MMDetection, and PaddlePaddle.

Selected Publications [Google Scholar]

(* Equal contribution; Project lead; Corresponding author)

Conference Papers, Journal Articles, and Preprints

Honors and Awards

Competitions

  • Winner of the CVPR 2022 Large-Scale Video Object Segmentation Challenge: Video Instance Segmentation

  • Runner-up in CVPR 2021 FGVC8 iNaturalist Challenge

  • Runner-up in ICCV 2019 WIDER Face and Person Challenge: Face Detection

  • Kaggle Competitions Master, 2018

Professional Activities