Portrait
Yifu Luo
Ph.D. Student
Technical University of Munich
About Me

I am a Ph.D. student at Technical University of Munich, majoring in computer science and supervised by Prof. Dr. Gjergji Kasneci. Prior to that,I recieved both my bachelor and master from Tsinghua University, majoring in Artificial Intelligence.

My research interests focus on the post-training paradigm (mainly as RL and OPD), for both LLMs and diffusion models. I have several top-tier publications as well as frontier lab internship experiences in this area.

I am always actively looking for research internship opportunities. Feel free to reach out to me if there is a match!

I am also opened to collaboration, especially (1) if you can give me SSH access to a large GPU cluster, or (2) if you are seeking for some hands-on guidance. If you are interested, please drop me an email. I check it every day!

Research framework: post-training for generative models, including reinforcement learning and on-policy distillation for both LLMs and diffusion models. Curriculum Vitae
Education
  • Technical University of Munich
    Technical University of Munich
    Ph.D.
    Supervised by Prof. Dr. Gjergji Kasneci
    Computer Science
    2026 - 2029
  • Tsinghua University
    Tsinghua University
    Master
    Artificial Intelligence
    Supervised by Prof. Dr. Xueqian Wang
    2023 - 2026
  • Tsinghua University
    Tsinghua University
    Bachelor
    Electronic Engineering
    2018 - 2022
Experience
Honors & Awards
  • Outstanding Graduates, Tsinghua University
    2026
  • First Class Outstanding Student, Tsinghua University
    2025
News
2026
New homepage. In summary for previous experience, I got both of my master and bachelor from Tsinghua University. My research focuses on the RL post-training. I have several top-tier first-author papers at NeurIPS and ICML in this field, as well as several frontier labs internship experiences in this field, such as ByteDance Seed and Kling AI.
Jul 24
Selected Publications (view all )
Learning from the Self-future: On-policy Self-distillation for dLLMs
Learning from the Self-future: On-policy Self-distillation for dLLMs

Yifu Luo*, Zeyu Chen*, Haoyu Wang, Xinhao Hu, Yuxuan Zhang, Zhizhou Sha, Shiwei Liu (* equal contribution)

Preprint 2026

Propose the first OPSD approach for dLLMs. We reframe the OPSD formulation from prefix conditioning to the selfsuffix generation conditioning tailored specifically for dLLMs. We also shift dense supervision from the token-level to steplevel, aligning training with the iterative denoising nature of dLLMs.

Learning from the Self-future: On-policy Self-distillation for dLLMs

Yifu Luo*, Zeyu Chen*, Haoyu Wang, Xinhao Hu, Yuxuan Zhang, Zhizhou Sha, Shiwei Liu (* equal contribution)

Preprint 2026

Propose the first OPSD approach for dLLMs. We reframe the OPSD formulation from prefix conditioning to the selfsuffix generation conditioning tailored specifically for dLLMs. We also shift dense supervision from the token-level to steplevel, aligning training with the iterative denoising nature of dLLMs.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Yifu Luo*, Haoyuan Sun*, Xinhao Hu*, Penghui Du*, Keyu Fan, Bo Li, Sinan Du, Wan Xu, Zhiyu Chen, Bo Xia, Yongzhe Chang, Changqian Yu, Kun Gai, Tiantian Zhang, Xueqian Wang (* equal contribution)

Forty-Third International Conference on Machine Learning (ICML) 2026

Propose a chunk-level RL optimization for flow-matching text-to-image generation. We shift GRPO from the step-level to the chunk-level, aggregating consecutive steps into a coherent “chunk” to mitigate the sparse rewards bottleneck.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Yifu Luo*, Haoyuan Sun*, Xinhao Hu*, Penghui Du*, Keyu Fan, Bo Li, Sinan Du, Wan Xu, Zhiyu Chen, Bo Xia, Yongzhe Chang, Changqian Yu, Kun Gai, Tiantian Zhang, Xueqian Wang (* equal contribution)

Forty-Third International Conference on Machine Learning (ICML) 2026

Propose a chunk-level RL optimization for flow-matching text-to-image generation. We shift GRPO from the step-level to the chunk-level, aggregating consecutive steps into a coherent “chunk” to mitigate the sparse rewards bottleneck.

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

Yifu Luo*, Xinhao Hu*, Keyu Fan*, Haoyuan Sun*, Zeyu Chen, Bo Xia, Tiantian Zhang, Yongzhe Chang, Xueqian Wang (* equal contribution)

Advances in Neural Information Processing Systems 38 (NeurIPS) 2025 13 citations

Propose the first GRPO approach for masked generative models in text-to-image generation. We redefine the transition probability tailored specifically for masked generative models, and explore several useful strategies to further enhance our method.

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

Yifu Luo*, Xinhao Hu*, Keyu Fan*, Haoyuan Sun*, Zeyu Chen, Bo Xia, Tiantian Zhang, Yongzhe Chang, Xueqian Wang (* equal contribution)

Advances in Neural Information Processing Systems 38 (NeurIPS) 2025 13 citations

Propose the first GRPO approach for masked generative models in text-to-image generation. We redefine the transition probability tailored specifically for masked generative models, and explore several useful strategies to further enhance our method.

All publications