Biography

I am Shenzhi Wang (็Ž‹ๆ…Žๆ‰ง in Chinese), a Ph.D. candidate at LEAP Lab in the Department of Automation at Tsinghua University. My research focuses on post-training for foundation models, with an emphasis on reinforcement learning.

I am currently with Moonshot AI, where my recent work includes contributions to Kimi K3, focusing on scalable agentic post-training. Previously, I served as a research intern on Alibaba Qwen’s Post-training Team, where I worked on reinforcement learning, scalable post-training, and multimodal reasoning. I developed Beyond the 80/20 Rule and HopChain, with HopChain serving as one of Qwen3.5’s vision-language RLVR training tasks.

My research has received Google Scholar citation count citations. The Flexibility Trap received the ๐Ÿ† ICML 2026 Outstanding Paper Award (2 out of 23,918 submissions), while Beyond the 80/20 Rule and Absolute Zero rank among the top 5 and top 25 most-cited NeurIPS 2025 papers, respectively. My publications include 3 Oral and 2 Spotlight papers. I have also open-sourced Llama3-Chinese-Chat, which has accumulated 1M+ downloads and reached Hugging Face Trending #7, and Xwen-Chat, which surpassed the then state-of-the-art DeepSeek-V3 in chat performance.

Interests
  • Foundation Model Post-Training
  • Reinforcement Learning
  • LLMs, VLMs, and DLMs
Education
  • Ph.D. in Artificial Intelligence, 2021โ€“2027

    Department of Automation, Tsinghua University

  • B.Eng. in Computer Science and Technology, 2017โ€“2021

    SHENYUAN Honors College, Beihang University

News

Academic Service & Public Impact

Contact

For research discussions and collaboration, please feel free to reach out by email.

  • wangshenzhi99@gmail.com