I am Shenzhi Wang (็ๆ ๆง in Chinese), a Ph.D. candidate at LEAP Lab in the Department of Automation at Tsinghua University. My research focuses on post-training for foundation models, with an emphasis on reinforcement learning.
I am currently with Moonshot AI, where my recent work includes contributions to Kimi K3, focusing on scalable agentic post-training. Previously, I served as a research intern on Alibaba Qwen’s Post-training Team, where I worked on reinforcement learning, scalable post-training, and multimodal reasoning. I led Beyond the 80/20 Rule and HopChain, with HopChain serving as one of Qwen3.5’s vision-language RLVR training tasks.
My research has received citations. The Flexibility Trap received the ๐ ICML 2026 Outstanding Paper Award (2 out of 23,918 submissions), while Beyond the 80/20 Rule and Absolute Zero rank among the top 5 and top 25 most-cited NeurIPS 2025 papers, respectively. My publications include 3 Oral and 2 Spotlight papers. I have also open-sourced Llama3-Chinese-Chat, which has accumulated 1M+ downloads and reached Hugging Face Trending #7, and Xwen-Chat, which surpassed the then state-of-the-art DeepSeek-V3 in chat performance.
Ph.D. in Artificial Intelligence, 2021โ2027
Department of Automation, Tsinghua University
B.Eng. in Computer Science and Technology, 2017โ2021
SHENYUAN Honors College, Beihang University
For a complete list of publications, please see my Google Scholar profile.
For research discussions and collaboration, please feel free to reach out by email.