I am Shenzhi Wang (็ๆ ๆง in Chinese), a Ph.D. candidate at LEAP Lab in the Department of Automation at Tsinghua University. My research focuses on post-training for foundation models, with an emphasis on reinforcement learning.
I am currently with Moonshot AI, where my recent work includes contributions to Kimi K3, focusing on scalable agentic post-training. Previously, I served as a research intern on Alibaba Qwen’s Post-training Team, where I worked on reinforcement learning, scalable post-training, and multimodal reasoning. I developed Beyond the 80/20 Rule and HopChain, with HopChain serving as one of Qwen3.5’s vision-language RLVR training tasks.
My research has received citations. The Flexibility Trap received the ๐ ICML 2026 Outstanding Paper Award (2 out of 23,918 submissions), while Beyond the 80/20 Rule and Absolute Zero rank among the top 5 and top 25 most-cited NeurIPS 2025 papers, respectively. My publications include 3 Oral and 2 Spotlight papers. I have also open-sourced Llama3-Chinese-Chat, which has accumulated 1M+ downloads and reached Hugging Face Trending #7, and Xwen-Chat, which surpassed the then state-of-the-art DeepSeek-V3 in chat performance.
Ph.D. in Artificial Intelligence, 2021โ2027 (expected)
Department of Automation, Tsinghua University
B.Eng. in Computer Science and Technology, 2017โ2021
SHENYUAN Honors College, Beihang University
For a complete list of publications, please see my Google Scholar profile.



Synthesizes instance-grounded multi-hop visual reasoning data for RLVR, improving generalization across 20 of 24 benchmarks on Qwen3.5-35B-A3B and Qwen3.5-397B-A17B.

Identifies arbitrary-order generation as a bottleneck for dLLM reasoning and introduces JustGRPO, an ICML 2026 Outstanding Paper Award-winning method that preserves parallel decoding while achieving state-of-the-art results.

Shows that high-entropy minority tokens drive RLVR and develops selective policy-gradient updates that match or surpass full-gradient training; ranked among the five most-cited NeurIPS 2025 papers.

Post-trains Qwen2.5 into open 7B and 72B chat models, with Xwen-72B-Chat surpassing the then state-of-the-art DeepSeek-V3 in chat performance.

Delivers bilingual Chinese-English Llama 3 chat models with stronger role-play, tool-use, and math capabilities.

Learns a family of policies and state-adaptively selects improvement-constraint balances, consistently improving offline-to-online reinforcement learning across standard baselines.
For research discussions and collaboration, please feel free to reach out by email.