I am a research fellow at the Department of Mathematics, National University of Singapore. Previously, I obtained my PhD at NUS and I was privileged to be supervised by Prof. Vincent Y. F. Tan. Before that, I received my BS in Mathematics from Beijing Normal University, advised by Prof. Huajie Chen and Prof. Shihua Zhang.
Research
I am broadly interested in both the theoretical limits and empirical applications of reinforcement learning and online learning (e.g., multi-armed bandits), with a current focus on designing practical algorithms with provable guarantees for LLMs. Feel free to reach out if you share similar interests!
Practical Reinforcement Learning
- Designing practical and effective reinforcement learning algorithms to improve decision-making during model training and deployment.
Reinforcement Learning and Bandit Algorithms
- Characterizing the fundamental limits of decision-making under practical modeling assumptions, subject to realistic constraints (e.g., efficiency, nonstationarity, risk requirements etc.).
- Designing algorithms with provable guarantees that approach these fundamental performance limits.
News
- [2026-09] “SP²ec: Adaptive Self-Speculative Decoding for Vision-Language Models” accepted to NeurIPS 2026. We identify the single-peak structure in self-speculative decoding and propose a training-free adaptive algorithm to further boost the inference speed. This cannot be done without my amazing collaborators!
- [2026-09] “Rethinking “RL Generalizes, SFT Memorizes”: The Role of SFT Data” accepted to NeurIPS 2026. We revisit the generalization ability of SFT and RL training from the data perspective, including data coverage and data scale. Great thanks to the wonderful collaborators!
- [2026-07] “On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits” accepted to EWRL 2026. We study a variant of the regret minimization problem in stochastic MAB, where the learner is granted a free exploration time period before regret starts to accumulate. Joint work with Zixin and Vincent.
- [2026-06] Posted “Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition”. We study fixed-budget BAI in the Bayesian setting, allowing the learner to abstain at the end of the sampling process. This abstention option induces a fundamental transition in the error probability from polynomial to exponential decay. Joint work with Yuqi and Vincent.
- [2026-06] “Demystifying the Slash Pattern in Attention: The Role of RoPE” accepted to the first Foundations of Deep Generative Models (FoGen) workshop, ICML 2026.
- [2026-01] I PhinisheD my PhD study! Great thanks to the support from everyone, especially my supervisor Prof. Vincent Tan!
- [2025-12-22] Yunlong’s homepage was finally online! 🎉 Hello, world!
