Publications

I am broadly interested in both the theoretical limits and empirical applications of reinforcement learning and online learning (e.g., multi-armed bandits), with a current focus on designing practical algorithms with provable guarantees for LLMs. Feel free to reach out if you share similar interests!

Conferences

SP²ec: Adaptive Self-Speculative Decoding for Vision-Language Models
Yuqi Huang, Xingyao Li, Yunlong Hou, Fengzhuo Zhang, Jiachun Pan and Vincent Y. F. Tan.
Conference on Neural Information Processing Systems (NeurIPS), 2026.
TL;DR: We exploit the single-peak structure in self-speculative decoding to further boost the inference speed.
Yunlong Hou, Fengzhuo Zhang, Yuan Cheng, Jiachun Pan, Xingyao Li and Zhuoran Yang.
Conference on Neural Information Processing Systems (NeurIPS), 2026; The first Foundations of Deep Generative Models workshop, ICML 2026; 2026 INFORMS Annual Meeting.
TL;DR: We revisit the generalization ability of SFT and RL training from the data perspective.
Yunlong Hou, Fengzhuo Zhang, Cunxiao Du, Xuan Zhang, Jiachun Pan, Tianyu Pang, Chao Du, Vincent Y. F. Tan and Zhuoran Yang.
International Conference on Machine Learning (ICML), 2025
TL;DR: A training-free approach to adaptively select the draft hyperparameter to improve inference efficiency.
Yunlong Hou, Vincent Y. F. Tan and Zixin Zhong.
Conference on Neural Information Processing Systems (NeurIPS), 2024
TL;DR: Best Arm Identification under the piecewise-stationary environment, where the best arm has the best average performance.
Yunlong Hou, Vincent Y. F. Tan and Zixin Zhong.
International Conference on Machine Learning (ICML), 2023
TL;DR: Regret minimization with any-time risk constraint.

Journals

Yuqi Huang, Yunlong Hou and Vincent Y. F. Tan.
Under Review
TL;DR: Fixed-budget BAI in the Bayesian setting, with the option of abstention at the end of the sampling process.
Yunlong Hou, Zixin Zhong and Vincent Y. F. Tan.
Under review; European Workshop on Reinforcement Learning (EWRL), 2026.
TL;DR: Regret Minimization with a free exploration time period at the front.
Yuan Cheng, Fengzhuo Zhang, Yunlong Hou, Cunxiao Du, Chao Du, Tianyu Tang, Aixin Sun and Zhuoran Yang.
Under review; The first Foundations of Deep Generative Models workshop, ICML 2026
TL;DR: The slash pattern in attention is caused by RoPE.
Yunlong Hou, Vincent Y. F. Tan and Zixin Zhong.
IEEE Transactions on Information Theory (IEEE TIT), Volume 69, Issue 4, April 2023, doi: 10.1109/TIT.2022.3222231.
TL;DR: Best Arm Identification with risk constraint (e.g., variance).