Research
My research interest is statistical and optimization theory for foundation models. I combine theoretical analysis with controlled experiments to study training, inference, and deployment mechanisms of modern learning systems.
Optimization Theory of Foundation Models
This line of work examines how data distributions, optimizer geometry, and architectural structure jointly shape foundation-model training. It develops theoretical explanations for differences in optimizer performance and characterizes how optimization choices influence learned representations.
- Relative Generalization Invariance of LLM Pretraining.Under review.
- Why Muon Outperforms Adam: A Curvature Perspective.Advances in Neural Information Processing Systems (NeurIPS), 2026.
- Muon Learns More Robust and Transferable Features than Adam.Under review.
- Muon Outperforms Adam in Tail-End Associative Memory Learning.International Conference on Learning Representations (ICLR), 2026.
- Demystifying the Slash Pattern in Attention: The Role of RoPE.Foundations of Deep Generative Models Workshop at the International Conference on Machine Learning (ICML), 2026.
Statistical Foundations of Reasoning and In-Context Learning
This research develops statistical theories of how language models use demonstrations and reasoning traces to make predictions. By connecting in-context learning and chain-of-thought prompting to Bayesian inference, it characterizes their mechanisms, performance guarantees, and limitations.
- Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods.Journal of Machine Learning Research (JMLR), 2026, to appear.
- What and How Does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization.International Conference on Artificial Intelligence and Statistics (AISTATS), 2025.
- Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework.Journal of Machine Learning Research (JMLR) 27(135):1–51, 2026.
- Rethinking “RL Generalizes, SFT Memorizes”: The Role of SFT Data.Advances in Neural Information Processing Systems (NeurIPS), 2026.
- Demystifying Classifier-Free Guidance for Auto-Regressive Image Generation.Scientific Understanding of Foundation Models Workshop at Proceedings of the Conference on Language Modeling (COLM), 2026. Oral Presentation.
Advances in Neural Information Processing Systems (NeurIPS), 2026.
Game-Theoretic and Economic Foundations of AI Deployment
This research develops theory for the equilibria analysis in large populations of heterogeneous users. In AI deployment, it examines how model performance and computational costs shape users’ personalization choices, equilibrium congestion, and platform design.
- Learning Regularized Graphon Mean-Field Games with Unknown Graphons.Journal of Machine Learning Research (JMLR), 25(372):1–95, 2024.
- Learning Regularized Monotone Graphon Mean-Field Games.Conference on Neural Information Processing Systems (NeurIPS), 2023.
- Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion.ACM Conference on Economics and Computation (EC), 2026.
Under review at Operations Research.
Statistical Foundations of Efficient and Reliable AI Deployment
This line of work studies how to improve the efficiency and reliability of autoregressive generation with theoretical guarantees.
- BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms.International Conference on Machine Learning (ICML), 2025.
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation.AAAI Conference on Artificial Intelligence (AAAI), 2026.
- LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation.Transactions on Machine Learning Research (TMLR), 2025.
Statistical Machine Learning Theory
This line of work studies the statistical foundations of learning graphical models from data.
- Active-LATHE: An Active Learning Algorithm for Boosting the Error Exponent for Learning Homogeneous Ising Trees.IEEE Transactions on Information Theory, 69(4):2537–2555, 2023.
- Robustifying Algorithms of Learning Latent Trees with Vector Variables.Conference on Neural Information Processing Systems (NeurIPS), 2021.