Research

My research interest is statistical and optimization theory for foundation models. I combine theoretical analysis with controlled experiments to study training, inference, and deployment mechanisms of modern learning systems.

Optimization Theory of Foundation Models

This line of work examines how data distributions, optimizer geometry, and architectural structure jointly shape foundation-model training. It develops theoretical explanations for differences in optimizer performance and characterizes how optimization choices influence learned representations.

Statistical Foundations of Reasoning and In-Context Learning

This research develops statistical theories of how language models use demonstrations and reasoning traces to make predictions. By connecting in-context learning and chain-of-thought prompting to Bayesian inference, it characterizes their mechanisms, performance guarantees, and limitations.

Game-Theoretic and Economic Foundations of AI Deployment

This research develops theory for the equilibria analysis in large populations of heterogeneous users. In AI deployment, it examines how model performance and computational costs shape users’ personalization choices, equilibrium congestion, and platform design.

Statistical Foundations of Efficient and Reliable AI Deployment

This line of work studies how to improve the efficiency and reliability of autoregressive generation with theoretical guarantees.

Statistical Machine Learning Theory

This line of work studies the statistical foundations of learning graphical models from data.