Publications

* indicates equal contribution; † indicates project lead.

Preprints

  1. 2026Relative Generalization Invariance of LLM Pretraining. F. Zhang*,†, S. Wang*, S. Li*, T. Ruan, J. He, I. Tsang, T. Pang, C. Du, T. Zhang, Z. Yang. Under review.
  2. 2026Muon Learns More Robust and Transferable Features than Adam. T. Ruan*, F. Zhang*,†, S. Wang*, S. Zhang, D. Bergemann, Z. Yang. Under review.
  3. 2026Demystifying the Slash Pattern in Attention: The Role of RoPE. Y. Cheng*, F. Zhang*,†, Y. Hou*, C. Du, C. Du, T. Pang, A. Sun, Z. Yang. Foundations of Deep Generative Models Workshop at the International Conference on Machine Learning (ICML), 2026.

Refereed Journal Articles

  1. 2026Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods. X. Hu, F. Zhang, S. Chen, Z. Yang. Journal of Machine Learning Research (JMLR), 2026, to appear.
  2. 2026Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework. J. Wang*, F. Zhang*,†, X. Li, V. Y. F. Tan, T. Pang, C. Du, A. Sun, Z. Yang. Journal of Machine Learning Research (JMLR) 27(135):1–51, 2026.
  3. 2025LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation. X. Zhang, F. Zhang, C. Du, C. Du, T. Pang, W. Gao, M. Lin. Transactions on Machine Learning Research (TMLR), 2025.
  4. 2024Learning Regularized Graphon Mean-Field Games with Unknown Graphons. F. Zhang, V. Y. F. Tan, Z. Wang, Z. Yang. Journal of Machine Learning Research (JMLR), 25(372):1–95, 2024.
  5. 2023Active-LATHE: An Active Learning Algorithm for Boosting the Error Exponent for Learning Homogeneous Ising Trees. F. Zhang, A. Tandon, V. Y. F. Tan. IEEE Transactions on Information Theory, 69(4):2537–2555, 2023.

Refereed Conference Proceedings

  1. 2026Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion. F. Zhang, Z. Yang, D. Bergemann. ACM Conference on Economics and Computation (EC), 2026.
    Under review at Operations Research.
  2. 2026Demystifying Classifier-Free Guidance for Auto-Regressive Image Generation. Z. Zhou*, J. Pan*, F. Zhang*,†, D. Bergemann, Z. Yang. Scientific Understanding of Foundation Models Workshop at Proceedings of the Conference on Language Modeling (COLM), 2026. Oral Presentation.
    Advances in Neural Information Processing Systems (NeurIPS), 2026.
  3. 2026Rethinking “RL Generalizes, SFT Memorizes”: The Role of SFT Data. Y. Hou*, F. Zhang*,†, Y. Cheng*, J. Pan, X. Li, Z. Yang. Advances in Neural Information Processing Systems (NeurIPS), 2026.
  4. 2026Why Muon Outperforms Adam: A Curvature Perspective. S. Wang*, F. Zhang*,†, J. Li, D. Bergemann, Z. Yang. Advances in Neural Information Processing Systems (NeurIPS), 2026.
  5. 2026Muon Outperforms Adam in Tail-End Associative Memory Learning. S. Wang*, F. Zhang*,†, J. Li*, C. Du, C. Du, T. Pang, Z. Yang, M. Hong, V. Y. F. Tan. International Conference on Learning Representations (ICLR), 2026.
  6. 2026Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation. X. Li, F. Zhang, C. Du, H. Ji. AAAI Conference on Artificial Intelligence (AAAI), 2026.
  7. 2026LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification. P. Yang, C. Du, F. Zhang, H. Wang, T. Pang, C. Du, B. An. Annual Meeting of the Association for Computational Linguistics (ACL), 2026.
  8. 2025When Attention Sink Emerges in Language Models: An Empirical View. X. Gu, T. Pang, C. Du, Q. Liu, F. Zhang, C. Du, Y. Wang, M. Lin. International Conference on Learning Representations (ICLR), 2025.
  9. 2025BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms. Y. Hou*, F. Zhang*,†, C. Du*, X. Zhang*, J. Pan, T. Pang, C. Du, V. Y. F. Tan, Z. Yang. International Conference on Machine Learning (ICML), 2025.
  10. 2025What and How Does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization. Y. Zhang*, F. Zhang*, Z. Yang, Z. Wang. International Conference on Artificial Intelligence and Statistics (AISTATS), 2025.
  11. 2025Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network Approximation. X. Chen, F. Zhang, K. Yan, L. Zhao. International Conference on Learning Representations (ICLR), 2025.
  12. 2025Enhancing Multi-Text Long Video Generation Consistency without Tuning: Time-Frequency Analysis, Prompt Alignment, and Theory. X. Li, F. Zhang, J. Pan, Y. Hou, V. Y. F. Tan, Z. Yang. Building Physically Plausible World Models Workshop at the International Conference on Machine Learning (ICML), 2025.
  13. 2024From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems. J. He, S. Chen, F. Zhang, Z. Yang. International Conference on Machine Learning (ICML), 2024.
  14. 2023Learning Regularized Monotone Graphon Mean-Field Games. F. Zhang, V. Y. F. Tan, Z. Wang, Z. Yang. Conference on Neural Information Processing Systems (NeurIPS), 2023.
  15. 2022Relational Reasoning via Set Transformers: Provable Efficiency and Applications to MARL. F. Zhang, B. Liu, K. Wang, V. Y. F. Tan, Z. Yang, Z. Wang. Conference on Neural Information Processing Systems (NeurIPS), 2022.
  16. 2021Robustifying Algorithms of Learning Latent Trees with Vector Variables. F. Zhang, V. Y. F. Tan. Conference on Neural Information Processing Systems (NeurIPS), 2021.