I am an Applied Scientist II at Amazon Store Foundational AI, where I build foundation models and large language models for shopping search, retrieval, and recommendation. I received my Ph.D. from the Graduate Group in Applied Mathematics (GGAM) at UC Davis in March 2022.
My research centers on reinforcement learning and policy optimization for large language models, and on stochastic and zeroth-order optimization.
You can find me on Google Scholar, OpenReview, and LinkedIn.
Research Impact
My research spans the post-training of large language models with reinforcement learning and the optimization theory that underlies modern machine learning.
Large language model post-training and reinforcement learning (RLVR). My recent work develops post-training methods for large language models, spanning supervised fine-tuning and reinforcement learning with verifiable rewards (RLVR). Representative contributions include SFT Doesn’t Always Hurt General Capabilities (ICLR 2025), which characterizes when domain-specific fine-tuning preserves versus erodes a model’s general ability; ST-PPO: Stabilized Off-Policy Proximal Policy Optimization for Multi-Turn Agents Training (COLM 2026), for training multi-turn LLM agents; Reward-Wise Value Estimation for Multi-Reward Optimization in LLMs (ICML 2026); and Reinforcement Learning in Inference Time: A Perspective from Successive Policy Iterations (NeurIPS 2025). In applied settings, these methods support small-language-model pretraining and post-training — on-policy distillation and RLVR — for large-scale shopping generative tasks such as query rewriting and ranking.
My first-authored paper Zeroth-Order Algorithms for Nonconvex-Strongly-Concave Minimax Problems with Improved Complexities (Journal of Global Optimization, 2022; earlier NeurIPS 2020 Workshop and arXiv versions) develops derivative-free (zeroth-order) algorithms with improved iteration and query complexity for nonconvex-strongly-concave minimax problems, where gradients are unavailable and only function evaluations can be queried. Cited across its versions by work at NeurIPS, ICML, ICLR, JMLR, AISTATS, TMLR, and SIAM Journal on Optimization, it has contributed to two further lines of research.
Min-max optimization theory. The paper’s complexity analysis and single-loop zeroth-order scheme are a building block for nonconvex saddle-point theory, extended by later work on single-loop algorithms without strong concavity (AISTATS 2022), duality-gap and Polyak–Łojasiewicz rates (JMLR, 2023), Nash-equilibrium search and derivative-free projection methods (SIAM Journal on Optimization 2021, 2024), and robust zeroth-order minimax methods (NeurIPS 2024).
Zeroth-order LLM training. The paper’s core idea — estimating descent directions from forward evaluations alone, without backpropagation — now underpins memory-efficient LLM fine-tuning, and it is cited by MeZO (NeurIPS 2023), DPZero (ICML 2023), Addax (ICLR 2024), HiZOO (ICLR 2025), MUZO (EMNLP 2025), and PaZO (NeurIPS).
Representative citing work across its NeurIPS 2020 Workshop, arXiv, and journal versions (NeurIPS, ICML, ICLR, JMLR, AISTATS, TMLR, SIAM Journal on Optimization, EMNLP, and IJCAI):
| Year | Venue | Citing paper |
|---|---|---|
| 2026 | NeurIPS | PaZO: Preconditioned Accelerated Zeroth-Order Optimization for Fine-Tuning LLMs |
| 2026 | NeurIPS | Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-Training |
| 2026 | NeurIPS | Stochastic Regret Guarantees for Online Zeroth- and First-Order Bilevel Optimization |
| 2026 | AISTATS | ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models |
| 2026 | JMLR | A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems |
| 2025 | NeurIPS | Zeroth-Order Optimization Finds Flat Minima |
| 2025 | ICLR | Adversarial Machine Unlearning |
| 2025 | ICLR | Second-Order Fine-Tuning without Pain for LLMs: A Hessian-Informed Zeroth-Order Optimizer (HiZOO) |
| 2025 | EMNLP | MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Models |
| 2025 | IJCAI | Sharpness-Aware Zeroth-Order Optimization for Graph Transformers |
| 2025 | TMLR | Min-Max Optimisation for Nonconvex-Nonconcave Functions Using a Random Zeroth-Order Extragradient Algorithm |
| 2025 | TMLR | A Framework for Finding Local Saddle Points in Two-Player Zero-Sum Black-Box Games |
| 2024 | NeurIPS | Robust and Faster Zeroth-Order Minimax Optimization: Complexity and Applications |
| 2024 | ICML | Delving into the Convergence of Generalized Smooth Minimax Optimization |
| 2024 | ICLR | Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models |
| 2024 | SIAM J. Opt. | Derivative-Free Alternating Projection Algorithms for General Nonconvex-Concave Minimax Problems |
| 2023 | NeurIPS | Fine-Tuning Language Models with Just Forward Passes (MeZO) |
| 2023 | NeurIPS | SimFBO: Towards Simple, Flexible and Communication-Efficient Federated Bilevel Learning |
| 2023 | ICML | DPZero: Private Fine-Tuning of Language Models without Backpropagation |
| 2023 | ICML | Communication-Efficient Federated Hypergradient Computation via Aggregated Iterative Differentiation |
| 2023 | AISTATS | Stochastic Gradient Descent-Ascent: Unified Theory and New Efficient Methods |
| 2023 | AISTATS | No-regret Sample-Efficient Bayesian Optimization for Finding Nash Equilibria with Unknown Utilities |
| 2023 | JMLR | Fast Objective & Duality Gap Convergence for Nonconvex-Strongly-Concave Min-Max Problems with PL Condition |
| 2023 | JMLR | Zeroth-Order Alternating Gradient Descent Ascent Algorithms for a Class of Nonconvex-Nonconcave Minimax Problems |
| 2022 | ICML | Gradient-Free Method for Heavily Constrained Nonconvex Optimization |
| 2022 | AISTATS | Faster Single-Loop Algorithms for Minimax Optimization without Strong Concavity |
| 2022 | AISTATS | Zeroth-Order Methods for Convex-Concave Min-Max Problems: Applications to Decision-Dependent Risk Minimization |
| 2022 | SIAM J. Opt. | New First-Order Algorithms for Stochastic Variational Inequalities |
| 2021 | NeurIPS | Global Convergence to Local Minmax Equilibrium in Classes of Nonconvex Zero-Sum Games |
| 2021 | AISTATS | Direct-Search for a Class of Stochastic Min-Max Problems |
| 2021 | AISTATS | AdaGDA: Faster Adaptive Gradient Descent Ascent Methods for Minimax Optimization |
| 2021 | SIAM J. Opt. | Efficient Search of First-Order Nash Equilibria in Nonconvex-Concave Smooth Min-Max Problems |
| 2020 | JMLR | Accelerated Zeroth-Order and First-Order Momentum Methods from Mini to Minimax Optimization |
Experience
At Amazon Store Foundational AI, I have worked on:
- 2022–2023 — RoBERTa-based behavior foundation models and graph models for search/ads retrieval and recommendation.
- 2024 — Instruction-following LLM-based embedding models for retrieval and recommendation.
- 2024 — Generative recommender for customer representation (Synergen).
- 2025 — Small language model pretraining, midtraining, and post-training (on-policy distillation, RLVR) for shopping generative tasks such as query rewrite, final-stage ranking, and shopping mission generation.
- 2025 — Foundation model post-training on user behavior data.
- 2026 — Personalized search agent for user-context retrieval, based on SQL and semantic search.
Publications
You can also find my articles on my Google Scholar profile.
Large Language Models and Reinforcement Learning
Reward-Wise Value Estimation for Multi-Reward Optimization in Large Language Models
Published in ICML 2026 Workshop on Decision-Making from Offline Datasets to Online Adaptation, 2026
Quan Wei, Zhongruo Wang, Chenliang Li, Xi Chen, Oana Frunza, Yang Katie Zhao, Mingyi Hong.
ST-PPO: Stabilized Off-Policy Proximal Policy Optimization for Multi-Turn Agents Training
Published in COLM 2026, 2026
Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong.
Reinforcement Learning in Inference Time: A Perspective from Successive Policy Iterations
Published in NeurIPS 2025 Workshop on Reasoning and Planning for Large Language Models, 2025
Xinnan Zhang, Chenliang Li, Siliang Zeng, Jiaxiang Li, Zhongruo Wang, Songtao Lu, Alfredo Garcia, Mingyi Hong.
GmNet: Revisiting Gating Mechanisms From A Frequency View
Published in ICLR 2025, 2025
Yifan Wang, Xu Ma, Yitian Zhang, Zhongruo Wang, Sung-Cheol Kim, Vahid Mirjalili, Vidya Renganathan, Yun Fu.
SFT Doesn’t Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
Published in ICLR 2025, 2025
Jiacheng Lin*, Zhongruo Wang*, Kun Qian, Tian Wang, Arvind Srinivasan, Hansi Zeng, Ruochen Jiao, Xie Zhou, Jiri Gesi †, Dakuo Wang, Yufan Guo, Kai Zhong, Weiqi Zhang, Sujay Sanghavi, Changyou Chen, Hyokun Yun, Lihong Li.
* Equal contribution.
† Jiri Gesi has zero contribution.
Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond
Published in arXiv:2408.11338, 2024
Minghao Liu, Zonglin Di, Jiaheng Wei, Zhongruo Wang, Hengxiang Zhang, Ruixuan Xiao, Haoyu Wang, Jinlong Pang, Hao Chen, Ankit Shah, Hongxin Wei, Xinlei He, Zhaowei Zhao, Haobo Wang, Lei Feng, Jindong Wang, James Davis, Yang Liu.
GT2Vec: Large Language Models for Knowledge Graph Augmented Text Embedding
Published in KDD 2024 Workshop on Structured Knowledge for Large Language Models, 2024
Jiacheng Lin, Kun Qian, Haoyu Han, Nurendra Choudhary, Tianxin Wei, Zhongruo Wang, Sahika Genc, Edward W. Huang, Sheng Wang, Karthik Subbian, Danai Koutra, Jimeng Sun.
Recommendation Systems
Synergen: Contextualized Generative Recommender for Unified Search and Recommendation
Published in arXiv:2509.21777, 2025
V. R. Gao, C. Xue, M. Versage, X. Zhou, Zhongruo Wang, C. Li, et al.
Optimization
Zeroth-Order Algorithms for Nonconvex-Strongly-Concave Minimax Problems with Improved Complexities
Published in Journal of Global Optimization, 87(2):709-740, 2023
Zhongruo Wang (first author), Krishnakumar Balasubramanian, Shiqian Ma, Meisam Razaviyayn. An earlier version appeared at the NeurIPS 2020 Workshop.
New Algorithms for Riemannian Optimization and Minimax Problems with Machine Learning
Published in Ph.D. Dissertation, University of California, Davis, 2022
Zhongruo Wang.
A Manifold Proximal Linear Method for Sparse Spectral Clustering with Application to Single-Cell RNA Sequencing Data Analysis
Published in INFORMS Journal on Optimization, 2021
Zhongruo Wang (first author), Bingyuan Liu, Shixiang Chen, Shiqian Ma, Lingzhou Xue, Hongyu Zhao.
