cv
Education, research, industry experience, and publications.
Basics
| Name | Ruiqi Lai |
| Label | PhD Student, Nanyang Technological University |
| RUIQI003@e.ntu.edu.sg | |
| Phone | (+65) 81776559 |
| Url | https://lrq619.github.io |
| Summary | Research interests: LLM inference, serverless computing, autoscaling, and distributed systems. |
Education
-
2023.08 - Present PhD Student
Nanyang Technological University
- Research interests: LLM inference, serverless computing, autoscaling, and distributed systems.
-
2019.09 - 2023.08 Bachelor of Engineering
Shanghai Jiao Tong University, University of Michigan–SJTU Joint Institute
Electrical and Computer Engineering
- Relevant coursework: Data Structures, Algorithms, Digital Circuits, Computer Architecture, Operating Systems, GPU Parallel Computing, and Distributed Cloud Computing.
Work
- Starting Jan 2027
Incoming Research Intern (Hired; Not Yet Started)
NVIDIA
Santa Clara, CA, USA
- Planned research focus: inference optimization.
- 2026.03 - Present
Research Intern
Alibaba, Platform Technology – AI Training Engine
Beijing, China
- Independently designed and implemented the Diffusion Transformer (DiT) model abstraction layer in ROLL, introducing unified interfaces for diffusion model forward passes, sampling, trajectory generation, and intermediate training states. Provided consistent execution interfaces for training and inference engines, decoupling model implementations from backend systems and substantially improving ROLL’s adaptability to DiT workloads.
- Implemented full support for FlowGRPO and DiffNFT post-training algorithms and completed end-to-end, multi-node, multi-GPU training of the 20B-parameter Qwen-Image model. Across approximately 400 GPU-hours of experiments, improved aggregate scores on multiple established DiT post-training datasets from approximately 0.3 to 0.9. Human evaluation confirmed substantial improvements in rendered text accuracy, semantic consistency, and overall visual quality.
- Built an end-to-end post-training pipeline for Qwen-Image, integrating FSDP2-based distributed training with vLLM-Omni-based parallel inference and rollout generation. Enabled the complete workflow from model loading and distributed training to online sampling, reward computation, and post-training inference validation.
- Released as the Diffusion RL support in ROLL v0.4.0, with the contribution credited in the official release notes.
- Led the Spotlight research project, jointly optimizing system infrastructure and algorithms for DiT post-training.
- 2021.12 - 2022.05
Software Engineering Intern
ByteDance
Shanghai, China
- Contributed to distributed file system development in C/C++ for ByteDance’s Volcengine cloud services and added a new data node to a Hadoop cluster.
Projects
- 2026.03 - 2026.06
Spotlight
PhD Student · Beijing, China
- Proposed Spotlight, a system that uses low-cost spot GPUs for reinforcement learning post-training of Diffusion Transformers, reducing training costs to one-seventh of the baseline.
- Designed dynamic seed exploration to offload the search for high-quality seeds to spot GPUs during training, along with elastic sequence parallelism that adjusts the parallelism degree to spot GPU availability in real time.
- Built a preemption-aware scheduler with tensor checkpointing to bound lost computation when spot instances are preempted.
- 2024.10 - 2025.11
TokenScale
PhD Student · Singapore
- Proposed TokenScale, a proactive autoscaling system for LLM inference with prefill/decode disaggregation, enabling timely and accurate resource scaling based on token-level workload dynamics. Accepted to ACM SoCC’26.
- Implemented a cluster-level LLM inference control plane in approximately 6,000 lines of Go, including a traffic gateway, autoscaler, and router with support for advanced routing and autoscaling policies.
- Demonstrated operation on a 32-node AMD/NVIDIA cluster, improving performance by up to 36% while reducing costs.
- 2023.08 - 2024.08
Liquid
PhD Student · Beijing, China
- Proposed and implemented Liquid on top of vLLM, an autoscaling mechanism for LLM inference that dynamically adjusts parallelism to allocate resources across multiple dimensions and improve overall resource utilization.
Publications
-
2026 SPOTLIGHT: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training
arXiv preprint
Ruiqi Lai, Dakai An, Wei Gao, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Dmitrii Ustiugov, Wei Wang.
-
2026 TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
ACM Symposium on Cloud Computing (SoCC’26) — Accepted
Ruiqi Lai, Hongrui Liu, Siyu Cao, Siyang Shao, Yixin Zhang, Chengzhi Liu, Dmitrii Ustiugov.
-
2025 Manage the Workloads, not the Cluster: Designing a Control Plane for Large-Scale AI Clusters.
EuroMLSys@ASPLOS/EuroSys’25
Ruiqi Lai, Siyu Cao, Leqi Li, Luo Mai, Dmitrii Ustiugov.