TokenScale
Proactive autoscaling for disaggregated LLM serving, driven by token-level workload dynamics.
October 2024 – November 2025 · Singapore
TokenScale is a proactive autoscaling system for LLM inference with prefill/decode disaggregation. It enables timely and accurate resource scaling based on token-level workload dynamics. The work was accepted to ACM SoCC’26.
- Implemented a cluster-level LLM inference control plane in approximately 6,000 lines of Go, including a traffic gateway, autoscaler, and router.
- Supported advanced routing and autoscaling policies for disaggregated serving.
- Demonstrated operation on a 32-node AMD/NVIDIA cluster, improving performance by up to 36% while reducing costs.