Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling

Chen, Lequn; Deng, Weixin; Canumalla, Anirudh; Xin, Yu; Zhuo, Danyang; Philipose, Matthai; Krishnamurthy, Arvind

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2308.07470 (cs)

[Submitted on 14 Aug 2023 (v1), last revised 28 Feb 2024 (this version, v2)]

Title:Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling

Authors:Lequn Chen, Weixin Deng, Anirudh Canumalla, Yu Xin, Danyang Zhuo, Matthai Philipose, Arvind Krishnamurthy

View PDF HTML (experimental)

Abstract:Having large batch sizes is one of the most critical aspects of increasing the accelerator efficiency and the performance of DNN model inference. However, existing model serving systems cannot achieve adequate batch sizes while meeting latency objectives as these systems eagerly dispatch requests to accelerators to minimize the accelerator idle time. We propose Symphony, a DNN serving system that explores deferred batch scheduling to optimize system efficiency and throughput. Further, unlike other prior systems, Symphony's GPU usage is load-proportional: it consolidates workloads on the appropriate number of GPUs and works smoothly with cluster auto-scaling tools. Symphony consists of two core design points. First, Symphony defines a schedulable window in which a batch of inference requests can be dispatched. This window is computed in order to improve accelerator efficiency while meeting the request's SLO. Second, Symphony implements a scalable, low-latency, fine-grained coordination scheme across accelerators to dispatch and execute requests in the schedulable window. Through extensive scheduler-only benchmarks, we demonstrate that Symphony can schedule millions of requests per second and coordinate thousands of GPUs while also enabling robust autoscaling that adapts to workload changes. Symphony outperforms prior systems by achieving 5x higher goodput when given the same number of GPUs and 60% reduction in GPUs when given the same workload.

Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
Cite as:	arXiv:2308.07470 [cs.DC]
	(or arXiv:2308.07470v2 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2308.07470

Submission history

From: Lequn Chen [view email]
[v1] Mon, 14 Aug 2023 21:46:37 UTC (849 KB)
[v2] Wed, 28 Feb 2024 21:40:00 UTC (563 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators