ARK: Avoiding Routing Collisions for KV Cache Transfer in Disaggregated LLM Inference
Hung-Chun Lin, Ting-Wei Hsu, Chung-En Ho, Ahmed Saeed
Proceedings of the ACM SIGCOMM 2026 Conference, pp. 2067–2073, 2026.
Recommended citation: Hung-Chun Lin, Ting-Wei Hsu, Chung-En Ho, and Ahmed Saeed. "ARK: Avoiding Routing Collisions for KV Cache Transfer in Disaggregated LLM Inference." Proceedings of the ACM SIGCOMM 2026 Conference, pp. 2067–2073, 2026. DOI: 10.1145/3789240.3828750. View paper
ARK coordinates network paths for KV-cache transfers in disaggregated LLM inference. It reserves distinct paths for concurrent, long-lived transfers by choosing source ports that map to different spine switches, reducing routing collisions without switch modifications or receiver-side packet reordering.
