NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference 事件

Name: NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference
Start: 2026-06-03

PRODUCT_LAUNCH2026-06-03影响: MEDIUM

NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference arXiv:2606.03910v1 Announce Type: cross Abstract: Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget. Current schedulers route on compute load and prefix-cache locality alone, ignoring the topological distance and dynamic congestion between prefill and decode instances. We close this gap

人工智能

关系图谱

NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference 事件

相关公司查看全部 (8)

相关人物查看全部 (1)

相关产品查看全部 (10)

相关技术查看全部 (10)

相关报道查看全部 (1)