Tag:inference
All the articles with the tag "inference".
Evaluating Narwhal: Adaptive Prefill/Decode Allocation Under Rapidly Changing Workloads
Updated:On Kimi-K3 with six TP8 engines, Narwhal delivered 76.18% of chat/document requests within the latency SLO, against 56.76% for Dynamo Planner and 51.66% for Ray Serve LLM. The main trade-off was an 11-second p95 TTFT tail on mixed traffic.
The Price of Anarchy in Disaggregated Inference
I split NVIDIA Dynamo's prefill and decode into three competing games and measured the Price of Anarchy on a 3-node B200 cluster. While the GPUs had headroom, no router tuning moved the needle; the moment they saturated, one parameter was the gap between a 1-second tail and a 28-second one. So I built a 270-line controller that watches for that moment and flips the switch, without touching Dynamo's core.
Athrael.net