The story so far
Long document prompts and long generated answers put pressure on different parts of an inference fleet. Prefill processes the prompt, while decode generates the answer. As traffic shifts between chat, document questions, and long generation, startup prefill/decode allocations can become a poor fit, leading to latency spikes and underutilized resources.
In my Price of Anarchy experiments a few months back, I examined routing under saturation with fixed prefill/decode splits. As I discovered limitations of static allocation under changing workloads, I went looking for dynamic alternatives. That’s when I came across Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al., which proposed adaptive scheduling mechanisms for disaggregated LLM inference. And thus Narwhal was born.
I built Narwhal, an adaptive disaggregated inference framework that reassigns running engines between prefill and decode while model weights stay loaded. Existing requests finish on their assigned engines, and new requests follow the revised allocation.
Why Narwhal? Well, the first time I got the core algorithm working and watched the system adapt in real time, it looked like the chart above. Sort of like a narwhal, right? At least that’s what I thought at the time. Anyway, several hundred hours of work later, I open-sourced Narwhal under the Apache 2.0 license.
The next step was to evaluate it under rapidly changing workloads and to compare its performance against existing frameworks. I limited the comparison to Dynamo Planner and Ray Serve LLM, which were the only frameworks I found with mature adaptive scheduling mechanisms.
TLDR: Compared against Dynamo Planner and Ray Serve LLM on identical hardware, Narwhal achieved the highest completion and SLO-qualified rates across both dynamic workloads. With prefix caching on, it delivered 76.18% of chat/document requests within the SLO, against 56.76% for Dynamo and 51.66% for Ray. Prefix caching lifted Narwhal’s SLO-qualified share from 40.30% to 76.18% on chat/document and from 43.88% to 45.63% on the mixed workload. Answer quality was nearly identical across all three frameworks for completed requests, with gaps in overall answer quality coming from completion rates. The main trade-off was an 11-second p95 TTFT tail on mixed traffic, which I aim to address in future work.
Keeping score
I wanted to know which framework got the most work done within the latency budget as the workload changed. Counting completed requests alone tells only half the story. A system might finish most requests but exceed latency limits on half of them. Another might reject requests aggressively, which can keep latencies low on paper.
For each request, I measured two latency metrics:
- TTFT (time to first token): at or below 2.1 seconds, measured from the scheduled send time so the figure includes load-generator lag.
- TPOT (time per output token): at or below 45 ms. I averaged this over tokens after the first to keep TTFT out of the figure. The average gives requests a fair chance even if some tokens are slower.
A request counted toward the SLO only if it met both limits and finished naturally with a nonempty answer and at least two output tokens.
Tables across the evaluation report two counts. Completed counts every successful response, including truncated ones. SLO-qualified counts only those that also finished naturally and met both latency limits.
The playing field
Each Narwhal run used v0.1.0, and replayed a fixed request schedule using AIPerf. I tested each framework against two workloads, each with prefix caching off and on. So, six configurations per workload, twelve runs in total. Along with outcomes and latency, I evaluated each framework’s admission and queueing behavior.
Note: All three frameworks ran with the settings listed on the serving configurations page. Each one received the same evaluation size and request schedule. I built that schedule around rapidly changing workloads to stress-test how each framework adapts under pressure, which is the scenario I designed Narwhal for.
| Input | Evaluated value |
|---|---|
| Model | moonshotai/Kimi-K3 |
| Fleet | Six TP8 engines |
| Frameworks | Narwhal, Dynamo Planner, Ray Serve LLM |
| Load client | AIPerf 0.12.0, streaming chat, four client workers |
| Scheduling | Fixed request schedule |
| Prefix caching | Separate off/on configurations for every framework and workload |
| Client timeout | 361 seconds |
| Evaluation size | 12 configurations, 172,014 offered requests in total |
The chat/document workload alternated between chat and document phases, starting and ending with chat, for five document phases in total. Each configuration received the same 14,250 scheduled requests: 6,810 chat and 7,440 document. I chose this structure to see how the frameworks handled the shift between prefill-heavy and decode-heavy phases.
The mixed workload combined five request classes: short chat, standard chat, conversation history, document, and long generation. It ran through nine consecutive load phases over one hour, with each configuration receiving 14,419 requests. Of those, 6,310 were document requests and 8,109 were chat or long-generation requests.
Completion and SLO delivery
With prefix caching enabled, Narwhal completed the highest share of offered requests and had the highest SLO-qualified share in both workloads:
| Workload | Framework | Completed / offered | SLO-qualified / offered |
|---|---|---|---|
| Chat/document | Narwhal | 13,074 / 14,250 (91.75%) | 10,855 / 14,250 (76.18%) |
| Chat/document | Dynamo Planner | 9,380 / 14,250 (65.82%) | 8,088 / 14,250 (56.76%) |
| Chat/document | Ray Serve LLM | 10,059 / 14,250 (70.59%) | 7,361 / 14,250 (51.66%) |
| Mixed | Narwhal | 13,973 / 14,419 (96.91%) | 6,579 / 14,419 (45.63%) |
| Mixed | Dynamo Planner | 8,131 / 14,419 (56.39%) | 3,135 / 14,419 (21.74%) |
| Mixed | Ray Serve LLM | 13,569 / 14,419 (94.11%) | 5,353 / 14,419 (37.12%) |
One canonical run per configuration. SLO qualification requires natural completion, a nonempty answer, at least two output tokens, TTFT ≤2.1 seconds, and TPOT ≤45 ms.
On chat/document, Narwhal delivered 2,767 more SLO-qualified requests than Dynamo and 3,494 more than Ray. On the mixed workload, it delivered 3,444 more than Dynamo and 1,226 more than Ray.
The mixed workload also exposed a large gap between Narwhal’s completion and qualification counts. It completed 13,973 requests, or 96.91% of offers, but only 6,579 qualified. The full 14,419 offers break down into four disjoint groups:
- 6,579 SLO-qualified responses.
- 6,503 completed responses that reached the output cap.
- 891 naturally completed responses that failed one or more remaining qualification checks.
- 446 failures or refusals.
Output caps accounted for most of the gap between completion and qualification. Capped answers made up 46.54% of Narwhal’s completions, 47.52% of Dynamo’s, and 47.05% of Ray’s. I kept the output limits fixed, since raising them would change the amount of decode work offered to the fleet.
The chat/document workload had far fewer capped completions: 459 for Narwhal, 238 for Dynamo, and 199 for Ray. The same qualification rule therefore removed a much smaller share of completed work in that workload.
Latency and overload
The chat/document medians were close. With caching enabled, completed-request TTFT p50 was 0.631 seconds for Narwhal, 0.627 for Dynamo, and 0.676 for Ray. Completion rates ranged from 65.82% to 91.75%, though, so those medians describe different surviving populations. The tails were further apart.
All timings below cover completed requests with prefix caching enabled. TPOT p95 is the 95th percentile of each request’s average time per output token.
| Workload | Framework | Completed / offered | TTFT p50 | TTFT p95 | TTFT p99 | TPOT p95 |
|---|---|---|---|---|---|---|
| Chat/document | Narwhal | 13,074 / 14,250 | 0.631 s | 3.601 s | 14.096 s | 36.174 ms |
| Chat/document | Dynamo Planner | 9,380 / 14,250 | 0.627 s | 7.943 s | 14.288 s | 36.129 ms |
| Chat/document | Ray Serve LLM | 10,059 / 14,250 | 0.676 s | 4.112 s | 6.833 s | 51.602 ms |
| Mixed | Narwhal | 13,973 / 14,419 | 0.384 s | 11.268 s | 14.203 s | 36.256 ms |
| Mixed | Dynamo Planner | 8,131 / 14,419 | 1.405 s | 3.520 s | 4.667 s | 36.687 ms |
| Mixed | Ray Serve LLM | 13,569 / 14,419 | 0.489 s | 1.183 s | 1.917 s | 49.747 ms |
On chat/document, Narwhal had the lowest TTFT p95 and Ray the lowest p99. On the mixed workload, Ray had the lowest value at both tail percentiles, with a TTFT p95 of 1.183 seconds against Narwhal’s 11.268 seconds. Narwhal had the lower median there, 0.384 seconds versus Ray’s 0.489, and the higher SLO-qualified count. It also had a much longer TTFT tail, with its admission-queue wait reaching 9.34 seconds at p95 on that workload.
Ray’s TPOT p95 was above the 45 ms threshold in both workloads: 51.602 ms on chat/document and 49.747 ms on mixed. On the mixed workload, 2,744 of its 13,569 completed requests missed the TPOT limit, against 107 that missed the TTFT limit.
Failures and refusals by category:
| Workload, caching on | Framework | Failed | Refused | Recorded failure behavior |
|---|---|---|---|---|
| Chat/document | Narwhal | 765 | 411 | All 765 failures were queue expiry before first output |
| Chat/document | Dynamo Planner | 47 | 4,823 | Streams cut mid-output by three Planner scale-downs |
| Chat/document | Ray Serve LLM | 4,191 | 0 | 4,130 service-unavailable errors, 61 incomplete response payloads |
| Mixed | Narwhal | 51 | 395 | Queue expiry and separate overload refusals |
| Mixed | Dynamo Planner | 11 | 6,277 | Streamed errors and separate HTTP 529 refusals |
| Mixed | Ray Serve LLM | 850 | 0 | HTTP 503 service-unavailable errors |
Ray Serve LLM shows no refusals because it doesn’t refuse requests explicitly. It returns 503s instead, so the evaluator counted those as failures.
Put side by side, these show three different ways of handling overload. Dynamo turned away 43.53% of mixed requests with HTTP 529 refusals, Ray returned 503s, and Narwhal held requests in its admission queue. On chat/document, 737 of its 765 queue expiries were chat requests. Across the whole workload, its chats qualified at 58.16%, against 92.66% for documents.
Delivered answer quality
I scored the document responses using normalized token-overlap F1 against reference answers. A reported value of 48% means a mean F1 of 0.48.
I calculated three averages to show what happened as I applied delivery requirements:
- Completed-only F1 averages F1 over completed document requests, including output-cap completions.
- All-offer F1 divides the sum of completed-document F1 by every offered document request.
- SLO-delivered F1 divides the sum of F1 from SLO-qualified document requests by every offered document request.
These are the cache-on results. All-offer and SLO-delivered values divide by 7,440 document offers on chat/document and 6,310 on the mixed workload.
| Workload | Framework | Completed-only F1 | All-offer F1 | SLO-delivered F1 |
|---|---|---|---|---|
| Chat/document | Narwhal | 48.84% | 48.41% | 45.42% |
| Chat/document | Dynamo Planner | 48.66% | 40.06% | 39.09% |
| Chat/document | Ray Serve LLM | 48.35% | 43.08% | 36.45% |
| Mixed documents | Narwhal | 47.98% | 46.06% | 37.88% |
| Mixed documents | Dynamo Planner | 48.28% | 24.93% | 17.61% |
| Mixed documents | Ray Serve LLM | 48.08% | 44.41% | 28.35% |
The three frameworks produced answers of nearly equal quality when a request completed. Completed-only F1 spanned 0.30 points on mixed documents and 0.49 points on chat/document, so the all-offer gaps come from completion rates. Dynamo completed 3,259 of 6,310 mixed document requests (51.65%), which roughly halved its score, from 48.28% to 24.93%. Narwhal completed 6,058 (96.01%) and Ray 5,828 (92.36%).
The SLO filter hit Ray hardest. It removes completed answers that missed the TTFT or TPOT limit or hit the output cap, and it took 16.06 points off Ray’s mixed-document score. Narwhal lost 8.18 points and Dynamo 7.32.
I also scored four other request classes in the mixed workload. Across all five classes and all 14,419 mixed offers, all-offer F1 was 28.38% for Narwhal, 16.13% for Dynamo, and 27.50% for Ray.
Prefix caching
I ran the cache-off configurations to see how much the delivery result changed when prefix reuse was unavailable. vLLM’s prefix caching reuses key/value state for shared prompt prefixes, so the engine doesn’t have to prefill that part of the input again.
For Narwhal, the observed change depended heavily on the workload:
| Workload | Prefix caching | Completion | SLO-qualified | TTFT p50, completed |
|---|---|---|---|---|
| Chat/document | Off | 91.54% | 40.30% | 2.290 s |
| Chat/document | On | 91.75% | 76.18% | 0.631 s |
| Mixed | Off | 97.09% | 43.88% | 0.576 s |
| Mixed | On | 96.91% | 45.63% | 0.384 s |
One canonical run per configuration. SLO-qualified shares use all 14,250 or 14,419 offers, and TTFT medians use completed requests. I computed point changes from request counts, so they can differ slightly from the rounded table values.
On chat/document, caching added only 30 completed requests (0.21 points) but 5,112 qualified requests (35.87 points). Median TTFT dropped from 2.290 seconds, above the 2.1-second budget, to 0.631 seconds.
The gain came from document requests. Narwhal completed almost all document offers either way (7,377 without caching, 7,375 with it), but qualified document requests rose from 1,758 to 6,894. Qualified chats slipped from 3,985 to 3,961.
The mixed workload gained much less: 252 more qualified requests (1.75 points) and 26 fewer completed ones (0.18 points). Median TTFT dropped from 0.576 to 0.384 seconds, but p95 stayed at 11.2 to 11.3 seconds in both runs.
All three frameworks had a higher qualified share with caching enabled in both workloads. Narwhal had the smallest mixed-workload increase but kept the highest qualified share:
SLO-qualified share with caching off and on, all three frameworks
| Workload | Framework | Cache off | Cache on |
|---|---|---|---|
| Chat/document | Narwhal | 40.30% | 76.18% |
| Chat/document | Dynamo Planner | 26.75% | 56.76% |
| Chat/document | Ray Serve LLM | 15.87% | 51.66% |
| Mixed | Narwhal | 43.88% | 45.63% |
| Mixed | Dynamo Planner | 14.87% | 21.74% |
| Mixed | Ray Serve LLM | 25.45% | 37.12% |
Scope and provenance
The results cover Kimi-K3 on six TP8 engines with these SLO thresholds and the settings on the serving configurations page. Because of resource constraints, each of the twelve configurations was run once.
The benchmarks show clear separation on some metrics, while other margins remain narrow. These include the chat/document TTFT p50 gap of under 5 ms and the sub-point spread in completed-only F1. Readers should interpret the results cautiously.
Subsequent Narwhal updates may produce different results.
Full results tables and configuration details:
Narwhal’s source and documentation describe the current implementation.
If you work on disaggregated serving, or you’ve tried adaptive prefill/decode allocation on your own workloads, I’d love to hear about it. Reach out on GitHub, LinkedIn, or via email.
Athrael.net