Serving configurations
These are the serving settings for the twelve runs in the Narwhal evaluation, run from 14 to 18 September 2026. Each system ran once per workload with prefix caching off and once with it on.
Shared settings
| Setting | Value |
|---|---|
| Model | moonshotai/Kimi-K3 |
| Fleet | Six engines, tensor parallelism 8 |
| TTFT SLO | 2.1 seconds |
| TPOT SLO | 45 milliseconds |
| Load client | AIPerf 0.12.0, streaming chat, fixed schedule, four client workers |
| Token counts | Server-reported |
| Client timeout | 361 seconds |
| Prefix caching | A separate run with caching off and with caching on |
Every run used these AIPerf arguments:
aiperf profile --model moonshotai/Kimi-K3 --endpoint-type chat --streaming \
--custom-dataset-type mooncake_trace --fixed-schedule --use-server-token-count \
--tokenizer builtin --ui none --workers-max 4 --export-level raw \
--request-timeout-seconds 361 --export-http-trace
Narwhal
All four Narwhal runs used Narwhal v0.1.0, which is older than the current implementation on GitHub.
| Setting | Value |
|---|---|
| Engine | vLLM 0.29.0, NIXL 1.4.1 |
| KV transfer | NixlConnector, kv_both role, pull transfer |
| Precision | BF16 |
| KV cache manager | Hybrid |
| Speculative decoding | Disabled |
| Starting allocation | Chat/document 1 prefill, 5 decode. Mixed 3 prefill, 3 decode |
| Minimum per role | 1 prefill engine, 1 decode engine |
| Admission | Open |
| Prefill concurrency | 4 |
| Decode concurrency | 38 |
| Queue | Capacity 4, timeout 0.8 seconds |
| Handoff timeout | 25 seconds from prefill to decode |
| Attempts per request | 2 |
| Request timeouts | 360 seconds overall, 11 seconds prefill, 13 seconds to first token |
The allocation controller’s tuning parameters are not included.
Dynamo Planner
The two chat/document runs used these settings.
| Setting | Value |
|---|---|
| Dynamo | 1.5.0.dev20260906 for frontend and planner, 1.4.2 for workers |
| Worker libraries | vLLM 0.29.0, NIXL 1.4.1, PyTorch 2.13.0, Transformers 5.17.0 |
| Mode | Disaggregated prefill and decode |
| KV transfer | NixlConnector, kv_both role |
| Engine limits | Maximum model length 1,048,576, 16,384 maximum batched tokens, block size 128 |
| GPU memory utilization | 0.92 |
| Scheduling | Hybrid KV cache manager, asynchronous scheduling |
| Request limits | Prefill 4, decode 38, queue 4 |
| Planner allocation | Opens at 3 prefill and 3 decode engines, scales between 2 and 6 engines |
| Planner scaling | Load and throughput scaling on, targeting TTFT 2,100 ms and ITL 45 ms |
| Prefix caching | --enable-prefix-caching or --no-enable-prefix-caching |
Ray Serve LLM
All four Ray runs used adaptive autoscaling with these settings.
| Setting | Value |
|---|---|
| Application | ray.serve.llm prefill/decode app with one prefill deployment, one decode deployment and one ingress |
| Replicas | 1 to 5 per deployment. Prefill starts at 1 and decode at 5 |
| Traffic start | After scaling down to 1 prefill and 1 decode replica |
| Autoscaling target | 2 ongoing requests per replica |
| Autoscaling timing | 30-second upscale delay, 600-second downscale delay |
| Autoscaling metrics | 10-second interval, 30-second look-back |
| Queues | 4 queued requests per deployment |
| Ingress | 1 replica, 96 ongoing requests |
| KV transfer | NixlConnector, kv_both role |
| Engine | Tensor parallelism 8, language-model-only mode |
| Engine limits | Maximum model length 1,048,576, 16,384 maximum batched tokens, block size 128 |
| GPU memory utilization | 0.92 |
| KV cache manager | Hybrid |
| Parsers | Kimi-K3 reasoning and tool parsers |
| Timeouts | 360-second server request deadline, 361-second client timeout, no client retries |
| Prefix caching | enable_prefix_caching: true or false |
Athrael.net