Skip to content
Athrael.net logoAthrael.net

« Evaluating Narwhal

Serving configurations

These are the serving settings for the twelve runs in the Narwhal evaluation, run from 14 to 18 September 2026. Each system ran once per workload with prefix caching off and once with it on.

Shared settings

SettingValue
Modelmoonshotai/Kimi-K3
FleetSix engines, tensor parallelism 8
TTFT SLO2.1 seconds
TPOT SLO45 milliseconds
Load clientAIPerf 0.12.0, streaming chat, fixed schedule, four client workers
Token countsServer-reported
Client timeout361 seconds
Prefix cachingA separate run with caching off and with caching on

Every run used these AIPerf arguments:

aiperf profile --model moonshotai/Kimi-K3 --endpoint-type chat --streaming \
  --custom-dataset-type mooncake_trace --fixed-schedule --use-server-token-count \
  --tokenizer builtin --ui none --workers-max 4 --export-level raw \
  --request-timeout-seconds 361 --export-http-trace

Narwhal

All four Narwhal runs used Narwhal v0.1.0, which is older than the current implementation on GitHub.

SettingValue
EnginevLLM 0.29.0, NIXL 1.4.1
KV transferNixlConnector, kv_both role, pull transfer
PrecisionBF16
KV cache managerHybrid
Speculative decodingDisabled
Starting allocationChat/document 1 prefill, 5 decode. Mixed 3 prefill, 3 decode
Minimum per role1 prefill engine, 1 decode engine
AdmissionOpen
Prefill concurrency4
Decode concurrency38
QueueCapacity 4, timeout 0.8 seconds
Handoff timeout25 seconds from prefill to decode
Attempts per request2
Request timeouts360 seconds overall, 11 seconds prefill, 13 seconds to first token

The allocation controller’s tuning parameters are not included.

Dynamo Planner

The two chat/document runs used these settings.

SettingValue
Dynamo1.5.0.dev20260906 for frontend and planner, 1.4.2 for workers
Worker librariesvLLM 0.29.0, NIXL 1.4.1, PyTorch 2.13.0, Transformers 5.17.0
ModeDisaggregated prefill and decode
KV transferNixlConnector, kv_both role
Engine limitsMaximum model length 1,048,576, 16,384 maximum batched tokens, block size 128
GPU memory utilization0.92
SchedulingHybrid KV cache manager, asynchronous scheduling
Request limitsPrefill 4, decode 38, queue 4
Planner allocationOpens at 3 prefill and 3 decode engines, scales between 2 and 6 engines
Planner scalingLoad and throughput scaling on, targeting TTFT 2,100 ms and ITL 45 ms
Prefix caching--enable-prefix-caching or --no-enable-prefix-caching

Ray Serve LLM

All four Ray runs used adaptive autoscaling with these settings.

SettingValue
Applicationray.serve.llm prefill/decode app with one prefill deployment, one decode deployment and one ingress
Replicas1 to 5 per deployment. Prefill starts at 1 and decode at 5
Traffic startAfter scaling down to 1 prefill and 1 decode replica
Autoscaling target2 ongoing requests per replica
Autoscaling timing30-second upscale delay, 600-second downscale delay
Autoscaling metrics10-second interval, 30-second look-back
Queues4 queued requests per deployment
Ingress1 replica, 96 ongoing requests
KV transferNixlConnector, kv_both role
EngineTensor parallelism 8, language-model-only mode
Engine limitsMaximum model length 1,048,576, 16,384 maximum batched tokens, block size 128
GPU memory utilization0.92
KV cache managerHybrid
ParsersKimi-K3 reasoning and tool parsers
Timeouts360-second server request deadline, 361-second client timeout, no client retries
Prefix cachingenable_prefix_caching: true or false