Tag:benchmark
All the articles with the tag "benchmark".
Evaluating Narwhal: Adaptive Prefill/Decode Allocation Under Rapidly Changing Workloads
Updated:On Kimi-K3 with six TP8 engines, Narwhal delivered 76.18% of chat/document requests within the latency SLO, against 56.76% for Dynamo Planner and 51.66% for Ray Serve LLM. The main trade-off was an 11-second p95 TTFT tail on mixed traffic.
You too can run the Vidore Benchmark with less than 32GB of GPU VRAM
Quick, practical notes to run the Vidore benchmark smoothly on a single 32GB GPU: dtype, batch size, and common OOM fixes.
Athrael.net