Tag:llm
All the articles with the tag "llm".
Evaluating Narwhal: Adaptive Prefill/Decode Allocation Under Rapidly Changing Workloads
Updated:On Kimi-K3 with six TP8 engines, Narwhal delivered 76.18% of chat/document requests within the latency SLO, against 56.76% for Dynamo Planner and 51.66% for Ray Serve LLM. The main trade-off was an 11-second p95 TTFT tail on mixed traffic.
The Most Beautiful RAG: Starring Colnomic, Qdrant, Minio and Friends
Updated:Introducing the first project in my little-scripts monorepo - A simple, yet beautiful RAG implementation using Colnomic, Qdrant and Nomic
Down the Rabbit Hole - One step closer to Production Grade GraphRAG
After my initial experiment with GraphRAG using Qdrant, Neo4j, and Ollama, I took on a journey to build a more dynamic and context-aware system. This post dives into the details of how I constructed a dynamic ontology for NLP GraphRag.
GraphRAG with Qdrant, Neo4j, and Ollama (Using Qwen2.5:3b and Nomic text embeddings)
I've been playing with a new approach to RAG systems - combining vector search with knowledge graphs for more contextual, relationship-aware answers. Here's what I've built, how it works, and why you might want to try it yourself.
Athrael.net