Why Can a Long Prompt Stall Active LLM Generation?
A vLLM report shows decode speed dropping to 0.5–5 tok/s during long prefills. Learn what the data proves, what it does not, and how to diagnose it.
Staff Context Architecture & RAG Researcher at Eyestech. Chennai-based distributed systems engineer (IIT Madras alumni). Researches long-context KV-cache memory structures, hierarchical context compaction, and vector retrieval economics.
A vLLM report shows decode speed dropping to 0.5–5 tok/s during long prefills. Learn what the data proves, what it does not, and how to diagnose it.
Ahead of DevDay 2026, OpenAI’s “o” leaks as an always-on personal AI assistant. Inside the -o email daemon, background loops, and the battle against Gemini and Grok.
An empirical systems audit of HNSW vs. IVF-PQ indexing across 100 million 1536-dimensional vectors. Benchmarking Qdrant, OpenSearch, Milvus, and Faiss on RAM footprint, Recall@10, QPS throughput, and NVMe disk-spilling economics.