A small, educational vLLM implementation for exploring LLM inference, scheduling, KV-cache management, and GPU execution.