The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.
Like htop, but for AI coding agents. Monitor Claude Code & Codex CLI sessions, tokens, context window, rate limits, and ports in real-time.