Open-source pure-Python LLM inference framework with token-level scheduling and async I/O — designed for lightweight high-throughput serving on commodity GPUs.