LightLLM

Open-source pure-Python LLM inference framework with token-level scheduling and async I/O — designed for lightweight high-throughput serving on commodity GPUs.

View on AIWEBTOOLS.AI