High-throughput inference engine forked from vLLM \u2014 powers PygmalionAI with support for 100+ model architectures.