Open-source fast inference engine for Transformer models in C++ with Python bindings — 2-4x faster than PyTorch with quantization, batching, and CPU/GPU support.