ExLlamaV2

Blazing-fast inference engine for quantized LLMs on consumer GPUs. EXL2 format pushes 4-bit and lower with state-of-the-art tokens/sec on a single RTX card.

View on AIWEBTOOLS.AI