FlexGen

High-throughput generative inference of LLMs on a single commodity GPU \u2014 run OPT-175B on a 16GB card.

View on AIWEBTOOLS.AI