Medusa Speculative

Multi-head speculative decoding for 2-3× LLM inference speedup — fully open source.

View on AIWEBTOOLS.AI