Medusa Speculative
Multi-head speculative decoding for 2-3× LLM inference speedup — fully open source.
View on AIWEBTOOLS.AI