Open-source heterogeneous CPU+GPU LLM inference — run 671B DeepSeek V3 on a single 24GB GPU at usable speeds. Pushes the limit of consumer hardware inference.