Fast, lightweight inference engine (3MB binary) that wraps llama.cpp with an OpenAI-compatible API for edge deployment.