Semantic cache for LLM queries that reuses similar past responses — cuts token cost by up to 100x and improves latency for ChatGPT, LangChain, and OpenAI apps.