Open-source LLM evaluation framework from Confident AI — Pytest-style unit tests, metrics and CI/CD checks for agents, RAG pipelines and prompts.