DeepEval

EvaluationTesting

Open-source LLM evaluation framework. Unit-test your LLM outputs with 14+ metrics including hallucination, toxicity, and RAG.

Visit website View on GitHub
Built with
PythonLLM
Share this part