Promptfoo
Test and evaluate LLM prompts and RAG pipelines. Red-teaming, A/B testing, and CI integration for prompt quality.
Test and evaluate LLM prompts and RAG pipelines. Red-teaming, A/B testing, and CI integration for prompt quality.
Open-source LLM evaluation framework. Unit-test your LLM outputs with 14+ metrics including hallucination, toxicity, and RAG.
Holistic Evaluation of Language Models. Evaluates LLMs across 42 scenarios and 7 metrics for a complete capability picture.
Beyond the Imitation Game benchmark. 200+ tasks designed to probe LLM capabilities beyond standard NLP benchmarks.
Benchmark for evaluating LLMs as agents across 8 distinct environments including OS, web browser, and database tasks.
Open-source data labeling platform. Annotate text, images, audio, and video for training and evaluating AI models.
Framework for evaluating RAG pipelines. Measures faithfulness, answer relevancy, context recall, and more without ground truth.