OpenAI Evals
Framework for evaluating LLMs and LLM systems. Define custom evaluation tasks and run them against any model.
Framework for evaluating LLMs and LLM systems. Define custom evaluation tasks and run them against any model.
Open-source LLM evaluation framework. Unit-test your LLM outputs with 14+ metrics including hallucination, toxicity, and RAG.
Test and evaluate LLM prompts and RAG pipelines. Red-teaming, A/B testing, and CI integration for prompt quality.
Benchmark for evaluating LLMs as agents across 8 distinct environments including OS, web browser, and database tasks.
Open-source data labeling platform. Annotate text, images, audio, and video for training and evaluating AI models.
Framework for evaluating RAG pipelines. Measures faithfulness, answer relevancy, context recall, and more without ground truth.
Open-source data curation platform for LLMs. Collect human feedback, curate training data, and evaluate model outputs collaboratively.