Open-source data curation platform for LLMs. Collect human feedback, curate training data, and evaluate model outputs collaboratively.
Holistic Evaluation of Language Models. Evaluates LLMs across 42 scenarios and 7 metrics for a complete capability picture.
Framework for evaluating RAG pipelines. Measures faithfulness, answer relevancy, context recall, and more without ground truth.
Beyond the Imitation Game benchmark. 200+ tasks designed to probe LLM capabilities beyond standard NLP benchmarks.
Open-source LLM evaluation framework. Unit-test your LLM outputs with 14+ metrics including hallucination, toxicity, and RAG.
Test and evaluate LLM prompts and RAG pipelines. Red-teaming, A/B testing, and CI integration for prompt quality.
Benchmark for evaluating LLMs as agents across 8 distinct environments including OS, web browser, and database tasks.
Open-source data labeling platform. Annotate text, images, audio, and video for training and evaluating AI models.
Open platform for evaluating LLMs through blind pairwise comparisons. The de-facto LLM community ranking.
Framework for evaluating LLMs and LLM systems. Define custom evaluation tasks and run them against any model.