OpenAI Evals

EvaluationBenchmark

Framework for evaluating LLMs and LLM systems. Define custom evaluation tasks and run them against any model.

Visit website View on GitHub
Built with
PythonOpenAI
Share this part