AgentBench

EvaluationBenchmark

Benchmark for evaluating LLMs as agents across 8 distinct environments including OS, web browser, and database tasks.

Visit website View on GitHub
Built with
PythonLLM
Share this part