BIG-Bench

EvaluationBenchmark

Beyond the Imitation Game benchmark. 200+ tasks designed to probe LLM capabilities beyond standard NLP benchmarks.

Visit website View on GitHub
Built with
PythonTensorFlowJAX
Share this part