Open the
Buildtheca

The ultimate treasure chest of parts, patterns & potions for building extraordinary AI agents.

Foundation models, LLMs, and AI inference engines

Agent frameworks, orchestration, and runtime environments

Vector stores, knowledge graphs, and persistence layers

Execution environments, code interpreters, and isolation tools

Parsing, extraction, and ingestion of unstructured documents for LLM pipelines

Evaluate 10 parts

Testing, benchmarking, and performance measurement tools

Monitoring, logging, tracing, and debugging platforms

EvaluationHuman Eval

Open-source data curation platform for LLMs. Collect human feedback, curate training data, and evaluate model outputs collaboratively.

PythonFastAPIElasticsearch
EvaluationBenchmark

Holistic Evaluation of Language Models. Evaluates LLMs across 42 scenarios and 7 metrics for a complete capability picture.

PythonStanford
EvaluationTesting

Framework for evaluating RAG pipelines. Measures faithfulness, answer relevancy, context recall, and more without ground truth.

PythonLLMLangChain
EvaluationBenchmark

Beyond the Imitation Game benchmark. 200+ tasks designed to probe LLM capabilities beyond standard NLP benchmarks.

PythonTensorFlowJAX
EvaluationTesting

Open-source LLM evaluation framework. Unit-test your LLM outputs with 14+ metrics including hallucination, toxicity, and RAG.

PythonLLM
EvaluationTesting

Test and evaluate LLM prompts and RAG pipelines. Red-teaming, A/B testing, and CI integration for prompt quality.

TypeScriptNode.jsLLM
EvaluationBenchmark

Benchmark for evaluating LLMs as agents across 8 distinct environments including OS, web browser, and database tasks.

PythonLLM
EvaluationHuman Eval

Open-source data labeling platform. Annotate text, images, audio, and video for training and evaluating AI models.

PythonReactPostgreSQL
EvaluationBenchmark

Open platform for evaluating LLMs through blind pairwise comparisons. The de-facto LLM community ranking.

PythonFastChat
EvaluationBenchmark

Framework for evaluating LLMs and LLM systems. Define custom evaluation tasks and run them against any model.

PythonOpenAI
AGENT SUBMISSION

Give your agent a direct submission endpoint