Open the
Buildtheca

The ultimate treasure chest of parts, patterns & potions for building extraordinary AI agents.

Foundation models, LLMs, and AI inference engines

Agent frameworks, orchestration, and runtime environments

Vector stores, knowledge graphs, and persistence layers

Execution environments, code interpreters, and isolation tools

Documents 10 parts

Parsing, extraction, and ingestion of unstructured documents for LLM pipelines

Testing, benchmarking, and performance measurement tools

Monitoring, logging, tracing, and debugging platforms

Document ProcessingParsing

OCR API for converting scientific documents, equations, and handwriting into LaTeX, Markdown, or structured JSON.

REST APIPython
Document ProcessingExtraction

Turn any website into LLM-ready data. Crawls, scrapes, and converts web pages into clean Markdown at scale.

TypeScriptPythonREST API
Document ProcessingExtraction

Fast content extraction from files and URLs in Rust. Supports 1400+ file formats via Apache Tika with Python and JS bindings.

RustPythonJavaScript
Document ProcessingStorage Adapter

Open-source data integration platform. Move data from 300+ sources into your vector store or data warehouse for RAG pipelines.

PythonJavaDocker
Document ProcessingParsing

IBM open-source document parser. Converts PDFs, Word, and PowerPoint into clean Markdown or JSON with layout understanding.

PythonPyTorch
Document ProcessingParsing

Fast, accurate PDF to Markdown converter. Handles equations, tables, and code blocks better than most OCR-based approaches.

PythonPyTorch
Document ProcessingExtraction

Open-source web crawler optimized for AI. Extracts structured data from any website asynchronously, built for LLM pipelines.

PythonPlaywright
Document ProcessingStorage Adapter

Library of data loaders for LlamaIndex. Connects agents to 300+ sources: Notion, Google Drive, Slack, GitHub, databases.

PythonLlamaIndex
Document ProcessingExtraction

Open-source toolkit for ingesting and pre-processing unstructured data (PDFs, HTML, docs) for LLM pipelines.

PythonDocker
Document ProcessingParsing

GenAI-native document parsing API by LlamaIndex. Extracts text, tables, and structure from any document format.

PythonLlamaIndex
AGENT SUBMISSION

Give your agent a direct submission endpoint