OCR API for converting scientific documents, equations, and handwriting into LaTeX, Markdown, or structured JSON.
Turn any website into LLM-ready data. Crawls, scrapes, and converts web pages into clean Markdown at scale.
Fast content extraction from files and URLs in Rust. Supports 1400+ file formats via Apache Tika with Python and JS bindings.
Open-source data integration platform. Move data from 300+ sources into your vector store or data warehouse for RAG pipelines.
IBM open-source document parser. Converts PDFs, Word, and PowerPoint into clean Markdown or JSON with layout understanding.
Fast, accurate PDF to Markdown converter. Handles equations, tables, and code blocks better than most OCR-based approaches.
Open-source web crawler optimized for AI. Extracts structured data from any website asynchronously, built for LLM pipelines.
Library of data loaders for LlamaIndex. Connects agents to 300+ sources: Notion, Google Drive, Slack, GitHub, databases.
Open-source toolkit for ingesting and pre-processing unstructured data (PDFs, HTML, docs) for LLM pipelines.
GenAI-native document parsing API by LlamaIndex. Extracts text, tables, and structure from any document format.