Extractous
Fast content extraction from files and URLs in Rust. Supports 1400+ file formats via Apache Tika with Python and JS bindings.
Fast content extraction from files and URLs in Rust. Supports 1400+ file formats via Apache Tika with Python and JS bindings.
Turn any website into LLM-ready data. Crawls, scrapes, and converts web pages into clean Markdown at scale.
IBM open-source document parser. Converts PDFs, Word, and PowerPoint into clean Markdown or JSON with layout understanding.
Open-source data integration platform. Move data from 300+ sources into your vector store or data warehouse for RAG pipelines.
Library of data loaders for LlamaIndex. Connects agents to 300+ sources: Notion, Google Drive, Slack, GitHub, databases.
OCR API for converting scientific documents, equations, and handwriting into LaTeX, Markdown, or structured JSON.
Open-source web crawler optimized for AI. Extracts structured data from any website asynchronously, built for LLM pipelines.