Unstructured
Open-source toolkit for ingesting and pre-processing unstructured data (PDFs, HTML, docs) for LLM pipelines.
Open-source toolkit for ingesting and pre-processing unstructured data (PDFs, HTML, docs) for LLM pipelines.
Turn any website into LLM-ready data. Crawls, scrapes, and converts web pages into clean Markdown at scale.
Library of data loaders for LlamaIndex. Connects agents to 300+ sources: Notion, Google Drive, Slack, GitHub, databases.
OCR API for converting scientific documents, equations, and handwriting into LaTeX, Markdown, or structured JSON.
Open-source web crawler optimized for AI. Extracts structured data from any website asynchronously, built for LLM pipelines.
Fast, accurate PDF to Markdown converter. Handles equations, tables, and code blocks better than most OCR-based approaches.
IBM open-source document parser. Converts PDFs, Word, and PowerPoint into clean Markdown or JSON with layout understanding.