AI for systematic literature review — search 125 million papers, extract data into tables, and screen at scale.
Data Extraction
Pulling structured records out of documents, web pages and other unstructured sources.
17 apps, 3 skills and 5 MCP servers tagged Data Extraction.
Apps
Turn any website into clean markdown or structured JSON — crawling, JavaScript rendering, and extraction in one API.
Web data extraction APIs for AI agents — 75+ ready-made scrapers for LinkedIn, Amazon, Google Maps, Reddit and X, priced per result.
Deep-research AI agent that builds cited reports or structured datasets, scaling depth to your budget.
One API to search, scrape, crawl, map and monitor the web, returning clean structured text that AI agents and RAG pipelines can use directly.
Construction takeoffs, estimates and invoices in one browser workflow — upload the PDF plans, measure on scale, and let Onyx AI draft the numbers.
Describe the data you want in plain English and BrowserAct builds a reusable scraper that runs in a real browser, with proxies and CAPTCHA handled.
Rindler signs in to the web portals your team already uses and finishes the task you describe in a sentence — pulling records, downloading invoices, checking statuses, submitting forms.
Cloud browser automation for AI agents: describe a workflow in plain English and Airtop compiles it into a deterministic agent that logs in, browses and acts at scale.
One REST API that scrapes, crawls, and extracts the web into clean Markdown and schema-validated JSON for AI agents.
AI built for the rigour of finance — run structured analysis across thousands of documents with auditable answers.
Collaborative legal AI for review, research, and drafting — built to work inside the documents lawyers already use.
A marketplace of thousands of ready-made scrapers and automation tools, plus the cloud to run your own.
ETL for unstructured data — turn PDFs, slides, emails, and images into clean, AI-ready structured output.
A spreadsheet that enriches itself — pull data from 100+ sources, research with AI, and run go-to-market plays.
A B2B database of hundreds of millions of contacts plus the sequencing, dialling, and AI to work them.
Mathpix converts math and science from images to LaTeX.
Skills
Skill: PaddleOCR Document Parsing
by PaddlePaddle
Turn PDFs and scanned images into structured Markdown/JSON: cell-level tables, LaTeX formulas, correct multi-column order.
Skill: Spreadsheets (XLSX)
by Anthropic
Open, create, and edit spreadsheets — formulas, formatting, charts, data cleaning, and CSV/TSV conversion.
Skill: PDF
by Anthropic
Read, create, and manipulate PDF files — extract text and tables, merge, split, watermark, fill forms, encrypt, and OCR scanned pages.
MCP servers
Apify's official MCP server — turn thousands of ready-made scrapers into tools your assistant can call.
Official Crawlbase MCP server giving agents live web access — fetch any URL as raw HTML, clean Markdown or a screenshot, with JS rendering and proxy rotation handled for you.
MCP: Tavily
by Tavily
Web search and content extraction optimized for AI agents — fast, relevance-ranked results with source content.
MCP: Firecrawl
by Firecrawl
Search, scrape, crawl, and extract structured data from the web — with JS rendering, batch scraping, and LLM-powered extraction.
MCP: Fetch
by Model Context Protocol
Fetch a URL and convert its contents to clean Markdown for efficient LLM consumption.
Related tags
Tags that appear alongside this one, ranked by how often.