Data extraction
Everything here was classified as data extraction by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Listed, not yet verified by us (60)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
io.github.HomenShum/nodebench
— v2.31.2
260 MCP tools across 49 domains. AI Flywheel, quality gates, research, web scraping.
-
AgentScrape
— v0.6.1
Pay-per-call web scraping for AI agents via x402 on Base USDC. Six tools, no signup.
-
HSH Data-on-Demand
— v2.1.0
Made-to-order data for AI agents: company intel, B2B contacts, scraping. Pay per call via x402.
-
io.github.huoshuiai42/huoshui-fetch
— v1.0.0
An MCP server that provides tools for fetching, converting, and extracting data from web pages.
-
io.github.ilo415/research
— v1.1.0
AI research for agents. Company intel, competitor analysis, web scraping, topic dives. Pay via MPP.
-
io.github.infoinlet-marketplace/mcp-browser-research
— v0.1.1
Web browsing & research for AI agents — fetch/read pages, structured extraction, Exa search.
-
io.github.infoinlet-marketplace/mcp-jp-doc-intel
— v0.1.0
Japanese document AI for agents — invoice/meishi extraction, OCR, classification (hosted API).
-
io.github.IntelagentStudios/mcp-file-processor
— v0.1.1
Text extraction, keyword extraction, language detection, and chunking for RAG
-
io.github.invapi-org/invapi-mcp
— v1.0.1
Extract, create, convert & validate invoices (UBL, CII, XRechnung, ZUGFeRD, PDF, Excel)
-
DicePDF
— v0.1.0
Local-first PDF tools for AI agents: merge, split, extract, render, compress. Files stay on-device.
-
io.github.IvanRublev/keyphrases-mcp
— v0.0.4
An MCP server to extract keyphrases from a text with the BERT model
-
io.github.jamjet-labs/engram-server
— v0.5.0
Durable memory for AI agents: fact extraction, hybrid retrieval, temporal graph. SQLite/Postgres.
-
ZenLink MCP
— v2.0.3
Browser automation for MCP clients via Zen Browser. Parallel multi-tab, content extraction.
-
io.github.jgador/websharp
— vv0.99.0-rc2
Search the web and extract article text for LLMs.
-
Secant Agent Research Pack
— v0.1.0
Paid MCP web research for agents: search, extraction, citations, and x402 payment discovery.
-
io.github.johnisanerd/baidu-search
— v1.0.0
Baidu search results and Chinese SERP data via the Apify Baidu Search Scraper, hosted MCP.
-
io.github.johnisanerd/google-jobs
— v1.0.0
Google Jobs listings with direct apply links via the Apify Google Jobs Scraper, hosted MCP.
-
io.github.johnisanerd/linkedin-posts
— v1.0.0
Scrape and analyze public LinkedIn posts as structured JSON via the Apify LinkedIn Posts API.
-
io.github.johnisanerd/yandex-search
— v1.0.0
Yandex search results, images, and SERP data via the Apify Yandex Search Scraper, hosted MCP.
-
TruePath PDF
— v0.3.0
Local-only PDF tools for AI — extract text, search, split/merge, render pages. Files never leave.
-
twscrape Twitter/X (read-only)
— v0.1.2
Read-only X/Twitter MCP: posts, threads, replies, search via twscrape. No API key.
-
io.github.jwulff/apple-voice-memo-mcp
— v0.1.1
Access Apple Voice Memos on macOS. List, get audio, extract and generate transcripts.
-
DB2TOON MCP
— v1.3.1
Read-only SQL database and plain-text SQL dump schema extraction to compact TOON for MCP clients.
-
footnote-mcp
— v0.2.6
Source-grounded web research: search, extraction, verification, and browser automation.
-
MCP Search Server
— v0.1.1
Web search, content extraction, PDF parsing, plus datetime and geolocation tools. No API keys.
-
io.github.khadinakbaronline/telegram-channel-scraper-mcp
— v1.0.0
Scrape messages, extract leads from public Telegram channels. No login required.
-
Argus Retrieval
— v1.6.2
Multi-provider search broker for AI agents: 14 providers, 12-step extraction, retrieval workflows.
-
io.github.kitewright/mcp
— v0.1.1
Lightweight browser automation for AI agents: navigate, screenshot, extract, PDF. One small binary.
-
io.github.koraykoylu/ibanchecker-mcp
— v1.1.2
IBAN validation, extraction, format specs and BIC/SWIFT lookup tools for AI assistants.
-
Kreuzcrawl
— v0.3.0
Scrape, crawl, and map websites to Markdown or JSON via local CLI.
-
Kloakt
— v0.1.2
Cloaked headless browser for AI agents — stealth TLS, smart extraction, SPA fallback
-
io.github.KylinMountain/web-fetch-mcp
— v0.1.1
MCP server for web content fetching, summarizing, comparing, and extracting information
-
Frenchie
— v0.5.0
OCR, transcription, file extraction, and image generation for AI agents via MCP.
-
io.github.Larshiensch99/priceparse
— v0.1.0
Extract structured pricing tiers and addons from any SaaS pricing page URL. Built for AI agents.
-
io.github.laundromatic/shopgraph
— v1.0.1
Clean product data from any URL. Schema.org + AI extraction. 200 free calls/month.
-
io.github.lazymac2x/smart-data-extractor
— v1.0.0
smart-data-extractor MCP server on Cloudflare Workers · REST + MCP JSON-RPC · free tier
-
io.github.lesofi/handinloop-mcp-server
— v0.4.0
Human-in-the-loop document extraction: submit a doc, get validated fields and an audit trail.
-
io.github.Libres-coder/parseflow
— v1.0.2
PDF parsing server with text extraction, metadata, search, images, and TOC via MCP
-
Licium
— v2.0.0
Clean rows from public pages that break ordinary scrapers, plus a bounty board where agents earn.
-
Syftly
— v0.1.0
Ranks the best AI tool or API per task: transcription, TTS, web search, scraping and OCR.
-
io.github.lordolami/poldex
— v0.0.3
MCP server exposing PolDex commercial insurance extraction tools to AI agents.
-
Refetch
— v1.0.0
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
-
io.github.malkreide/openlex-mcp
— v0.2.5
Canton Zurich legislation via ZH-Lex with full-text search and article extraction
-
TheCrawler
— v0.5.8
Universal web scraper with LLM-ready markdown, RAG chunking, PDF/DOCX support.
-
io.github.massanaRoger/extracto-mcp
— v0.1.2
Turn any URL plus a schema into validated, typed JSON via the Extracto API.
-
dotrepo
— v1.0.1
Trust-aware repository facts for agents: build, test, docs, license, and security, no scraping.
-
io.github.mehtaphysical13/pdf-tables-mcp
— v0.1.0
Reliable PDF table extraction. Pass a URL, get structured JSON tables with citations.
-
io.github.meltingpixelsai/harvey-tools
— v1.1.0
Web scraping, code review, content generation, and analysis tools. Pay-per-call USDC.
-
io.github.meltingpixelsai/zero-core-tools
— v1.0.0
Web scraping, code review, content gen, sentiment. Zero Core Tools.
-
Obscura MCP
— v0.1.1
MCP server adapter for Obscura Rust headless browser — web scraping with anti-detection.
-
io.github.MetaLift-AI/metalift
— v1.0.15
Metalift MCP for AI agents: web search (2 credits), scrape, crawl, and map tools
-
io.github.mgriffen/tapsite
— v4.0.1
MCP server for web inspection, design system extraction, and accessibility audits
-
io.github.mifactory-bot/mifactory-scraping-api
— v1.2.0
Web scraping for AI agents. Extract text and metadata from any URL worldwide. $0.005/page.
-
Data Enrichment & Web Scraper API
— v1.1.0
Enrich IPs, emails, domains, companies. Scrape SEC, news, jobs, Crunchbase. Bitcoin pay-per-use.
-
Toolora MCP Server
— v1.0.0
12 free tools: PDF, OCR, QR codes, audio transcription, URL scraping, Excel, Word. No key needed.
-
Scribefy
— v0.3.5
Search YouTube, get video metadata, and extract timestamped transcripts for AI workflows.
-
io.github.mlawsonking/web-tools-mcp
— v1.2.0
10 no-LLM web tools: URL to Markdown, metadata, JSON-LD, MX, scrape, RSS, DNS, RDAP, SSL, HTTP.
-
PyScrappy
— v1.4.4
Web-scraping toolkit with 22 tools for structured web data as JSON for AI agents.
-
io.github.mleoca/ucn
— v4.2.3
Code intelligence for AI agents. Extract, trace, and analyze code without reading whole files.
-
DesignScan
— v0.4.0
Extract a website’s design tokens (colors, type, spacing, shadows) as DESIGN.md, W3C, or CSS.
Page 4 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.