Data extraction
Everything here was classified as data extraction by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. This makes current maintenance visible; it does not prove quality or fit for your task.
Listed, not yet verified by us (45)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
MailFixture
— needs an account, v0.1.0
Test-inbox API for email and SMS: create inboxes, long-poll messages, extract OTPs and links.
-
com.mart402/extract
— 23 functions, v0.2.0
Web extraction, OCR (Japanese-strong), invoice and company data for AI agents. x402, no API key.
-
Scraper API
— needs an account, v0.1.0
One API for public web data across social, directories and real estate, as clean JSON.
-
com.scrapingant/mcp
— 3 functions, v1.0.0
Cloud-based web access with real browsers and JS rendering by ScrapingAnt
-
com.scraptik/tiktok
— 19 functions, v0.2.1
Unofficial TikTok API & scraper: creator analytics, video data, comments, search. x402, no API key.
-
Docu-Scan MCP
— needs an account, v1.0.0
PDF and document extraction via Google Document AI. Free trial available.
-
stagenth · 文档解析
— 4 functions, v1.0.0
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
-
stagenth · PDF 工具箱
— 14 functions, v1.1.0
14 PDF tools: merge, split, compress, watermark, encrypt, organize, metadata, page extraction.
-
stagenth · 网页数据
— 4 functions, v1.0.0
Web scraping to clean Markdown with JS rendering, multi-page crawl, structured extract, sitemaps.
-
com.thenextgennexus/energy-commodity-mcp
— needs an account, v1.0.0
DISCONTINUED 2026-05-23. Indeed.com Cloudflare + DataDome anti-bot returns HTTP 403 on all scraping
-
com.thenextgennexus/environmental-esg-mcp
— needs an account, v1.0.0
⚠️ DEPRECATED — G2's Cloudflare anti-bot defeats free scraping. Not commercially viable without paid
-
com.thenextgennexus/google-maps-mcp-server
— 3 functions, v1.0.0
Local business lead extraction with email + phone enrichment from Google Maps.
-
com.thenextgennexus/healthcare-fda-intelligence-mcp
— needs an account, v1.0.0
Scrape the FAA Aircraft Registry by N-number — aircraft make, model, year, owner name, address, airw
-
com.thenextgennexus/review-intelligence-mcp-server
— 3 functions, v1.0.0
G2, Trustpilot, Yelp reviews with sentiment and theme extraction across sources.
-
com.thenextgennexus/sec-corporate-events-mcp
— needs an account, v1.0.0
Extract papers from ArXiv — titles, abstracts, authors, categories & PDF links. Monitor new AI, phys
-
com.thenextgennexus/web-scraping-mcp-server
— 3 functions, v1.0.0
Generic URL crawl + HTML extraction — fallback for sites without dedicated MCPs.
-
com.thenextgennexus/youtube-media-mcp-server
— 3 functions, v1.0.0
YouTube video search with transcript extraction as first-class output.
-
com.vevang/agents
— 6 functions, v1.0.0
Hire Vevang's AI agents, pay-per-call in USDC on Base via x402: video, visibility, verify, extract
-
reapx App Store reviews - ratings and review history
— needs an account, v1.0.0
Scrape App Store reviews and rating history by app, country, rating or date. Pay per row.
-
reapx books - Open Library editions and authors
— needs an account, v1.0.0
Scrape Open Library book editions, authors, subjects and identifiers. Pay per row.
-
reapx clinical trials - ClinicalTrials.gov study records
— needs an account, v1.0.0
Scrape ClinicalTrials.gov studies by condition, sponsor, phase, status or location. Pay per row.
-
reapx crypto - CoinGecko coin prices and markets
— needs an account, v1.0.0
Scrape CoinGecko coin prices, market caps, volumes and exchange listings. Pay per row.
-
reapx music - Discogs releases, artists and labels
— needs an account, v1.0.0
Scrape Discogs music releases, artists, labels, formats and catalogue numbers. Pay per row.
-
reapx Docker containers - Hub images, tags and pulls
— needs an account, v1.0.0
Scrape Docker Hub container images, tags, pulls and publisher metadata. Pay per row.
-
reapx Federal Register - rules, notices and orders
— needs an account, v1.0.0
Scrape Federal Register rules, proposed rules, notices and orders by agency or date. Pay per row.
-
reapx GitHub repos - repository metadata and activity
— needs an account, v1.0.0
Scrape GitHub repository metadata, stars, forks, topics, licences and activity. Pay per row.
-
reapx jobs - listings from Greenhouse, Workable, Reed and more
— needs an account, v1.0.0
Scrape job listings from Greenhouse, Workable, Reed, RemoteOK and other boards. Pay per row.
-
reapx fediverse - Mastodon posts and instance timelines
— needs an account, v1.0.0
Scrape Mastodon posts, profiles, hashtags and instance timelines across the fediverse. Pay per row.
-
reapx npm packages - registry metadata and downloads
— needs an account, v1.0.0
Scrape npm package metadata, versions, downloads, dependencies and maintainers. Pay per row.
-
reapx openFDA - drug recalls, labels and adverse events
— needs an account, v1.0.0
Scrape openFDA drug recalls, enforcement reports, labels and adverse events. Pay per row.
-
reapx scientific papers - arXiv, OpenAlex and Crossref
— needs an account, v1.0.0
Scrape arXiv, OpenAlex and Crossref papers by author, topic, journal or DOI. Pay per row.
-
reapx SEC filings - EDGAR company filings and full-text search
— needs an account, v1.0.0
Scrape SEC EDGAR filings by company, form type, date, or full text. Pay per row.
-
reapx Shopify - store products and app listings
— needs an account, v1.0.0
Scrape Shopify store products, variants, prices and app store listings. Pay per row.
-
reapx Stack Overflow - questions, answers and tags
— needs an account, v1.0.0
Scrape Stack Overflow questions, answers, tags, scores and accepted status. Pay per row.
-
ScrapeNest
— needs an account, v1.0.0
Web search, scraping, RAG answers with citations, and translation as MCP tools.
-
toolvend - DNS, WHOIS & domain tools
— 10 functions, v0.1.2
DNS, WHOIS/RDAP, DMARC/SPF, LEI, sitemap, web extract, VAT, QR. Paid per call in USDC, no signup.
-
Caliper
— 10 functions, v0.1.1
Geometry and CAD file metadata extraction for STL, OBJ, PLY, PCD, LAS/LAZ, glTF/GLB.
-
webclaw
— v0.6.17
Turn any URL into clean markdown/JSON for AI agents. Self-hostable web content extraction.
-
Scrappycoco
— needs an account, v1.0.0
Route web, social, and filings data tasks to the best available scraper.
-
PDF Kit
— v1.0.2
AI-powered PDF tools: fill forms, merge, extract data, and split PDFs
-
io.github.Cal-Dev-Tech/perpage-mcp-server
— v0.1.3
Give AI agents clean, LLM-ready web data — scrape any URL to markdown or extract structured JSON.
-
io.github.CedricKouma/facturx-toolkit
— needs an account, v0.1.0
Validate, extract, repair and generate French Factur-X / EN16931 invoices via AgentForge API
-
DocRocket
— needs an account, v1.0.0
Generate on-brand proposals, reports, and contracts instantly. Auto-extracts brand from any URL.
-
Zyntra - Temp e-mails MCP
— needs an account, v1.0.0
Disposable inboxes for AI agents: create, wait for delivery, and extract email content or links.
-
Agent Utility API
— 8 functions, v1.1.0
x402 pay-per-call tools: company enrichment, PDF extraction, Amazon/KDP data, YouTube transcripts.
Repository gone or archived (15)
The registry still lists these, but the linked repository returns 404 or the owner archived it. Shown so you do not spend time discovering that yourself.
-
Iteration Layer
— repository gone, endpoint not answering, v1.0.0
Composable APIs for document extraction, image transformation, and document & sheet generation.
-
Text2Event
— repository gone, needs an account, v1.0.0
Extracts calendar events from natural-language text, with .ics and calendar links.
-
Wauldo
— repository gone, needs an account, v1.1.0
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
-
dev.urlsnap/urlsnap-mcp
— repository gone, v1.0.2
Screenshot, PDF, and markdown extraction from any URL via Claude Desktop and Claude Code.
-
io.github.AceDataCloud/mcp-webextrator
— repository gone, needs an account, v2026.8.1.0
MCP server for web extraction and rendering via AceDataCloud WebExtrator
-
Archiet
— repository gone, v0.1.1
Archiet MCP server: scaffold full-stack apps from PRD text and extract capabilities.
-
io.github.atakanelik34/neo-x402-mcp
— repository gone, endpoint not answering, v1.2.0
8-tool AI web intelligence suite: search, scrape, screenshot, SEO, docs, crypto, code.
-
ClipForge
— repository gone, needs an account, v1.0.0
Trim, watermark, extract audio, and convert video to 9:16 vertical via API or MCP tools.
-
io.github.CodyWatters/riveter
— repository gone, v0.1.2
MCP server for Riveter's enrichment, scraping, and monitoring API
-
io.github.CSOAI-ORG/meok-ai-reflection-mcp
— repository gone, v1.0.7
MEOK AI Labs - ai-reflection MCP server extracted from SOV3
-
io.github.CSOAI-ORG/meok-bft-governance-mcp
— repository gone, v1.0.7
MEOK AI Labs - bft-governance MCP server extracted from SOV3
-
io.github.CSOAI-ORG/meok-neural-health-monitor-mcp
— repository gone, v1.0.4
MEOK AI Labs - neural-health-monitor MCP server extracted from SOV3
-
io.github.CSOAI-ORG/meok-quantum-scoring-mcp
— repository gone, v1.0.5
MEOK AI Labs - quantum-scoring MCP server extracted from SOV3
-
LinkMeta MCP
— repository gone, v1.2.0
Extract URL metadata (Open Graph, Twitter Cards, JSON-LD) from any URL
-
Erudite Intelligence x402 Services
— repository gone, endpoint not answering, v1.0.0
80 paid MCP tools via x402 USDC micropayments. Crypto data, web scraping, auditing, and more.
Page 3 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.