Data extraction
Everything here was classified as data extraction by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Worked on in the last 90 days (19)
Someone pushed a commit recently. This is the shortlist worth trying first.
-
ClicheFactory Document Intelligence
— last commit 2026-06-17, v0.1.8
Extract structured JSON from PDFs, images, DOCX, XLSX, CSV, EML attachments, and DSPy pipelines.
-
io.github.CSOAI-ORG/keyword-extractor-ai-mcp
— last commit 2026-06-15, v1.0.4
keyword-extractor-ai-mcp MCP server by MEOK AI Labs
-
IRONCLAW BTC Node
— last commit 2026-06-12, v1.0.1
Real BTC full node: fees, mempool, txs, portfolio, trace, whales, SEC, scraping, Reddit via x402.
-
GST Validator
— last commit 2026-06-11, v0.1.1
Validate Indian GSTINs locally (Verhoeff), extract PAN, identify issuing state.
-
ai.smithery/oxylabs-oxylabs-mcp
— last commit 2026-06-08, v1.13.1
Fetch and process content from specified URLs using the Oxylabs Web Scraper API.
-
ShadowCrawl
— last commit 2026-06-07, v2.0.0-rc
Stealth scraping & search. Bypasses Cloudflare, DataDome & LinkedIn via Cyborg HITL approach.
-
ShadowCrawl
— last commit 2026-06-07, v2.3.0
Rust MCP stealth scraper: anti-bot search/scrape with CDP fallback + HITL non-robot.
-
io.github.Godalo-ai/godalo
— last commit 2026-06-04, v0.0.3
Affiliate product search for AI agents. Merchant feeds, not web scraping. Works with any MCP client.
-
Amazon Scraper API
— last commit 2026-06-02, v0.1.4
Scrape Amazon products, search, and async batch ASIN lookups across 20 marketplaces
-
io.github.bamchi/scrapi
— last commit 2026-05-27, v2.0.6
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
-
io.github.erikirby/web-network-tools
— last commit 2026-05-20, v1.0.0
DNS lookup, WHOIS, SSL checker, and HTTP status tools. Uses public APIs only, no scraping.
-
com.mcparmory/agentql
— last commit 2026-05-12, v1.0.8
Query webpages and extract structured data using natural language
-
com.mcparmory/apify
— last commit 2026-05-12, v1.0.5
Build, deploy, and run web scraping and automation actors in the cloud
-
com.mcparmory/firecrawl
— last commit 2026-05-12, v1.0.3
Scrape, crawl, and extract structured data from websites at scale
-
com.mcparmory/google-search
— last commit 2026-05-12, v1.0.2
Scrape Google search results with SERP data, ads, and knowledge panels
-
com.mcparmory/parallel
— last commit 2026-05-12, v1.0.2
Search the web, extract content, and run distributed tasks with AI-powered automation
-
com.mcparmory/pdfco
— last commit 2026-05-12, v1.0.2
Extract data, edit, convert, and parse PDF documents with OCR and AI
-
com.mcparmory/scrapingant
— last commit 2026-05-12, v1.0.1
Scrape webpages, handle JavaScript and CAPTCHA, extract structured data
-
OpenBrand
— last commit 2026-05-12, v0.1.2
Extract brand assets (logos, colors, backdrop images, brand name) from any website URL
Exists, but quiet (23)
No commit in the last 90 days. Often fine for something small and finished, risky for anything you need fixed.
-
Afterpaths
— last commit 2026-04-27, v0.2.5
Session memory for AI coding agents. Search past sessions, extract rules, and track what works.
-
UnCorreoTemporal
— last commit 2026-04-19, v0.1.1
Temporary email for AI agents: create inboxes, wait for emails, extract OTPs, verify signups.
-
Docpick
— last commit 2026-04-15, v0.1.2
Schema-driven document extraction with local OCR + LLM. Document in, Structured JSON out.
-
MarkGrab
— last commit 2026-04-15, v0.1.2
Universal web content extraction — any URL to LLM-ready markdown. HTML, YouTube, PDF, DOCX.
-
AIR SDK
— last commit 2026-04-10, v0.2.12
Collective intelligence for browser automation agents. Site capabilities, selectors, and extraction.
-
LionScraper MCP + CLI + HTTP API (Node)
— last commit 2026-04-08, v1.0.6
LionScraper Node: MCP stdio + CLI + HTTP API bridging AI to LionScraper browser extension.
-
LionScraper MCP + CLI + HTTP API (Python)
— last commit 2026-04-08, v1.0.6
LionScraper Python: MCP stdio + CLI + HTTP API bridging AI to LionScraper browser extension.
-
ai.smithery/rainbowgore-stealthee-mcp-tools
— last commit 2026-04-05, v1.14.0
Spot pre-launch products before they trend. Search the web and tech sites, extract and parse pages…
-
AiPayGen — 65+ AI Tools as an MCP Server
— last commit 2026-04-03, v1.9.7
65+ AI tools as MCP: research, write, code, scrape, translate, RAG, agent memory, workflows
-
io.github.agenson-horrowitz/document-parser
— last commit 2026-04-02, v1.0.8
Parse and extract structured data from various document formats (PDF, Word, HTML).
-
io.github.agenson-horrowitz/web-content-extractor
— last commit 2026-04-02, v1.0.8
Extract and process web content into clean, structured formats optimized for LLMs.
-
io.github.blipemail/email
— last commit 2026-03-24, v0.1.3
MCP server for Blip disposable email — create inboxes, receive emails, extract OTP codes
-
io.github.BrightWayAI/video-analyzer
— last commit 2026-03-10, v0.1.2
Analyze videos: extract frames, transcribe audio, generate storyboard breakdowns.
-
io.github.fredpsantos33/iteratools
— last commit 2026-03-06, v1.0.5
40+ pay-per-use tools for AI agents: search, TTS, QR, PDF, scraping, image gen. x402.
-
io.github.gigabrain-observer/google-docs-mcp-server
— last commit 2026-03-02, v0.1.1
Google Docs MCP server with full tab support, markdown extraction, and batch updates.
-
ESG MCP Servers
— last commit 2026-02-28, v0.1.2
31 MCP tools for ESG data extraction, PDF processing, vector search, and EU regulation analysis.
-
io.github.baixianger/camoufox-mcp
— last commit 2026-02-20, v1.0.0
Anti-detection browser automation with Camoufox - stealth Firefox for web scraping
-
io.github.Fato07/log-analyzer-mcp
— last commit 2026-01-16, v0.4.2
AI-powered log analysis - parse, search, extract errors across 9+ formats
-
ai.smithery/huuthangntk-claude-vision-mcp-server
— last commit 2025-10-08, v1.0.0
Analyze images from multiple angles to extract detailed insights or quick summaries. Describe visu…
-
ai.smithery/blacklotusdev8-test_m
— last commit 2025-09-26, v1.14.0
Greet anyone by name with a friendly hello. Scrape webpages to extract content for quick reference…
-
ai.smithery/BadRooBot-test_m
— last commit 2025-09-20, v1.14.0
Send quick greetings, scrape website content, and generate text or images on demand. Perform web s…
-
ai.smithery/ImRonAI-mcp-server-browserbase
— last commit 2025-09-12, v2.0.0
Automate cloud browsers to navigate websites, interact with elements, and extract structured data.…
-
ai.smithery/kwp-lab-rss-reader-mcp
— last commit 2025-09-10, v1.0.0
Track and browse RSS feeds with ease. Fetch the latest entries from any feed URL and extract full…
Listed, not yet verified by us (13)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
Reka
— v0.1.10
Understand your videos with Reka AI — search, ask questions, and extract insights.
-
Voxplo
— v0.1.1
Give AI agents a phone: outbound AI calls that return a summary, transcript, and extracted fields.
-
Prowl MCP
— v1.4.0
MCP server: 360+ pay-as-you-go research tools (SEO, ads, SERP, scraping) + prowl_analyze reports
-
Spider Cloud
— v1.0.0
Crawl, scrape, search the web, and automate browsers at scale with anti-bot bypass.
-
Sofya
— v1.27.0
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
-
AgenticTotem Web Extractor
— v1.0.5
AI web extraction: send URLs + a JSON Schema, get clean structured data. Pay-per-use via x402.
-
AllPDFMagic
— v1.0.0
PDF tools + invoice extraction, bank statement parsing, GST reconciliation & GSTIN validation.
-
Audioscrape Audio Intelligence
— v1.0.3
Search speech in podcasts, government meetings, and your own audio: speakers, entities, timestamps.
-
BedrockNews
— v1.0.0
GRIN-scored news for agents: discover by verdict, pull full analysis, traverse extraction graphs.
-
com.borisinc/tools
— v1.0.0
Pay-per-call AI tools over x402: web research, summarization, structured extraction (USDC, Base).
-
DeckExtract
— v1.0.0
Download DocSend and Papermark decks as PDF or PPTX, including email-gated and protected links.
-
IDLE Protocol
— v1.0.5
Web scraping, inference, DNS, monitoring from residential IPs — pay per request in USDC
-
com.geekflare/mcp
— v0.3.6
Geekflare MCP server for scraping, search, screenshots, broken links, DNS lookup, and network tools.
Repository gone or archived (5)
The registry still lists these, but the linked repository returns 404 or the owner archived it. Shown so you do not spend time discovering that yourself.
-
ai.filegraph/document-processing
— repository gone, v1.0.1
Extract text from documents, manipulate PDFs, and perform OCR on images.
-
ai.smithery/arjunkmrm-fetch
— repository gone, v1.0.0
Fetch web pages and extract exactly the content you need. Select elements with CSS and retrieve co…
-
ai.smithery/arjunkmrm-scrapermcp_el
— repository gone, v1.15.0
Extract and parse web pages into clean HTML, links, or Markdown. Handle dynamic, complex, or block…
-
ai.smithery/luminati-io-brightdata-mcp
— repository gone, v1.0.0
One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.…
-
com.docimprint/api
— repository gone, v1.0.4
AI document intelligence: extract, summarize, claim-check, notarize, and signed action receipts.
Page 2 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.