Data extraction
Everything here was classified as data extraction by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Listed, not yet verified by us (60)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
io.github.mshegolev/prometheus-mcp
— v0.1.0
Prometheus MCP — query metrics, inspect alerts, and explore scrape targets (read-only).
-
io.github.MukundaKatta/agentcast
— v0.1.0
Structured-output enforcer: extract and validate JSON from messy LLM text.
-
io.github.MukundaKatta/html-to-markdown-mcp
— v0.1.0
Convert HTML to Markdown or strip to plain text. For web-scraping agents.
-
io.github.mysleekdesigns/crawlforge-mcp-server
— v5.0.0
Web scraping, crawling, deep research & autonomous extraction — 27 MCP tools, clean Markdown/JSON
-
io.github.n24q02m/wet-mcp
— v3.6.0
MCP server for web search, content extraction, academic research, and library docs.
-
pdfmux
— v1.8.7
PDF-to-Markdown extraction that audits its own output and flags any extractor's silent drops.
-
Documentation Extractor
— v1.0.0
Extract llms.txt from any docs site - Mintlify Docusaurus GitBook parser
-
XActions
— v3.4.4
X/Twitter automation MCP: scrape, post, schedule, analyze, engage. No API key required.
-
io.github.Nizoka/pdfnative-mcp
— v1.5.0
PDF native MCP server: generate, validate, sign PAdES, embed, extract. AI-powered.
-
SchemaSure — structured extraction
— v0.3.0
Extract schema-valid JSON from text/HTML or document images. Pay per call via x402 V2.
-
Colors-LE
— v2.2.2
Extract colors from stylesheets and code, with their notation and position.
-
Dates-LE
— v2.2.2
Extract dates and timestamps from logs, data files and code, with their format and position.
-
Numbers-LE
— v2.2.2
Extract numeric values from config files, data files and plain text.
-
Paths-LE
— v2.2.2
Extract file and directory paths from config files and code, with their kind and position.
-
Regex-LE
— v2.2.2
Extract regular expressions from code, each with a ReDoS safety verdict.
-
Scrape-LE
— v2.2.2
Analyse robots.txt content and report whether a path may be crawled.
-
Strings-LE
— v2.2.2
Extract string values from config files, data files and plain text.
-
URLs-LE
— v2.2.2
Extract URLs from documentation, configuration and code, with their protocol and position.
-
io.github.ofershap/markdown
— v1.0.1
Markdown MCP — search, extract sections, list headings, find code blocks.
-
io.github.ofershap/scraper
— v1.0.1
Web scraping MCP — extract clean markdown, links, and metadata from any URL.
-
io.github.olamide-olaniyan/sociavault-mcp
— v2.0.0
Scrape social data from TikTok, Instagram, YouTube, X, LinkedIn, Reddit, and ad libraries.
-
Olostep MCP Server
— v1.0.17
Search, scrape, and crawl the web for AI agents. Batch scraping and answers with citations.
-
Alembica MCP
— v0.3.4
MCP server for Alembica validation, extraction, cost estimation, and schema queries.
-
io.github.openagentemail/mcp
— v0.4.0
Unlimited agent mailboxes — create identities, read/wait for mail, extract OTPs, send email.
-
io.github.Optisol-Business/db-metadata-extractor-mcp
— v0.1.6
Extract database metadata from PostgreSQL, Snowflake, SQL Server, BigQuery, and Oracle.
-
Outscraper MCP Server
— v0.2.3
Outscraper MCP business discovery, Maps intelligence, enrichment, reviews, and contact data.
-
io.github.Pakvothe/i1n
— v1.4.7
Agent-complete localization: extract strings, push keys, AI-translate, pull type-safe TS types.
-
io.github.patwalls/superhighway-mcp
— v1.2.0
Web search for AI agents — 5 tools: search, news, images, scrape, research. x402/USDC or API key.
-
BAILII UK Case Law
— v1.0.1
Search UK case law on BAILII — court judgments with section extraction. Runs locally.
-
PDF4me
— v1.0.0
PDF & document automation via the PDF4me API — convert, OCR, extract, edit, secure.
-
io.github.Perufitlife/multi-scraper-mcp
— v1.0.1
14 web scrapers as MCP tools: Reddit, Amazon, Google Maps, Yelp, YouTube, Indeed & more.
-
io.github.peter-schout/bol-ai
— v1.0.0
Extract structured data from Bills of Lading: parties, ports, containers, incoterms. EU-hosted.
-
io.github.pgalyen1987/gate402-mcp
— v0.5.0
Pay-per-call agent APIs over x402: web scraping, token compression, and semantic cache.
-
Decodo
— v0.1.0
Decodo MCP — wraps the Decodo Web Scraping API (decodo.com, formerly
-
Diffbot
— v0.1.0
Diffbot MCP — Knowledge Graph company enrichment + web content extraction (diffbot.com)
-
Oxylabs
— v0.1.0
Oxylabs MCP — Oxylabs Web Scraper API (oxylabs.io)
-
Scrapingdog
— v0.1.0
Scrapingdog MCP — wraps Scrapingdog (scrapingdog.com), a proxy-based web
-
Wiki Private Law
— v0.1.0
Extractive legal answers and semantic search over the public wiki.private.law corpus
-
io.github.PrometheusAgency/bluesky-scraper
— v0.1.4
Scrape Bluesky posts, profiles, followers, threads and keyword search. Clean JSON, pay per result.
-
io.github.PrometheusAgency/shopify-app-store-scraper
— v0.1.1
Scrape Shopify App Store apps, full review histories and catalog. E-commerce market research.
-
io.github.ProntoHQ/mcp-pronto
— v0.1.0
B2B sales intelligence: find companies, extract leads, enrich contacts with emails/phones.
-
Fitter
— v1.7.0
Turn any website or API into structured JSON with LLM-authored declarative scraping configs.
-
GroundAPI
— v1.0.0
Real-time data API for AI Agents: stocks, weather, forex, logistics, search, scrape, news, IP.
-
io.github.rachadele/biolit
— v0.1.25
LLM-assisted biomedical literature screening and extraction for PubMed, GEO, and preprints.
-
io.github.rchanllc/joltsms-sms
— v1.0.2
Provision real-SIM US phone numbers, receive SMS, and extract OTP codes for AI agents.
-
io.github.RECERQA/rq-scan
— v1.2.1
MCP Server for RQ-SCAN - AI-powered document OCR and data extraction platform
-
io.github.reefapi/reefapi-mcp
— v1.0.0
One MCP for 160+ live web-data APIs — clean JSON from sites that block scrapers.
-
cadloop Slicer
— v0.1.0
Find printer profiles, check bed fit, slice to .gcode.3mf and extract the G-code.
-
Robot Resources Scraper
— v0.1.2
Token compression for web content — 70-80% reduction from HTML to clean markdown
-
Video to Audio Converter MCP
— v0.1.0
Video to Audio Converter is a free, browser-based tool that extracts audio tracks from common video
-
io.github.rog0x/web
— v1.0.1
Web scraping, search, monitoring, and HTML-to-markdown for AI agents
-
Memory OS AI
— v3.0.1
Adaptive memory for AI agents — FAISS search, chat extraction, cross-project linking
-
io.github.RubenTay/agent-vending-factory
— v0.1.1
Pay-per-call MCP tools via x402 USDC: ZAR prices, data extraction, Python sandbox, SA flights.
-
Content Intelligence API
— v2.4.0
9 MCP tools: extract, analyze, research, compare, monitor, brief. Pay-per-call x402 or subscribe.
-
CoreWise
— v1.0.0
Extract structured insights from videos, podcasts, articles, and PDFs with multi-model AI
-
io.github.ryudi84/web-scraper
— v1.0.0
MCP server for web scraping: URLs, links, metadata, CSS extraction.
-
io.github.Sahib-Sawhney-WH/looking-glass-mcp
— v3.1.1
AI-native browser for agents. 71 tools with self-healing, semantic extraction, vault CLI.
-
reMarkable MCP Server
— v1.0.0
Access your reMarkable tablet - read documents, browse files, extract text and OCR
-
Arachne MCP
— v1.1.0
Web scraping, browser RAG, vision, transcription, and 20 AI tools. Universal data intelligence.
-
io.github.Samuelfmedeiros/arachne-mcp
— v1.1.2
Arachne MCP — scraping, visão, OCR, RAG + VRT visual (15 tools).
Page 5 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.