Data extraction
Everything here was classified as data extraction by keyword match against the maintainer's own description, so treat the grouping as a starting point rather than a verdict.
The list is ordered by the most recent commit, not by stars. A popular project that stopped in 2024 is not a better answer than a smaller one shipped last week.
Listed, not yet verified by us (60)
Published to the registry, but we have not yet checked its repository. Treat the entry as the maintainer’s claim only.
-
io.github.Sandip124/wisegit
— v0.1.1
Extracts decision intent from git history and protects intentional code from AI modification.
-
io.github.sathvic-kollu/techtenstein-pdf
— v1.0.2
PDF text and table extraction plus metadata. Supports OCR for scanned documents.
-
ScrapeBadger
— v0.1.1
Twitter/X scraping API for AI agents. Get profiles, tweets, trends, and more.
-
io.github.ScrapeGraphAI/scrapegraph-mcp
— v1.0.1
AI-powered web scraping and data extraction capabilities through ScrapeGraph API
-
io.github.scrapercity/scrapercity-cli
— v1.0.1
B2B lead generation MCP server - 20+ scrapers, email finder, skip trace, and more.
-
OpenGraph.io MCP Server
— v1.3.5
MCP server for OpenGraph.io API - fetch OG data, screenshots, scrape, and generate images
-
io.github.selvage-lab/selvage
— v0.4.1
An LLM-based code review MCP server with AST-powered smart context extraction
-
x402 Web Extract
— v1.0.1
Pay-per-use web extract, token prices, and wallet balances via x402 USDC micropayments.
-
ClawPage
— v0.2.1
Extract and structure any web page into clean JSON. Free tier: 10/day.
-
Skyvern
— v1.0.23
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
-
Cyberbro MCP Server
— v0.0.5
MCP server for Cyberbro IOC extraction, enrichment and reputation analysis.
-
HTML to Markdown MCP Server
— v0.3.0
Converts HTML to Markdown with auto-summary and section extraction. Supports Playwright.
-
Suparse Document Processing
— v1.3.0
Extract structured data from PDFs and documents with Suparse AI OCR.
-
Evocrawl MCP Server
— v4.0.1
MCP server for Evocrawl — search, scrape, and interact with the web.
-
Payx Tavily Search
— v0.1.0
FastMCP server for Tavily search API with web search and content extraction
-
ReadGZH
— v1.0.0
Read WeChat public account articles via MCP. 99.89% anti-scraping success, 50-87% token compression.
-
io.github.syedtaj7/kitbag-mcp
— v1.1.2
Zero-config MCP server bundling 50+ utility tools: converters, OCR, scraping, and formatters.
-
Talonic
— v0.1.73
Extract structured, schema-validated data from PDFs, scans, images, spreadsheets, and forms.
-
io.github.tarunlnmiit/inbox-to-action
— v0.2.11
Agentic Gmail triage: classify, summarize, extract tasks, draft replies. Never sends.
-
io.github.tathagat22/plumb-mcp
— v0.13.2
Two-way Figma MCP: extract + verify design-to-code, or generate on-brand Figma pages from a prompt.
-
io.github.teamsincetoday/newsletter-commerce-mcp
— v0.1.3
Extract shoppable products from newsletter content with affiliate signals.
-
io.github.teamsincetoday/podcast-commerce-mcp
— v0.1.3
Extract product mentions from podcast transcripts with affiliate signals. Free tier 200/day.
-
io.github.teamsincetoday/recipe-commerce-mcp
— v0.1.3
Extract shoppable ingredients from recipe content with affiliate signals.
-
io.github.terradeed/terradeed-mcp-server
— v0.1.4
TerraDeed MCP server — x402 web scraping & structured extraction. Pay-per-call with USDC on Base.
-
Kognitrix AI
— v1.0.0
8 AI services via MCP: content, code, image, docs, data extraction, translation, SEO, email.
-
PDFTools
— v1.0.0
Extract text and tables from PDFs, merge documents, and get metadata.
-
io.github.therealMrFunGuy/test-mailbox
— v1.0.0
Disposable email inboxes for testing — create addresses, wait for messages, extract links.
-
io.github.therealtimex/browser-use
— v0.7.10
AI browser automation - navigate, click, type, extract content, and run autonomous web tasks
-
io.github.thunderbit-com/thunderbit-mcp-server
— v1.0.4
AI-powered web scraping MCP. Distill pages to Markdown or extract structured data via JSON Schema.
-
io.github.thunderbit-open/thunderbit-mcp-server
— v1.0.3
AI-powered web scraping MCP. Distill pages to Markdown or extract structured data via JSON Schema.
-
io.github.thuupx/lunge
— v1.3.1
Agent-native API client: execute/test REST/GraphQL/WS/SSE with assertions, extraction, collections
-
io.github.tiliondev/fortress
— v0.1.3
Stealth browser for AI agents: fetch pages behind Cloudflare/DataDome/CAPTCHA, extract clean data.
-
Aetheris-MCP — x402 Web Scraping MCP Server
— v1.0.0
x402-gated web scraping MCP server. Pay 0.01 USDC per page on Base. Returns clean Markdown via JSDOM
-
io.github.Tlalvarez/auxiliar-mcp
— v0.22.1
Web access for AI agents — search/scrape/extract/crawl via the auxiliar.ai gateway.
-
io.github.TN0123/one-shot-ui
— v0.7.0
Deterministic UI extraction and screenshot diffing for AI coding agents.
-
Mirror Memory MCP Server
— v1.0.0
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
-
io.github.tstockham96/engram
— v0.5.1
Intelligent agent memory with automatic extraction, consolidation, and bi-temporal recall.
-
Video URL Analyzer
— v1.5.1
Analyze YouTube/TikTok/Instagram videos: transcripts, AI insights, tutorial extraction
-
io.github.UltraStarz/x402-extract
— v0.1.2
Pay-per-call MCP server. $0.01 USDC extracts schema.org/Product JSON from any URL via x402.
-
CRW Web Scraper
— v0.29.0
Open-source web scraper for AI agents with scrape, crawl, and map tools
-
fastCRW
— v1.0.0
Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.
-
io.github.User0856/snaprender
— v1.5.4
Screenshot, extract content, and render web pages. PNG/JPEG/WebP/PDF, batch, signed URLs.
-
Web Content Extract Mcp
— v1.0.0
Web Content Extract Mcp connects AI agents to real public APIs via MCP. Tools include
-
io.github.Vektoris-AI/proofetch
— v0.0.2
Verified extraction: source-backed JSON from PDFs/URLs; honest null + signed receipt.
-
xportalx
— v1.0.0
AI agent gateway with web fetching, data extraction, crypto pricing, and x402 payments
-
io.github.virgilvox/vue-harvest
— v0.0.3
Extract reusable component libraries and design tokens from Vue applications
-
io.github.VOYAGER-Inc/excel-vision-mcp
— v1.2.1
Excel MCP with image extraction and write support — AI agents see and edit xlsx files
-
io.github.vpatser1/data-aggregator-mcp
— v1.1.4
Unified data server: stocks, crypto, news, weather, web scraping, FX rates in one MCP
-
reMarkable MCP Server
— v0.4.7
Access your reMarkable tablet - read documents, browse files, extract text and OCR handwritten notes
-
Web Resurrect
— v1.4.0
Resurrect expired domains: scrape Wayback, AI-rewrite, generate images, publish to WordPress.
-
io.github.webberdesign/webbersites-x402-data-api
— v1.0.0
45 pay-per-call AI agent tools: scraping, SEO, crypto data, lint, agent memory. x402 USDC on Base.
-
io.github.webberdesign/webbersites-x402-mcp
— v0.3.0
45 pay-per-call AI agent tools + wallet-owned agent memory. Scraping, SEO, lint, crypto. x402 USDC
-
WebScraping.AI
— v1.0.7
Web scraping tools with Chromium JS rendering, rotating proxies, and AI question answering.
-
FramesCLI
— v0.1.0
Make videos AI-readable. Extract timestamped frames and transcripts from any video for AI agents.
-
Crawlberg
— v1.1.4
Scrape, crawl, and map websites to Markdown or JSON via local CLI.
-
io.github.XJTLUmedia/ai-hr-management-toolkit
— v3.0.5
AI HR toolkit: 24 MCP tools for resume parsing, skill extraction & ATS management.
-
xpay All Tools
— v1.0.0
980+ AI tools as MCP servers. Finance, lead gen, scraping, dev tools, media, research, and more.
-
xpay Social Media
— v1.0.0
96 social media scraping tools. Twitter/X, LinkedIn, Instagram, TikTok, Reddit, YouTube.
-
xpay Web Scraping
— v1.0.0
35+ web scraping tools. Firecrawl, Bright Data, Jina, Olostep, ScrapeGraph, Notte, Riveter.
-
HackTricks MCP Server
— v1.3.4
Search and query HackTricks pentesting documentation with quick lookup and section extraction
Page 6 of 7
How this page is ordered
Entries are grouped by whether anyone is still working on them, using the date of the most recent push to the repository. They are not ordered by stars, because a star is a bookmark somebody left once and never took back.
Where we have not checked an entry yet, it says so rather than being mixed in with the verified ones.