ZBS Index What actually exists in applied AI, with the source next to it

mcp server

Firecrawl MCP Server

MCP server for Firecrawl — search, scrape, and interact with the web.

Description as published by the maintainer. Source

  • version 3.23.4
  • active
  • data extraction
  • retrieval

active — Most recent push to the repository was 2026-08-08. Dashed tags are derived by ZBS Index from the published description, not stated by the maintainer.

What this server can do

25 functions, named and described by the server itself. Parameter names are shown because they say more about what a function does than its name usually does.

firecrawl_agent(urls, prompt, schema)
Start an asynchronous web research job from a prompt, optional seed URLs, and an optional JSON schema. Use this for a requested synthesis across multiple sources when the task can wait for asynchronous completion. The agent can search, navigate, read pages, and assemble a structured result. This call returns only a job ID, not the research result. Read the job with `firecrawl_agent_status` until it reaches `completed` or `failed`; research commonly takes several minutes. If the job cannot finish within the task's available time, `firecrawl_search` and `firecrawl_scrape` can gather evidence synchronously. Required: prompt.
firecrawl_agent_status(id)
Retrieve progress or final results for a `firecrawl_agent` job ID. A `processing` response is non-terminal and does not contain the final research result. Check again after 15–30 seconds until the status is `completed` or `failed`; complex jobs can take several minutes. If the job cannot finish within the task's available time, use `firecrawl_search` and `firecrawl_scrape` to complete the requested output. Returns job status, progress information, and result data when completed. Required: id.
firecrawl_check_crawl_status(id)
Retrieve the current status, progress, and available results for an existing crawl ID. This only reads Firecrawl job state and does not start or modify the crawl. Required: id.
firecrawl_crawl(url, delay, limit, prompt, sitemap, webhook, excludePaths, includePaths, scrapeOptions, maxConcurrency, webhookHeaders, allowSubdomains, crawlEntireDomain, maxDiscoveryDepth, allowExternalLinks, ignoreQueryParameters, deduplicateSimilarURLs)
Start a multi-page crawl at a website URL, poll it to a terminal state, and return the final status and collected data. Scope can be bounded with include/exclude paths, depth, page limit, subdomain/external-link controls, sitemap handling, delay, and scrape options. Crawl results can be large; use conservative limits when full-site coverage is unnecessary. Webhooks and interactive scrape actions are unavailable in safe mode. Returns the crawl ID, status, and page data. Required: url.
firecrawl_developer_search(k, query, skills)
For a developer question — code behaviour, a library or framework, an API contract, an error message, or a known bug — search an index built for coding agents. The index covers GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Set skills to "only" to limit the search to agent-skill files. Returns ranked results with an ID, source type, URL, title, and the matched passages in markdown. Required: query.
firecrawl_extract(urls, prompt, schema, enableWebSearch, includeSubdomains, allowExternalLinks)
Extract structured information from one or more URLs with an optional natural-language prompt and JSON schema. It can include subdomains, follow external links, or use web search when those options are enabled. Use this for a defined structured result rather than full page content. Returns data shaped by the supplied schema or prompt. Required: urls.
firecrawl_interact(url, code, prompt, timeout, language, scrapeId, scrapeOptions)
Open or reuse a live browser session to navigate a page, click controls, fill fields, or run browser code. Provide either `url` or `scrapeId`, and either a natural-language `prompt` or executable `code`; code can run as Bash, Python, or Node with a bounded timeout. This acts on the live site, so actions such as form submission can create persistent external side effects. Returns execution output, stdout/stderr, exit status, and session viewing URLs.
firecrawl_interact_stop(scrapeId)
Stop the live interact session associated with a `scrapeId` and release its resources. Returns a success confirmation. Required: scrapeId.
firecrawl_map(url, limit, search, sitemap, includeSubdomains, ignoreQueryParameters)
Enumerate URLs indexed under one website through Firecrawl without fetching each page's content. Use this when the request asks for a site's URL inventory, when several relevant pages must be located, or when the desired page URL is unknown. An optional `search` term narrows the URL list, while sitemap, subdomain, query-parameter, and result-limit options control coverage. Returns matching URLs rather than page bodies. Retrieve one page with `firecrawl_scrape`; collect content across multiple pages with `firecrawl_crawl`. Required: url.
firecrawl_monitor_check(id, skip, limit, checkId, pageStatus)
Retrieve one monitor check and its page-level results, optionally filtered by page status. Pages report `same`, `new`, `changed`, `removed`, or `error`; configured goal judging can add a meaningful-change decision. Markdown tracking returns a unified text diff, JSON tracking returns field paths with previous/current values and a current snapshot, and mixed tracking returns both. Returns one page of results plus a `next` URL when more pages exist. Required: id, checkId.
firecrawl_monitor_checks(id, limit, offset, status)
List historical checks for a monitor, optionally filtered by status and bounded by a result limit. Returns one page of check summaries and pagination metadata. Required: id.
firecrawl_monitor_create(body, goal, name, page, email, pages, queries, timezone, maxResults, webhookUrl, includeDiffs, scheduleText, searchWindow, excludeDomains, includeDomains)
Create a recurring scrape, crawl, or search monitor that compares each check with its retained predecessor. The simple form accepts `page`/`pages` or `queries` plus a plain-language `goal`; the advanced `body` form controls targets, schedule, change-tracking formats, judging, retention, webhook, and notifications. In the simple form, a `goal` is required. If `queries` contains one or more non-empty values and is supplied with `page`/`pages`, `queries` create the search target and page targets are ignored. A monitor schedules future network checks and can send configured email or webhook notifications. Returns the created monitor.
firecrawl_monitor_delete(id)
Permanently delete a monitor by ID and stop its future schedule. This operation cannot be undone and returns deletion status. Required: id.
firecrawl_monitor_get(id)
Retrieve one monitor by ID, including its configuration and current state. This does not run or modify the monitor. Required: id.
firecrawl_monitor_list(limit, offset)
List monitors for the authenticated account with optional pagination controls. Returns one page of monitor records and pagination metadata.
firecrawl_monitor_run(id)
Queue an immediate check for a monitor outside its normal schedule. This starts network work for the monitor's configured targets and returns the queued check. Required: id.
firecrawl_monitor_update(id, body)
Patch an existing monitor by ID. The body can change its name, active/paused status, schedule, targets, goal, judging, webhook, notifications, or retention; these changes affect future scheduled checks. Returns the updated monitor. Required: id, body.
firecrawl_parse(proxy, maxAge, formats, parsers, filePath, redactPII, pdfOptions, contentType, excludeTags, includeTags, jsonOptions, queryOptions, storeInCache, onlyMainContent, zeroDataRetention, removeBase64Images, skipTlsVerification)
Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with `pdfOptions.maxPages`. Local MCP reads `filePath` from the server filesystem. Hosted MCP uses two calls: first provide `filePath` to receive upload instructions, upload locally, then call again with the returned `uploadRef`; do not send both fields together. Remote web URLs belong in `firecrawl_scrape`. Set `redactPII` to request redaction of personally identifiable information in the returned content. `zeroDataRetention` requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Required: filePath.
firecrawl_research_inspect_paper(paperId)
Retrieve canonical metadata for one paper ID, such as an arXiv, PMC, PMID, or DOI identifier. Returns the title, abstract, authors, categories, source IDs, and dates as markdown. Required: paperId.
firecrawl_research_read_paper(k, paperId, question)
Retrieve in-body passages from one paper that are relevant to a specific question. Full text is available only for indexed papers; `k` controls the number of passages. Returns matching passages or a notice when full text is unavailable. Required: paperId, question.
firecrawl_research_related_papers(k, mode, intent, rerank, seed_ids)
Find citation-graph candidates from one to ten `seed_ids`; the first ID is the primary seed and later IDs are anchors. `mode` defaults to `similar` (co-citation/bibliographic coupling); `citers` returns papers citing a seed and `references` papers cited by a seed. `intent` ranks candidates. Returns ranked candidates and the evaluated pool size. Required: seed_ids, intent.
firecrawl_research_search_github(k, query)
Search indexed public GitHub issue, pull-request, and README content. Returns ranked matches with repository, URL, snippet, and full matched markdown when available. Required: query.
firecrawl_research_search_papers(k, to, from, query, authors, categories)
For topics represented in the indexed corpus, search paper metadata and abstracts with a natural-language query. Optional author, category, and date filters constrain results. Returns ranked papers with canonical IDs, titles, authors, and abstracts. Required: query.
firecrawl_scrape(url, proxy, maxAge, mobile, actions, formats, parsers, profile, waitFor, location, lockdown, redactPII, pdfOptions, excludeTags, includeTags, jsonOptions, queryOptions, storeInCache, onlyMainContent, screenshotOptions, zeroDataRetention, removeBase64Images, skipTlsVerification)
Retrieve and extract content from one supplied URL through Firecrawl. Use this when the request identifies a page and needs its content or defined fields. It can return markdown, HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema; JSON is useful when the requested result has defined fields, while markdown preserves readable page content. This tool operates on a known page. For a set of pages use `firecrawl_crawl`, and to discover page URLs use `firecrawl_map` or `firecrawl_search`. Options include JavaScript render delay, cache age, main-content filtering, PII redaction, and lockdown cache-only retrieval. Browser actions may change the live page when interactive actions are enabled. Returns the selected content formats and page metadata. Required: url.
firecrawl_search(tbs, limit, query, filter, sources, location, categories, enterprise, highlights, scrapeOptions, excludeDomains, includeDomains)
Search web, news, or image sources and return ranked results. Operators include quoted phrases, `-term`, `site:host`, `inurl:term`, `intitle:term`, and `related:host`; the set is non-exhaustive. `includeDomains` and `excludeDomains` are mutually exclusive hostname filters; categories limit results to GitHub, research, PDF, or developer sources. For a programming question, add `categories: ["developer"]`. It searches an index of GitHub issues, merged pull requests, repository READMEs, and curated documentation sites, and returns the hits in `data.developer` beside the web results. `scrapeOptions` can attach extracted page content. Returns source-type result groups and usage metadata. Authenticated responses can include an `id` for optional search feedback. Required: query.

Last successful function declaration observed on . Source: pkg:npm/firecrawl-mcp. We list what the server declared; we do not call any of these functions.

Endpoint status observed on . Source: pkg:npm/firecrawl-mcp.

Signals

These are separate measurements of different things. They are deliberately not combined into one score, because a popularity number that mixes website traffic with saves and stars cannot be checked or acted on.

Signal Value What it measures Window Observed Source
GitHub stars 7,174 Number of GitHub accounts that bookmarked this repository since it was created. It is a bookmark count, not installs, not active users and not quality. cumulative, all time GitHub
Last commit 2026-08-08 Date of the most recent push to any branch. This is the strongest cheap indicator of whether the project is still maintained. point in time GitHub
Open issues 148 Open issues plus open pull requests, as GitHub counts them together. A high number can mean an active project or an abandoned one. as of fetch GitHub
Latest published version 3.23.4 Latest version string the maintainer published to the registry. as of fetch Model Context Protocol
Registry record last updated 2026-08-06 When the registry record was last updated by its maintainer. point in time Model Context Protocol
License MIT Licence GitHub detected in the repository. Detection can be wrong; the LICENSE file is authoritative. as of fetch GitHub
First listed in the MCP Registry 2026-08-06 Date this server was first published to the official MCP Registry. Not a usage or quality measure. point in time Model Context Protocol
repository status active The repository exists on GitHub and is not archived. This says nothing about how recently it was worked on. as of fetch GitHub
mcp tools declared 25 tools Number of functions the server declared when started and asked to list them. It says what the server offers an agent, not how well any of it works. as of probe npm
mcp endpoint status ok The server listed 25 functions when asked. as of probe npm
package install scripts none This package declares no install-time scripts, so installing it does not execute any of its code. as of probe npm

Where to get it

Related, by what their authors tagged them

  • ai.smithery/pinkpixel-dev-web-scout-mcp — last commit 2026-07-21, shares content-extraction, web-crawler, web-scraping
    Search the web and extract clean, readable text from webpages. Process multiple URLs at once to sp…
  • searchts — last commit 2026-08-04, shares content-extraction, web-scraping
    Keyless web access for AI agents: read bot-walled pages as Markdown, search, grab assets
  • DLBrowser — last commit 2026-07-22, shares web-crawler, web-scraping
    Self-healing web access for AI agents — recovers through bot blocks, captchas and JS pages.
  • Librecrawl — Technical SEO Audit MCP Server — last commit 2026-07-28, shares web-crawler
    Self-hosted technical SEO audit MCP. 50+ checks, WAF detection, ephemeral. Built on LibreCrawl.
  • Web Search — last commit 2026-07-29, shares content-extraction
    Keyless multi-engine web search, fetch, and clean-Markdown read for AI agents. No API keys.
  • io.github.brightdata/brightdata-mcp — last commit 2026-07-27, shares data-collection, web-scraping
    Bright Data's Web MCP server enabling AI agents to search, extract & navigate the web
  • ai.smithery/oxylabs-oxylabs-mcp — last commit 2026-06-08, shares data-collection
    Fetch and process content from specified URLs using the Oxylabs Web Scraper API.
  • Jitsu — last commit 2026-08-06, shares data-collection
    Manage Jitsu data pipelines: destinations, streams, connections, functions, live events.
  • io.github.cyanheads/survey-mcp-server — last commit 2025-10-20, shares data-collection
    MCP server for conducting dynamic, conversational surveys with structured data collection.
  • io.github.auxiliar-ai/auxiliar-mcp — archived, last commit 2026-07-11, shares search-api, web-scraping
    Eval-backed discovery for the auxiliar.ai gateway — the best web-access provider per job, measured.

These share tags the maintainers applied themselves, such as content-extraction, web-crawler, web-scraping, data-collection. Common tags like "mcp" or "ai" are ignored for this: agreeing with six hundred other projects is not a similarity.

This is not a recommendation and not a test result. It is a map of what the authors said their work is about.

How the author describes it

Topics the maintainer set on GitHub: batch-processing, claude, content-extraction, data-collection, firecrawl, firecrawl-ai, javascript-rendering, llm-tools, mcp, mcp-server, model-context-protocol, search-api, web-crawler, web-scraping.

This record as data

Every field on this page, with its source and observation date, is in the catalog JSON. Fetch the whole kind at once instead of parsing this HTML.

GET /api/v1/entries/mcp_server.json

Sources

  1. firecrawl/firecrawl-mcp-server on GitHub — GitHub, observed , trust tier 3.
  2. Official MCP Registry — Model Context Protocol, observed , trust tier 1.
  3. firecrawl-mcp started in an isolated container — npm, observed , trust tier 1.