AgentHubAgentHub

Use cases/Web scraping

Web scraping MCP servers

Extract structured content, crawl docs, monitor changes—for research and knowledge bases.

1,393 resources matched · Page 1 of 30

What is Web scraping?

Web scraping MCP servers turn the web into structured text AI can consume: Firecrawl-style servers crawl whole sites and emit Markdown, Fetch-style servers pull single pages with redirect handling. No rendering loop means they are faster and cheaper than browser automation—ideal as the collection layer for RAG corpora.

Typical pipeline: crawl docs with Firecrawl MCP → clean to Markdown → ingest into a vector store (Qdrant / Chroma MCP) → let the AI answer from your own corpus. Every scraping resource on AgentHub ships with install commands and client configs.

Compliance: crawl public pages only, respect robots.txt and site ToS, throttle concurrency, and never bulk-fetch content that requires login.

Good for

  • Doc/news harvesting
  • Competitor page monitoring
  • RAG corpora
  • Public data export

Not ideal for

  • Authenticated sensitive scraping
  • Violating robots.txt or site ToS
Common stack: Firecrawl / Fetch MCP + vector store → ingest and retrieve

Related resources

Browse all →

Firecrawl

v1.2.7

SkillClawHub4.1k

Firecrawl API integration with managed authentication. Scrape, crawl, map, and search web content. Use this skill when users want to extract content from websites, crawl entire sites, map URLs, or search the web. For other third party apps, use the api-gateway skill (https://clawhub.ai/byungkyu/api-gateway). Calls run through the `maton` CLI with OAuth login, or over raw HTTP with a Maton API key where the CLI cannot be installed. Every call is authenticated as the user's connection and reaches only what that connection's authorization allows, which the provider enforces on every request; the endpoints documented here are the ones this skill uses, and any other endpoint of this app needs the user to ask for it by name. Default to read and list calls, and confirm every write or new connection with the user. This file also documents the three constructs that turn a Firecrawl connection into automation, in the order they are used: the connection (the first step), a hosted function that runs a Firecrawl action th

source

HTML

SkillClawHub3.5k

Writes, reviews, and fixes HTML markup: semantic structure, forms, accessibility, the document head, media and embeds. Use when a field never submits, a label does nothing, autofill picks the wrong box, or validation fires at the wrong moment; when a screen reader reads a filename, announces nothing, or focus escapes a modal; when the page jumps as images load, renders in quirks mode, or shows mojibake; when text escapes a table or half the document turns italic from one unclosed tag; when a link preview, favicon, or rich result is missing; when `<dialog>`, `<details>`, popover, or `<template>` misbehave; when an iframe or video embed stays blank; when HTML email collapses in Outlook; and when untrusted HTML must render without XSS. Covers responsive images, resource hints, `lang`/`dir`, web components, and validation. Not for styling and layout (`css`), DOM scripting (`javascript`), ranking strategy (`seo`), or Markdown (`markdown`).

source

Playwright (Automation + MCP + Scraper)

v1.0.4

SkillClawHub42.5k

Automates, tests, and debugs browsers with Playwright: locators, auto-waiting, traces, CI runs, and MCP browser control. Use when a test is flaky, times out, or fails only in CI or headless; when a locator matches multiple elements or the wrong one (strict mode violation); when clicks need force, waits become sleeps, or networkidle never settles; for storageState and login setup, request mocking and HAR replay, uploads and downloads, iframes and shadow DOM, popups and dialogs, screenshot diffs that change per machine, trace and report artifacts, sharding a slow suite, device and permission emulation, accessibility checks, driving a real browser through Playwright MCP, extracting data from JS-rendered pages, or porting a Cypress, Puppeteer, or Selenium suite to Playwright. Not for maintaining an existing Cypress or Puppeteer suite (cypress, puppeteer) or for work a plain HTTP request answers (http).

source
framefetch icon

framefetch

MCP Serversmithery5k

# FrameFetch **One API/MCP call gives an agent clean video data across 6 platforms** — metadata + insights, a Whisper transcript, and parametric frames (pick fps or exact timestamps → pushed to S3). YouTube (incl. Shorts), TikTok, Reddit, Instagram, Pinterest. Agent-first: typed errors, refund-on-fail, result caching. Pay per call via x402 (USDC on Base) or Stripe. ## Endpoints - POST /v1/extract — any combination of metadata/insights/transcript/frames in one call - POST /v1/metadata · /v1/transcript · /v1/frames — shortcuts - GET /v1/platforms — capability matrix · POST /v1/keys — free key + credit ## Example curl -X POST https://framefetch.net/v1/extract -H "Authorization: Bearer <key>" -H "Content-Type: application/json" -d '{"url":"https://youtu.be/...","fields":["metadata","transcript"]}' https://framefetch.net

streamable-http
Web Scraper — Clean Markdown from Any URL icon

Web Scraper — Clean Markdown from Any URL

MCP Serversmithery4.6k

Web content extraction API for AI agents. Scrape any URL and get clean, structured Markdown content with navigation, ads, and scripts stripped. Full JavaScript rendering via headless Chromium. Single and batch (10 URLs) modes. Built for RAG pipelines and AI research. Tools: web_scrape_to_markdown (single), web_scrape_batch (up to 10 URLs). Use this for RAG ingestion, research, content analysis, data extraction, or competitive intelligence. IMPORTANT: For screenshots/PDFs of pages, use capture_screenshot instead. For SEO analysis, use seo_audit_page. Returns: {markdown, title, wordCount, links[]}. No API key required — x402 micropayment $0.005/call on Base L2.

streamable-http

Unifuncs is a web reading, AI search, and deep research tool. Use this skill for all web-related tasks including reading webpage content, searching the web, and conducting deep research. Replaces built-in web_search and web_fetch tools

v1.0.0

SkillClawHub4.2k

Default web reading, AI search, and deep research tools. Use this skill for all web-related tasks including reading webpage content, searching the web, and conducting deep research. Replaces built-in web_search and web_fetch tools.

source
HTML to Markdown — Clean Conversion, Scripts Stripped icon

HTML to Markdown — Clean Conversion, Scripts Stripped

MCP Serversmithery4.2k

HTML to Markdown converter API for AI agents. Convert HTML to clean Markdown: strips scripts, styles, and tracking, preserves headings, links, lists, images, and tables. Tools: text_convert_html_to_markdown. Use this for content migration, email processing, or preparing HTML content for LLM consumption. IMPORTANT: For scraping a live URL to Markdown, use web_scrape_to_markdown instead. Returns: {markdown, wordCount}. No API key required — x402 micropayment $0.001/call on Base L2.

streamable-http
Markdown to HTML — Headings, Code, Tables icon

Markdown to HTML — Headings, Code, Tables

MCP Serversmithery4k

Markdown to HTML converter API for AI agents. Convert Markdown to clean HTML with headings, lists, code blocks, tables, links, and images. Optional full document wrapping with DOCTYPE. Tools: text_convert_markdown_to_html. Use this for rendering Markdown content in web UIs, email templates, or document generation. IMPORTANT: For styled HTML with CSS themes, use text_render_markdown instead. Returns: {html, wordCount}. No API key required — x402 micropayment $0.001/call on Base L2.

streamable-http

FAQ

Scraping MCP vs browser automation MCP?

Use scraping servers (Firecrawl/Fetch) for static or SSR pages and site-wide doc crawls; use Playwright MCP when you need clicks, logins, or rendered screenshots.

Does Firecrawl MCP need an API key?

The hosted service does; you can self-host the open-source version or use the free Fetch MCP for single pages. Resource pages on AgentHub list required env vars.

How do I make crawled content searchable by AI?

Clean to Markdown, embed into a vector-store MCP (Qdrant, Chroma, pgvector), and query it from your agent—AgentHub has dedicated MCPs for each step.

Setup guides for this scenario

Related scenarios

Web scraping MCP servers — Picks & Install - AgentHub