AgentHubAgentHub

使用场景/网页抓取与采集

网页抓取 MCP 推荐

从网页提取结构化内容、爬取文档、监控变更。适合调研、内容聚合、知识库构建。

共匹配 512 个资源 · 第 2 / 11 页

什么是网页抓取与采集?

网页抓取类 MCP Server 把互联网内容变成 AI 可直接消费的结构化文本:Firecrawl 类服务负责整站爬取与 Markdown 转换,Fetch 类服务负责单页拉取与重定向处理。相比浏览器自动化,它们不渲染交互、速度快、token 消耗低,是构建 RAG 语料与知识库的首选采集层。

典型工作流:用 Firecrawl MCP 批量爬取文档站 → 清洗为 Markdown → 送入向量库(如 Qdrant / Chroma MCP)→ 再用检索 MCP 让 AI 基于自有语料回答。AgentHub 上每个抓取类资源都附带安装命令与客户端配置,可直接复制到 Cursor 或 Claude Code。

合规边界:只抓取公开页面,遵守目标站点 robots.txt 与服务条款;控制并发与频率,避免给对方服务造成压力;需要登录才能访问的内容不要自动化批量拉取。

适合做什么

  • 新闻/文档采集
  • 竞品页面监控
  • 构建 RAG 语料
  • 批量导出公开数据

不太适合

  • 登录后敏感数据爬取
  • 违反 robots.txt 或站点 ToS 的场景
常见组合: Firecrawl / Fetch MCP + 向量库 → 自动入库与检索

相关资源

查看全部 →

discourse-writing-html-css(Discourse HTML/CSS 编写)

SkillSkillsMP

为 Discourse 核心、插件、主题与主题组件编写和修复 HTML/CSS/SCSS。在编写或修改模板(.gjs/.hbs)、样式表(.scss)、组件标记、类名、响应式布局、FormKit/select-kit 样式或处理 CSS 回归时使用。涵盖 Discourse 的 BEM 加独立修饰符命名、CSS 自定义属性色彩体系(主题化与暗色模式)、模板/HTML 约定、CSS 修复模式,以及样式表的存放位置。

source

PPTX 与 HTML 保真度审查 pptx-html-fidelity-audit

SkillSkillsMP

将 python-pptx 导出结果与其源 HTML 演示文稿对照审查,识别布局/内容偏差(页脚溢出、内容裁切、缺失斜体/强调、样式丢失、节奏失衡的间距),并以严格的页脚导轨 + 光标流布局纪律重新导出。当用户的 .pptx 由 HTML 幻灯片生成并要求比较/审查/验证/修复导出时使用——包括「compare ppt with html」「fidelity audit」「fix the pptx」「ppt is cut off」「footer overlap」「italic missing in pptx」「re-export the deck」「pptx-html-fidelity-audit」等说法,或任何需要验证/修复 python-pptx 与 HTML 往返转换的场景。当用户并排展示 deck.html 与 deck.pptx 并调试视觉差异时也应触发。

source

PPTX 与 HTML 保真度审查 pptx-html-fidelity-audit

SkillSkillsMP

将 python-pptx 导出结果与其源 HTML 演示文稿对照审查,识别布局/内容偏差(页脚溢出、内容裁切、缺失斜体/强调、样式丢失、节奏失衡的间距),并以严格的页脚导轨 + 光标流布局纪律重新导出。当用户的 .pptx 由 HTML 幻灯片生成并要求比较/审查/验证/修复导出时使用——包括「compare ppt with html」「fidelity audit」「fix the pptx」「ppt is cut off」「footer overlap」「italic missing in pptx」「re-export the deck」「pptx-html-fidelity-audit」等说法,或任何需要验证/修复 python-pptx 与 HTML 往返转换的场景。当用户并排展示 deck.html 与 deck.pptx 并调试视觉差异时也应触发。

source

PDF 转 HTML

SkillSkillsMP

将 PDF 转换为单个自包含、可读性好的 HTML 文件,保留图片、表格、图表和阅读顺序——还可选择翻译为另一种语言同时保留全部图形。使用结构化抽取(PyMuPDF)、基于字号的版式推断、压缩的 base64 内联图片(单一可携带文件),以及强制的无头 Chrome 视觉校验。当有人想把 PDF 作为网页或干净文档来阅读、把 PDF 转成 HTML,或在保留图/表/图表的前提下翻译 PDF 时使用——例如“PDF 转 HTML”“把这个 PDF 转成中文网页版”“让这份报告变得可读”“翻译这份 PDF 但别丢图表”“我只想在手机上读这份 PDF”。与 doc-to-markdown(纯 Markdown 文本)和 pdf-creator(Markdown→PDF)不同——本技能产出带样式、忠实还原图片的 HTML 阅读体验。

source

google-search

v0.1.0

SkillClawHub4.6k

A Google Search API alternative and SERP API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no API key application. Use when the user wants programmatic Google search results as clean JSON, including Google's own operators — site:, filetype:, intitle:, and quoted exact phrases — plus pagination, language (hl), and country/region scoping. Also covers rank tracking input, competitive research, and search-result monitoring without Google's own Custom Search API quota and billing.

source

App Store API

v0.1.0

SkillClawHub4.5k

An Apple App Store API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no Apple Developer Program membership. Use when the user wants to search iOS apps by keyword in any country's storefront, fetch an app's or app bundle's full details, reviews sorted by recent/helpful, apps similar to a given app, search or fetch app bundles, or fetch a developer's app catalog. Also covers iOS app store optimization (ASO) research, competitor app monitoring, localized storefront comparison, and app discovery without Apple Developer Program access.

source

Yelp API

v0.1.0

SkillClawHub4.5k

A Yelp API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no Yelp Fusion API app approval. Use when the user wants to search local businesses by query and location sorted by rating or review count, fetch a business's full details by ID or by its Yelp URL handle/slug, or fetch a business's reviews. Also covers local business discovery, restaurant/service research, review sentiment input, and competitor monitoring for local businesses without Yelp Fusion API's app-approval process and daily call caps.

source

Google News API

v0.1.0

SkillClawHub4.5k

A Google News API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no RSS scraping. Use when the user wants keyword search across Google News scoped to a language edition, section headlines (world, business, technology, entertainment, sport, science, health, or a specific topic ID), the latest headlines, a list of supported language-region codes, or to decode a Google News redirect URL into the real article URL. Also covers news monitoring, headline aggregation, media tracking, and press-mention alerts without scraping Google News' RSS feeds directly.

source

Google Play

v0.1.0

SkillClawHub4.5k

A Google Play Store API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no Google Play Console access. Use when the user wants to search Android apps by keyword with price (free/paid) and country storefront filters, fetch an app's full details, reviews sorted by newest/rating/helpfulness, permissions, or data safety disclosure, list apps similar to a given app, or fetch a developer's app catalog. Also covers Android app store optimization (ASO) research, competitor app monitoring, review sentiment input, and app discovery without Google Play Console developer access.

source

Google Maps API

v0.1.0

SkillClawHub4.4k

A Google Maps API alternative and Google Places API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no Google Cloud billing account. Use when the user wants to search places by text query or anchor the search to a latitude/longitude coordinate, fetch a place's full details by its feature ID (fid), pull a place's reviews sorted by relevance/newest/rating, or look up a single review by ID. Also covers local business discovery, points-of-interest data, store-locator input, and review monitoring without Google Cloud Platform's API key setup, billing, and per-request pricing.

source

Youtube API

v0.1.0

SkillClawHub4.4k

A YouTube API alternative on fetcher.sh — pay-per-call in USDC via x402, or prepaid credits with a Bearer key, no OAuth and no daily quota. Use when the user wants to search YouTube videos with filters for upload date, duration, or sort order, search channels or playlists, look up a channel by ID, @handle, or custom URL path, list a channel's videos, shorts, or live streams, fetch a video's or short's details and comments, list a playlist's videos, pull posts under a hashtag, or check trending videos by region. Also covers YouTube data pipelines, channel monitoring, video analytics input, and trend tracking without the official YouTube Data API's daily quota limits.

source

常见问题

网页抓取 MCP 和浏览器自动化 MCP 怎么选?

静态或服务端渲染的页面、整站文档爬取用抓取类(Firecrawl/Fetch);需要点击、滚动、登录交互或抓渲染后截图,用浏览器自动化(Playwright MCP)。

Firecrawl MCP 需要 API Key 吗?

托管版需要;也可以自部署开源版,或用免费的 Fetch MCP 做单页抓取。资源详情页会标注分发方式与所需环境变量。

抓回来的内容如何喂给 AI 检索?

常见做法是 Markdown 清洗后写入向量库 MCP(Qdrant、Chroma、pgvector 等),再配合检索 Skill 或 Context7 类知识库 MCP 完成问答。

配套安装与配置教程

相关场景