Crawlgraph
v0.1.1
CrawlGraph MCP — backlink intelligence on the public Common Crawl webgraph
使用场景/网页抓取与采集
从网页提取结构化内容、爬取文档、监控变更。适合调研、内容聚合、知识库构建。
共匹配 889 个资源 · 第 3 / 19 页
网页抓取类 MCP Server 把互联网内容变成 AI 可直接消费的结构化文本:Firecrawl 类服务负责整站爬取与 Markdown 转换,Fetch 类服务负责单页拉取与重定向处理。相比浏览器自动化,它们不渲染交互、速度快、token 消耗低,是构建 RAG 语料与知识库的首选采集层。
典型工作流:用 Firecrawl MCP 批量爬取文档站 → 清洗为 Markdown → 送入向量库(如 Qdrant / Chroma MCP)→ 再用检索 MCP 让 AI 基于自有语料回答。AgentHub 上每个抓取类资源都附带安装命令与客户端配置,可直接复制到 Cursor 或 Claude Code。
合规边界:只抓取公开页面,遵守目标站点 robots.txt 与服务条款;控制并发与频率,避免给对方服务造成压力;需要登录才能访问的内容不要自动化批量拉取。
v0.1.1
CrawlGraph MCP — backlink intelligence on the public Common Crawl webgraph
v1.0.0
Scrape any URL into clean LLM-ready Markdown, or crawl a whole site.
v1.0.1
Google Trends for agents: trend verdict, stats, rising queries, regions, trending now.
v1.0.3
Telegram Channel Scraper - No Login, $0.30 per 1K Posts
v1.0.2
Zomato Restaurant Scraper - City, Cuisine, Contact
v1.0.0
YouTube transcripts, Google Hotels prices, Google Ads Transparency, Google Trends and Threads posts.
v1.0.0
Bounded crawl-readiness check for agent access: robots.txt, sitemap.xml, llms.txt, homepage.
v1.1.0
Public Instagram and TikTok data: profiles, posts, reels, stories, comments, hashtags and places.
v0.1.0
Audit webpage access for AI crawlers. API key and Premium or Partner plan required.
v1.0.0
One-call audit: robots.txt AI-bot rules, llms.txt, JSON-LD, meta tags, sitemap. Score plus fixes.
v1.0.1
Real-time data APIs for 50+ commerce, grocery, social and maps platforms over one OAuth MCP server
v1.0.0
Convert HTML to clean markdown locally. No network and no key.
v1.0.0
Fetch title, description, and open graph tags from a web page. No key required.
v1.0.1
Use this MCP server to Common Crawl index discovery for historical web captures.
v0.8.1
Web scraping + research toolkit for AI agents: scrape, crawl, monitor URLs with anti-bot bypass.
v0.1.0
HTML to clean Markdown for LLM and RAG pipelines: headings, lists, tables, links, code fences.
v1.0.0
Can AI crawlers read this site? GPTBot/ClaudeBot verdicts, robots vs edge blocking, 47k-domain index
v0.1.3
Checks robots.txt line by line and says what each AI crawler token actually controls - training, AI
v1.0.1
YouTube, TikTok, Instagram, X, Reddit, LinkedIn and Threads: public data in one schema.
v0.1.2
16 BEM naming rules read the HTML file you have open and 22 markup and SCSS skeletons write the styl
v0.36.0
面向 AI 智能体的开源网页抓取器,提供抓取、爬取与站点结构梳理工具。
v1.4.0
40 Apify public-data tools for leads, news, SEO, jobs, SEC, procurement, and registries.
v0.1.2
Google Ads by domain or brand: creatives, formats, dates; or bulk-check which domains advertise.
v1.0.0
Autonomous HTTP 402 Web Scraping Mesh on Base for $0.02 USDC.
v1.0.0
HTML numeric-entity count, input discarded
v0.1.1
Classic Outlook renders with the Word engine, the new Outlook with WebView2. Lint one email file aga
v0.3.1
Local-first public-data MCP server that writes inspectable records and working files.
v1.0.1
Fetch, crawl, and browse protected pages with anti-bot handling - renders in a real browser and
v1.1.0
Publish HTML5 games in seconds from an AI agent: audience, monetization, analytics.
v1.5.0
搜索品牌,并从 Brandfetch API 获取设计资源、企业资料与其他品牌背景信息。
v1.0.0
WAF-bypass web scraper and passive URL mapper for AI agents with zero pay-on-fail.
v0.2.0
Google, Bing, DuckDuckGo and Google Maps search results, free without an API key.
v1.0.0
Produtos da Shopee Brasil por busca ou loja: preço, Pix, nota, cidade, loja e vendidos
v0.1.0
Depop listings for searches or URLs, with alerts on new listings, price drops and sales.
v1.7.0
对访问日志中的爬虫流量做归类与诊断,免费无需密钥(原文功能描述为未填写的占位符)。
v1.2.2
面向 YouTube、TikTok 与 Instagram 的生产级字幕转写 API,带 AI 转写兜底。
v1.0.0
Apify MCP pinned to 30 Steadyfetch actors: ad, video and audio transcripts, trends, jobs, Amazon
v3.0.0
10 个按次付费的网页工具:Markdown 转换、CSS 抓取、链接提取、爬取、截图、PDF 与浏览器。
v1.0.0
Reddit posts, comments, search, users & subreddits as JSON via Apify. No Reddit API key. $2/1k
v1.0.0
HTML entity count, markup discarded
v1.0.0
HTML canonical link count, markup discarded
v1.0.0
html lang attribute shape, markup discarded
v1.0.0
HTML link rel count, markup discarded
v1.0.0
HTML script tag count, markup discarded
v1.0.0
HTML meta tag count, markup discarded
v1.0.0
HTML h1 count, markup discarded
v1.0.0
HTML script tag count, markup discarded
v1.0.0
HTML meta tag count, markup discarded
静态或服务端渲染的页面、整站文档爬取用抓取类(Firecrawl/Fetch);需要点击、滚动、登录交互或抓渲染后截图,用浏览器自动化(Playwright MCP)。
托管版需要;也可以自部署开源版,或用免费的 Fetch MCP 做单页抓取。资源详情页会标注分发方式与所需环境变量。
常见做法是 Markdown 清洗后写入向量库 MCP(Qdrant、Chroma、pgvector 等),再配合检索 Skill 或 Context7 类知识库 MCP 完成问答。