web scraping
17 MCP servers and Agent skills related to web scraping, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Playwright Scraper Skill
v1.2.0
io.clawhub.waisimon/playwright-scraper-skill
Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.
Web Scraping
v1.0.0
io.clawhub.zhangqixin9527/web-scraping
Extract structured information from websites using web_fetch for simple pages and browser automation for dynamic sites, login-gated flows, pagination.
Web Scraper
v0.1.1
io.clawhub.guifav/web-scraper
Web scraping and content comprehension agent — multi-strategy extraction with cascade fallback, news detection, boilerplate removal, structured metadata.
Scrape
v1.0.0
io.clawhub.ivangdavila/scrape
Legal web scraping with robots.txt compliance, rate limiting, and GDPR/CCPA-aware data handling.
Scrapling
v1.0.8
io.clawhub.zendenho7/scrapling
Adaptive web scraping framework with anti-bot bypass and spider crawling.
Reddit Scraper
v1.0.0
io.clawhub.javicasper/reddit-scraper
Read and search Reddit posts via web scraping of old.reddit.com. Use when Clawdbot needs to browse Reddit content - read posts from subreddits, search for topics, monitor specific communities. Read-only access with no posting or comments.
Scrapling Web Scraping
v1.0.0
io.clawhub.zhengxinjipai/scrapling-web-scraper
Zero-bot-detection web scraping for OpenClaw. Bypass Cloudflare, handle JavaScript-heavy sites, and adapt to website changes automatically.
Crawl4ai
v1.0.0
io.clawhub.codylrn804/crawl4ai
AI-powered web scraping framework for extracting structured data from websites. Use when Codex needs to crawl, scrape, or extract data from web pages using AI-powered parsing, handle dynamic content, or work with complex HTML structures.
Firecrawler
v1.0.0
io.clawhub.capt-marbles/firecrawler
Web scraping and crawling with Firecrawl API. Fetch webpage content as markdown, take screenshots, extract structured data, search the web, and crawl documentation sites. Use when the user needs to scrape a URL, get current web info, capture a screenshot, extract specific data from pages, or crawl docs for a framework/library.
Crawl4AI Web Scraper
v1.0.1
io.clawhub.angusthefuzz/crawl-for-ai
Full web page scraping with JavaScript rendering via local Crawl4AI instance, delivering clean markdown or detailed JSON including links and media.
Firecrawl CLI
v1.0.0
io.clawhub.yash-kavaiya/firecrawl-cli
Web scraping, crawling, searching, and browser automation via the Firecrawl CLI (firecrawl).
defuddle-web-cleaner
v1.0.0
io.clawhub.extrastu/defuddle
Extract and clean readable article content, metadata, and markdown from URLs or HTML for research, note taking, and web scraping.
Bright Data
v1.0.0
io.clawhub.meirkad/bright-data
Web scraping and search via Bright Data API. Requires BRIGHTDATA_API_KEY and BRIGHTDATA_UNLOCKER_ZONE. Use for scraping any webpage as markdown (bypassing bot detection/CAPTCHA) or searching Google with structured results.
Actionbook
v0.1.1
io.clawhub.adcentury/actionbook
Activate when the user needs to interact with any website — browser automation, web scraping, screenshots, form filling, UI testing, monitoring, or building AI agents. Provides pre-verified page actions with step-by-step instructions and tested selectors.
Bits Browser Automation
v1.0.0
io.clawhub.robbiethompson18/bits
Control browser automation agents via the Bits MCP server. Use when running web scraping, form filling, data extraction, or any browser-based automation task. Bits agents can navigate websites, click elements, fill forms, handle OAuth flows, and extract structured data.
Luma Event Manager
v2.1.1
io.clawhub.mariovallereyes/luma-event-manager
Luma Event Manager for Clawdbot — Discover events by topic or location, RSVP, view guest lists, and sync to Google Calendar. No API key required (web scraping), no Luma Plus subscription needed.
XPR Web Scraping
v0.2.11
io.clawhub.paulgnz/xpr-web-scraping
Tools for fetching and extracting cleaned text, metadata, and links from single or multiple web pages with format options and link filtering.