ocr
26 MCP servers and Agent skills related to ocr, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Markdown Converter
v1.0.0
io.clawhub.steipete/markdown-converter
Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
OCR - Local (No API Key)
v1.0.0
io.clawhub.shaw555/ocr-local
Extract text from images using Tesseract.js OCR (100% local, no API key required). Supports Chinese (simplified/traditional) and English.
Image Ocr
v1.0.0
io.clawhub.xejrax/image-ocr
Extract text from images using Tesseract OCR
Tencent COS
v1.1.9
io.clawhub.shawnminh/tencent-cos-skill
腾讯云对象存储(COS)和数据万象(CI)集成技能。覆盖文件存储管理、AI处理和知识库三大核心场景。 存储场景:上传文件到云端、下载云端文件、批量管理存储桶文件、获取文件签名链接分享、查看文件元信息、查询数据万象及子服务开通状态。 图片处理场景:图片质量评估打分、AI超分辨率放大、AI智能裁剪、二维码/条形码识别、添加文字水印、获取图片EXIF信息、缩放、裁剪、旋转、格式转换。 文档处理场景:Word/Excel/PPT等办公文档转PDF、文档预览。 媒体处理场景:视频智能封面提取、视频转码、视频截帧、获取媒体信息。 内容审核场景:图片/视频/音频/文本/文档内容审核,检测违规内容。 智能语音
QVeris Official
v1.0.9
io.clawhub.linfangw/qveris-official
QVeris is a capability discovery and tool calling engine.
Image Vision
v1.0.0
io.clawhub.cntuang/image-vision
Analyze and interpret images by describing content, extracting text, answering questions, comparing visuals, and extracting structured data from JPG, PNG.
DeepRead OCR
v1.1.0
io.clawhub.uday390/deepread-ocr
AI-native OCR platform that turns documents into high-accuracy data in minutes.
MinerU PDF Parser
v1.0.1
io.clawhub.easonai-5589/mineru
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Tesseract Ocr
v1.0.0
io.clawhub.whalefell/tesseract-ocr
Extract text from images using the Tesseract OCR engine directly via command line. Supports multiple languages including Chinese, English, and more.
mineru document extractor
v0.1.30
io.clawhub.mineru-extract/mineru-document-extractor
MinerU document extraction — convert PDFs, scanned documents, images, Word (DOC/DOCX), PowerPoint (PPT/PPTX), Excel (XLS/XLSX).
Nanonets OCR
v1.0.2
io.clawhub.shhdwi/docstrange
Document extraction API by Nanonets. Convert PDFs and images to markdown, JSON, or CSV with confidence scoring. Use when you need to OCR documents, extract invoice fields, parse receipts, or convert tables to structured data.
Screenshot Ocr
v1.0.0
io.clawhub.sxliuyu/screenshot-ocr
截图 OCR 识别工具。截图→自动识别文字→复制/保存,适合提取图片内容、表格数据、验证码。
pdf-ocr-layout
v1.0.2
io.clawhub.baokui/pdf-ocr-layout
基于智谱 GLM-OCR、GLM-4.7 及 GLM-4.6V 的多模态文档深度解析工具。 Use when: - 需要高精度提取文档(PDF/图片)中的表格并转换为 Markdown 格式 - 需要从文档页面中自动裁剪并提取插图、图表为独立文件 - 需要对提取的图表进行深度语义理解(基于 GLM-4.6V 视觉分析) - 需要对提取的表格数据进行逻辑分析(基于 GLM-4.7 文本分析) 核心架构: 1. 视觉提取:GLM-OCR 2. 语义理解:GLM-4.7 (纯文本/表格) + GLM-4.6V (多模态/图像)
v1.0.0
io.clawhub.yang1002378395-cmyk/pdf-processor-cn
Use this skill whenever the user wants to do anything with PDF files.
Pdf To Structured
v2.0.0
io.clawhub.datadrivenconstruction/pdf-to-structured
Extract structured data from construction PDFs. Convert specifications, BOMs, schedules, and reports from PDF to Excel/CSV/JSON. Use OCR for scanned documents and pdfplumber for native PDFs.
PaddleOCR Text Recognition
v2.0.0
io.clawhub.bobholamovic/paddleocr-text-recognition
Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs.
OCR with python
v1.0.0
io.clawhub.roamer-remote/ocr-python
Extract Chinese and English text from images and scanned PDFs, including documents like invoices and contracts, using PaddleOCR in Python.
Ocr Document
v1.0.0
io.clawhub.tanis90/ocr-document
OCR document extraction - extract text from scanned documents, photos, and images using OCR.
Nutrient OpenClaw
v1.3.0
io.clawhub.jdrhyne/nutrient-openclaw
Use the pinned Nutrient OpenClaw plugin to convert, OCR, extract, redact, watermark, sign, or inspect the last-known local credit record for documents. Route only OpenClaw document-processing requests to its declared tools. Treat processing as an external, credit-consuming DWS transfer that requires a bounded estimate and action-time confirmation for every invocation.
Image OCR Reader
v1.0.0
io.clawhub.igetmm/image-ocr-reader
Extract text from images using OCR with support for Chinese and English in common formats like jpg, png, and jpeg.
夸克扫描王 OCR文字识别 - yescan ocr universal
v1.0.14
io.clawhub.yescan-ai/yescan-ocr-universal
由夸克扫描王提供的高准取率 OCR 文字识别工具,支持印刷、手写、表格、多语言、公式等各种场景。支持图片、截图、扫描件中的文字提取,包括手写文档、表格内容、数学公式、商品图片等复杂场景。精准识别各类证件(身份证、社保卡、驾驶证、行驶证、港澳通行证、学位证等证件)及票据(增值税发票、火车票、英文发票等票据),同时支持医疗报告单、营业执照、习题题目等专业文档识别。
Japanese Tutor
v1.0.2
io.clawhub.chndranndr/japanese-tutor
Interactive Japanese learning assistant. Supports vocabulary, grammar, quizzes, roleplay, PDF/DOCX material parsing for study/homework help, and OCR translation.
MarkItDown Skill
v1.0.1
io.clawhub.karmanverma/markitdown-skill
OpenClaw agent skill for converting documents to Markdown. Documentation and utilities for Microsoft's MarkItDown library. Supports PDF, Word, PowerPoint, Excel, images (OCR), audio (transcription), HTML, YouTube.
Upstage Document Parse
v1.0.5
io.clawhub.upstage-deployment/upstage-document-parse
Parse documents (PDF, images, DOCX, PPTX, XLSX, HWP) into layout-aware markdown/HTML with tables, figures, headings.
Tencent Cloud COS
v1.2.0
io.clawhub.shawnminh/tencentcloud-cos
腾讯云对象存储(COS)和数据万象(CI)集成技能。覆盖文件存储管理、AI处理和知识库三大核心场景。 存储场景:上传文件到云端、下载云端文件、批量管理存储桶文件、获取文件签名链接分享、查看文件元信息、查询数据万象及子服务开通状态。 图片处理场景:图片质量评估打分、AI超分辨率放大、AI智能裁剪、二维码/条形码识别、添加文字水印、获取图片EXIF信息、缩放、裁剪、旋转、格式转换。 文档处理场景:Word/Excel/PPT等办公文档转PDF、文档预览。 媒体处理场景:视频智能封面提取、视频转码、视频截帧、获取媒体信息。 内容审核场景:图片/视频/音频/文本/文档内容审核,检测违规内容。 智能语音
Captcha Auto
v1.0.6
io.clawhub.annoyingc/captcha-auto
智能验证码自动识别 Skill - 混合模式(本地 Tesseract OCR + 阿里云千问 3 VL Plus)。支持两阶段输入框查找、安全隐私警告。用于网页自动化中的验证码识别、填写和提交。