transcription
20 MCP servers and Agent skills related to transcription, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Markdown Converter
v1.0.0
io.clawhub.steipete/markdown-converter
Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
Local Whisper
v1.0.0
io.clawhub.araa47/local-whisper
Local speech-to-text using OpenAI Whisper. Runs fully offline after model download. High quality transcription with multiple model sizes.
Youtube
v1.0.1
io.clawhub.grpaiva/youtube
Search YouTube videos, get channel info, fetch video details and transcripts using YouTube Data API v3 via MCP server or yt-dlp fallback.
YouTube Summarizer
v1.0.0
io.clawhub.abe238/youtube-summarizer
Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms. Detects YouTube URLs and provides metadata, key insights, and downloadable transcripts.
Faster Whisper
v1.5.1
io.clawhub.theplasmak/faster-whisper
Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription.
Audio
v1.0.1
io.clawhub.ivangdavila/audio
Process, enhance, and convert audio files with noise removal, normalization, format conversion, transcription, and podcast workflows.
TG Voice Whisper Transcriber
v1.0.0
io.clawhub.drones277/tg-voice-whisper
Automation skill for TG Voice Whisper Transcriber.
Transcribe audio files via OpenRouter using audio-capable models
v1.0.0
io.clawhub.obviyus/openrouter-transcribe
Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).
Speech To Text
v0.1.5
io.clawhub.okaris/speech-to-text
Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large.
AudioPod
v1.2.3
io.clawhub.rakesh1002/audiopod
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.
AssemblyAI advanced speech transcription
v1.0.1
io.clawhub.tristanmanchester/assemblyai-transcribe
Transcribe, diarise, translate, post-process, and structure audio/video with AssemblyAI.
Speech is Cheap Transcribe
v1.2.0
io.clawhub.ilyakam/asr
Fast, affordable automatic speech-to-text transcription supporting 100 languages, speaker diarization, word timestamps, and customizable output formats.
Audio Transcribe
v1.0.0
io.clawhub.aktheknight/audio-transcribe
Auto-transcribe voice messages locally using faster-whisper with selectable Whisper models, no API key required.
Elevenlabs Transcribe
v1.0.1
io.clawhub.paulasjes/elevenlabs-transcribe
Transcribe audio to text using ElevenLabs Scribe. Supports batch transcription, realtime streaming from URLs, microphone input, and local files.
MarkItDown Skill
v1.0.1
io.clawhub.karmanverma/markitdown-skill
OpenClaw agent skill for converting documents to Markdown. Documentation and utilities for Microsoft's MarkItDown library. Supports PDF, Word, PowerPoint, Excel, images (OCR), audio (transcription), HTML, YouTube.
Aliyun Asr
v1.0.10
io.clawhub.jixsonwang/aliyun-asr
Pure Aliyun ASR skill for voice message transcription, supports multiple channels including Feishu
it will help you to send voice messages to your AI Assistant and also can make it talk
v1.0.0
io.clawhub.amreahmed/elevenlabs-voice
Text-to-Speech and Speech-to-Text using ElevenLabs AI. Use when the user wants to convert text to speech, transcribe voice messages, or work with voice in multiple languages. Supports high-quality AI voices and accurate transcription.
Azure Ai Voicelive Py
v0.1.0
io.clawhub.thegovind/azure-ai-voicelive-py
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
Azure Ai Transcription Py
v0.1.0
io.clawhub.thegovind/azure-ai-transcription-py
Azure AI Transcription SDK for Python. Use for real-time and batch speech-to-text transcription with timestamps and diarization. Triggers: "transcription", "speech to text", "Azure AI Transcription", "TranscriptionClient".
Video Summary
v1.6.3
io.clawhub.lifei68801/video-summary
Video summarization for Bilibili, Xiaohongshu, Douyin, and YouTube. Extract insights from video content through transcription and summarization.