audio
99 MCP servers and Agent skills related to audio, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Markdown Converter
v1.0.0
io.clawhub.steipete/markdown-converter
Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
Openai Whisper Api
v1.0.0
io.clawhub.steipete/openai-whisper-api
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
Edge TTS
v2.0.0
io.clawhub.i3130002/edge-tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
cellcog
v2.0.21
io.clawhub.cellcog/cellcog
Any-to-any AI sub-agent — research, images, video, audio, music, podcasts, avatars, voice cloning, documents, spreadsheets, dashboards, 3D models, diagrams, and code in one request. Agent-to-agent protocol with multi-step iteration for high accuracy. #1 on DeepResearch Bench (Apr 2026) — deep reasoning meets all modalities, so all your work gets done, not just code.
ffmpeg-video-editor
v1.0.0
io.clawhub.mahmoudadelbghany/ffmpeg-video-editor
Generate FFmpeg commands from natural language video editing requests - cut, trim, convert, compress, change aspect ratio, extract audio, and more.
Cellcog
v2.0.15
io.clawhub.nitishgargiitd/cellcog
Any-to-any AI sub-agent — research, images, video, audio, music, podcasts, avatars, voice cloning, documents, spreadsheets, dashboards, 3D models, diagrams.
Video Transcript Downloader
v1.0.0
io.clawhub.steipete/video-transcript-downloader
Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site. Use when asked to “download this video”, “save this clip”, “rip audio”, “get subtitles”, “get transcript”, or to troubleshoot yt-dlp/ffmpeg and formats/playlists.
Songsee
v1.0.0
io.clawhub.steipete/songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
Wonda
v1.2.0
io.clawhub.degausai/wonda
Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation
Video Subtitles
v1.0.0
io.clawhub.ngutman/video-subtitles
Generate SRT subtitles from video/audio with translation support. Transcribes Hebrew (ivrit.ai) and English (whisper), translates between languages, burns subtitles into video. Use for creating captions, transcripts, or hardcoded subtitles for WhatsApp/social media.
Yt Dlp Downloader
v0.1.0
io.clawhub.apollo1234/yt-dlp-downloader-skill
Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp. Use when the user provides a video URL and wants to download it, extract audio (MP3), download subtitles, or select video quality. Triggers on phrases like "下载视频", "download video", "yt-dlp", "YouTube", "B站", "抖音", "提取音频", "extract audio".
Kokoro TTS
v0.1.0
io.clawhub.edkief/kokoro-tts
Generate spoken audio from text using the local Kokoro TTS engine. Use when the user asks to "say" something, requests a voice message, or wants text converted to speech.
OpenAI TTS
v1.0.0
io.clawhub.pors/openai-tts
Text-to-speech via OpenAI Audio Speech API.
Elevenlabs Tts
v2.4.0
io.clawhub.shaharsha/elevenlabs-tts
ElevenLabs TTS - the best ElevenLabs integration for OpenClaw.
Audio Generation
v1.0.17
io.clawhub.cellcog/audio-generation-cellcog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Three voice providers (OpenAI, ElevenLabs, MiniMax). Professional audio production from text prompts.
FFmpeg CLI
v1.0.0
io.clawhub.ascendswang/ffmpeg-cli
Process video and audio using FFmpeg CLI for transcoding, cutting, merging, audio extraction, thumbnails, GIFs, speed, filters, subtitles, and watermarks.
FFmpeg
v1.0.0
io.clawhub.ivangdavila/ffmpeg
Process video and audio with correct codec selection, filtering, and encoding settings.
Voice Transcribe
v1.0.1
io.clawhub.darinkishore/voice-transcribe
Transcribe audio files using OpenAI's gpt-4o-mini-transcribe model with vocabulary hints and text replacements. Requires uv (https://docs.astral.sh/uv/).
ListenHub
v0.6.0
io.clawhub.kkaticld/listenhub-ai
Turn ideas into podcasts, explainer videos, voice narration, and AI images via ListenHub.
Seedance Video Generation
v1.0.3
io.clawhub.jackycser/seedance-video-generation
Generate AI videos using ByteDance Seedance. Use when the user wants to: (1) generate videos from text prompts, (2) generate videos from images (first frame, first+last frame, reference images), or (3) query/manage video generation tasks. Supports Seedance 1.5 Pro (with audio), 1.0 Pro, 1.0 Pro Fast, and 1.0 Lite models.
Elevenlabs
v1.3.4
io.clawhub.odrobnik/elevenlabs
Text-to-speech, sound effects, music generation, voice management, and quota checks via the ElevenLabs API.
Audio Cog
v1.0.12
io.clawhub.nitishgargiitd/audio-cog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue.
Audio
v1.0.1
io.clawhub.ivangdavila/audio
Process, enhance, and convert audio files with noise removal, normalization, format conversion, transcription, and podcast workflows.
Fal.ai API
v0.1.0
io.clawhub.agmmnn/fal-ai
Generate images, videos, and audio via fal.ai API (FLUX, SDXL, Whisper, etc.)
TubeScribe
v1.1.8
io.clawhub.matusvojtek/tubescribe
YouTube video summarizer with speaker detection, formatted documents, and audio output. Works out of the box with macOS built-in TTS. Optional recommended tools (pandoc, ffmpeg, mlx-audio) enhance quality. Requires internet for YouTube access. No paid APIs or subscriptions. Use when user sends a YouTube URL or asks to summarize/transcribe a YouTube video.
Seedance 2.0 prompt-engineering skill
v2.0.0
io.clawhub.dandysuper/seedance-2-prompt-engineering-skill
Generate precise, timecoded Seedance 2.0 prompts integrating multimodal inputs with asset mapping for controlled 4-15s video creation and editing.
Tts
v1.0.0
io.clawhub.amstko/tts
Convert text to speech using Hume AI (or OpenAI) API. Use when the user asks for an audio message, a voice reply, or to hear something "of vive voix".
ACE Music - Free Suno Alternative Generate unlimited AI music for free using ACE-Step 1.5. Full songs with vocals, lyrics, any genre, any language. No subscription, no credits, no limits. The open-source Suno alternative, powered by ACE Music's free API.
v1.0.0
io.clawhub.fspecii/ace-music
Generate AI music using ACE-Step 1.5 via ACE Music's free API.
ElevenLabs Speech-to-Text
v1.0.0
io.clawhub.clawdbotborges/elevenlabs-stt
Transcribe audio files using ElevenLabs Speech-to-Text (Scribe v2).
Voice Reply
v1.0.0
io.clawhub.stolot0mt0m/voice-reply
Local text-to-speech using Piper voices via sherpa-onnx. 100% offline, no API keys required. Use when user asks for a voice reply, audio response, spoken answer, or wants to hear something read aloud. Supports multiple languages including German (thorsten) and English (ryan) voices. Outputs Telegram-compatible voice notes with [[audio_as_voice]] tag.
Voice
v1.0.1
io.clawhub.zhaov1976/voice
Convert text to speech using Microsoft Edge's TTS engine with customizable voices, direct playback, and automatic temporary file cleanup.
Google Gemini Media
v1.0.1
io.clawhub.xsir0/google-gemini-media
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
Voice Message
v1.0.4
io.clawhub.xmanrui/voice-message
Send voice messages across chat channels (Telegram, Discord, Feishu/Lark, Signal, WhatsApp.
notebooklm-cli
v0.1.0
io.clawhub.oconnell-carl/notebooklm-cli
Command-line interface to manage Google NotebookLM notebooks, sources, and generate audio, quizzes, reports, presentations.
Transcribe
v1.0.2
io.clawhub.javicasper/transcribe
Transcribe audio files to text using local Whisper (Docker). Use when receiving voice messages, audio files (.mp3, .m4a, .ogg, .wav, .webm), or when asked to transcribe audio content.
Music Generation
v1.0.0
io.clawhub.ivangdavila/music-generation
Generate AI music with optimized prompts, style control, and production-ready audio output.
Flyworks Avatar Video
v1.0.0
io.clawhub.linhui99/flyworks-avatar-video
Generate videos using Flyworks (a.k.a HiFly) Digital Humans. Create talking photo videos from images, use public avatars with TTS, or clone voices for custom audio.
Vocal Chat
v1.0.0
io.clawhub.rubenfb23/vocal-chat
Handles voice-to-voice conversations on WhatsApp. Automatically transcribes incoming audio and responds with local TTS audio. Use when the user wants to "talk" instead of type.
Transcribe audio files via OpenRouter using audio-capable models
v1.0.0
io.clawhub.obviyus/openrouter-transcribe
Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).
Speech To Text
v0.1.5
io.clawhub.okaris/speech-to-text
Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large.
AudioPod
v1.2.3
io.clawhub.rakesh1002/audiopod
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.
Qwen3-tts
v1.0.0
io.clawhub.paki81/qwen-tts
Local text-to-speech using Qwen3-TTS-12Hz-1.7B-CustomVoice. Use when generating audio from text, creating voice messages, or when TTS is requested. Supports 10 languages including Italian, 9 premium speaker voices, and instruction-based voice control (emotion, tone, style). Alternative to cloud-based TTS services like ElevenLabs. Runs entirely offline after initial model download.
TTS WhatsApp
v1.0.0
io.clawhub.hopyky/tts-whatsapp
Send high-quality text-to-speech voice messages on WhatsApp in 40+ languages with automatic delivery
Gemini STT
v1.1.0
io.clawhub.araa47/gemini-stt
Transcribe audio files using Google's Gemini API or Vertex AI
Video Messages from your openclaw
v0.1.2
io.clawhub.thewulf7/avatar-video-messages
Generate and send video messages with a lip-syncing VRM avatar. Use when user asks for video message, avatar video, video reply, or when TTS should be delivered as video instead of audio.
Audio Content Generator
v1.0.0
io.clawhub.udiedrichsen/audio-gen
Generate audiobooks, podcasts, or educational audio content on demand. User provides an idea or topic, Claude AI writes a script, and ElevenLabs converts it to high-quality audio. Supports multiple formats (audiobook, podcast, educational), custom lengths, and voice effects. Use when asked to create audio content, make a podcast, generate an audiobook, or produce educational audio. Returns MP3 audio file via MEDIA token.
声音克隆 Vidu Audio Clone
v1.3.22
io.clawhub.dlazyai/dlazy-vidu-audio-clone
Clone voice and generate new text reading audio with one click using Vidu Audio Clone. 使用 Vidu 声音克隆技术,通过参考音频一键复制音色并生成新文本的朗读音频。
视频生成 Seedance 2.0
v1.3.21
io.clawhub.dlazyai/dlazy-seedance-2-0
ByteDance's latest video generation model. Supports multi-modal reference (images, video, audio) to generate videos, as well as first/last frame and text-to-video modes. 字节跳动最新视频生成模型 Seedance 2.0,支持多模态参考(图片 + 视频 + 音频)生视频、首尾帧及文生视频,适合高质量多样化视频创作。