Audio
v1.0.1
io.clawhub.ivangdavila/audio
Process, enhance, and convert audio files with noise removal, normalization, format conversion, transcription, and podcast workflows.
“Audio” 共 118 个结果
v1.0.1
io.clawhub.ivangdavila/audio
Process, enhance, and convert audio files with noise removal, normalization, format conversion, transcription, and podcast workflows.
v1.3.16
io.clawhub.dlazyai/dlazy-vidu-audio-clone
Clone voice and generate new text reading audio with one click using Vidu Audio Clone. 使用 Vidu 声音克隆技术,通过参考音频一键复制音色并生成新文本的朗读音频。
v1.0.0
io.clawhub.aktheknight/audio-transcribe
Auto-transcribe voice messages locally using faster-whisper with selectable Whisper models, no API key required.
v0.2.2
io.clawhub.guoqiao/mlx-audio-server
Local 24x7 OpenAI-compatible API server for STT/TTS, powered by MLX on your Mac.
v1.1.0
io.clawhub.matrixy/audio-reply-skill
Generate audio replies using TTS. Trigger with "read it to me [public URL]" to fetch and read content aloud, or "talk to me [topic]" to generate a spoken res...
v1.0.17
io.clawhub.cellcog/audio-generation-cellcog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Three voice providers (OpenAI, ElevenLabs, MiniMax). Professional audio production from text prompts.
v1.3.8
io.clawhub.dlazyai/dlazy-kling-audio-clone
Generate customized speech that highly restores the timbre by uploading reference audio using Kling Audio Clone. 使用可灵 (Kling) 声音克隆模型,通过上传参考音频,生成高度还原该音色的定制语音。
v1.0.0
io.clawhub.obviyus/openrouter-transcribe
Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).
v1.3.18
io.clawhub.dlazyai/dlazy-audio-generate
Audio generation skill. Automatically selects the best dlazy CLI audio/TTS model based on the prompt. 音频生成技能。根据提示词自动选择最佳的 dlazy CLI 音频/TTS 模型。
v1.2.3
io.clawhub.rakesh1002/audiopod
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.
v1.2.0
io.clawhub.brokemac79/webchat-audio-notifications
Add browser audio notifications to Moltbot/Clawdbot webchat with 5 intensity levels - from whisper to impossible-to-miss (only when tab is backgrounded).
v1.0.0
io.clawhub.udiedrichsen/audio-gen
Generate audiobooks, podcasts, or educational audio content on demand. User provides an idea or topic, Claude AI writes a script, and ElevenLabs converts it to high-quality audio. Supports multiple formats (audiobook, podcast, educational), custom lengths, and voice effects. Use when asked to create audio content, make a podcast, generate an audiobook, or produce educational audio. Returns MP3 audio file via MEDIA token.
v1.0.12
io.clawhub.nitishgargiitd/audio-cog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Thre...
v1.1.0
io.clawhub.claudiodrusus/qr-gen
Generate QR codes from text, URLs, WiFi credentials, vCards, or any data. Use when the user wants to create a QR code, share a link as a scannable code, gene...
v1.0.0
io.clawhub.steipete/markdown-converter
Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
v2.0.21
io.clawhub.cellcog/cellcog
Any-to-any AI sub-agent — research, images, video, audio, music, podcasts, avatars, voice cloning, documents, spreadsheets, dashboards, 3D models, diagrams, and code in one request. Agent-to-agent protocol with multi-step iteration for high accuracy. #1 on DeepResearch Bench (Apr 2026) — deep reasoning meets all modalities, so all your work gets done, not just code.
v1.0.0
io.clawhub.steipete/openai-whisper-api
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
v1.0.0
io.clawhub.mahmoudadelbghany/ffmpeg-video-editor
Generate FFmpeg commands from natural language video editing requests - cut, trim, convert, compress, change aspect ratio, extract audio, and more.
v2.0.15
io.clawhub.nitishgargiitd/cellcog
Any-to-any AI sub-agent — research, images, video, audio, music, podcasts, avatars, voice cloning, documents, spreadsheets, dashboards, 3D models, diagrams,...
v1.0.0
io.clawhub.steipete/video-transcript-downloader
Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site. Use when asked to “download this video”, “save this clip”, “rip audio”, “get subtitles”, “get transcript”, or to troubleshoot yt-dlp/ffmpeg and formats/playlists.
v1.2.0
io.clawhub.degausai/wonda
Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation
v1.0.0
io.clawhub.steipete/songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
v2.0.0
io.clawhub.i3130002/edge-tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
v1.0.0
io.clawhub.pors/openai-tts
Text-to-speech via OpenAI Audio Speech API.