Voice Message
v1.0.4
io.clawhub.xmanrui/voice-message
Send voice messages across chat channels (Telegram, Discord, Feishu/Lark, Signal, WhatsApp, and others) using edge-tts for text-to-speech and ffmpeg for audi...
“Audio” 共 303 个结果
v1.0.4
io.clawhub.xmanrui/voice-message
Send voice messages across chat channels (Telegram, Discord, Feishu/Lark, Signal, WhatsApp, and others) using edge-tts for text-to-speech and ffmpeg for audi...
v0.1.0
io.clawhub.thegovind/azure-ai-voicelive-py
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
v0.1.2
io.clawhub.kalvinrv/seedance-v2
Seedance 2.0 Pro on RunComfy. Seedance 2.0 Pro (ByteDance Seedance v2) is a multi-modal cinematic short-form video model with native lip-sync audio. This ski...
v1.0.0
io.clawhub.gugic/inworld-tts
Text-to-speech via Inworld.ai API. Use when generating voice audio from text, creating spoken responses, or converting text to MP3/audio files. Supports multiple voices, speaking rates, and streaming for long text.
v1.0.11
io.clawhub.mogens9/ai-podcast
Generate AI podcast episodes from PDFs, text, notes, and links using MagicPodcast in OpenClaw. Creates natural two-person dialogue audio, supports custom lan...
v0.1.0
io.clawhub.apollo1234/yt-dlp-downloader-skill
Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp. Use when the user provides a video URL and wants to download it, extract audio (MP3), download subtitles, or select video quality. Triggers on phrases like "下载视频", "download video", "yt-dlp", "YouTube", "B站", "抖音", "提取音频", "extract audio".
v1.0.1
io.clawhub.xsir0/google-gemini-media
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
v1.0.0
io.clawhub.al-one/edge-tts-uvx
Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
v1.1.1
io.clawhub.legendarylibr/seiso
Unified media generation gateway for agents. Discover tools dynamically, choose API key or x402 auth, invoke image/video/audio/music/3D/training tools, and h...
v0.1.0
io.clawhub.agmmnn/fal-ai
Generate images, videos, and audio via fal.ai API (FLUX, SDXL, Whisper, etc.)
v1.0.1
io.clawhub.ivangdavila/podcast
Create and grow podcasts by planning episodes, producing audio or video, generating clips, and building audience across formats.
v1.3.15
io.clawhub.dlazyai/dlazy-jimeng-omnihuman-1-5
Generate realistic digital human broadcast videos from portrait images and audio/text using Jimeng OmniHuman 1.5. 使用即梦 (Jimeng) OmniHuman 1.5 模型,通过人像图片和音频/文本生成逼真的数字人播报视频。
v0.1.0
io.clawhub.kalvinrv/lipsync
Lip-sync a face to a specific audio track on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar from a portrai...
v1.1.8
io.clawhub.matusvojtek/tubescribe
YouTube video summarizer with speaker detection, formatted documents, and audio output. Works out of the box with macOS built-in TTS. Optional recommended tools (pandoc, ffmpeg, mlx-audio) enhance quality. Requires internet for YouTube access. No paid APIs or subscriptions. Use when user sends a YouTube URL or asks to summarize/transcribe a YouTube video.
v1.0.0
io.clawhub.abhishek-official1/clawvox
ClawVox - ElevenLabs voice studio for OpenClaw. Generate speech, transcribe audio, clone voices, create sound effects, and more.
v1.1.5
io.clawhub.gizmogremlin/voice-ai-voices
High-quality voice synthesis with 9 personas, 11 languages, and streaming using Voice.ai API.
v1.0.1
io.clawhub.paulasjes/elevenlabs-transcribe
Transcribe audio to text using ElevenLabs Scribe. Supports batch transcription, realtime streaming from URLs, microphone input, and local files.
v0.6.0
io.clawhub.kkaticld/listenhub-ai
Turn ideas into podcasts, explainer videos, voice narration, and AI images via ListenHub. Use when the user wants to "make a podcast", "create an explainer v...
v0.1.0
io.clawhub.oconnell-carl/notebooklm-cli
Command-line interface to manage Google NotebookLM notebooks, sources, and generate audio, quizzes, reports, presentations, and visual study materials progra...
v1.0.0
io.clawhub.liudu2326526/ffmpeg-master
Use when performing video/audio processing tasks including transcoding, filtering, streaming, metadata manipulation, or complex filtergraph operations with FFmpeg.
v0.1.5
io.clawhub.kalvinrv/happyhorse-1-0
HappyHorse 1.0 — text-to-video generation on RunComfy. HappyHorse 1.0 is currently #1 on Artificial Analysis Video Arena and produces native 1080p video with...
v0.1.0
io.clawhub.kalvinrv/kling-3-0
Kling 3.0 video generation on RunComfy. Kling 3.0 (also called Kling V3.0) is Kuaishou Technology's third-generation multi-shot video model with native synch...
v0.1.0
io.clawhub.kalvinrv/controlnet-pose
Pose-conditioned generation on RunComfy via the `runcomfy` CLI. Routes across Kling 2-6 Motion Control Pro / Standard (transfer the motion / blocking of a re...
v1.0.2
io.clawhub.javicasper/transcribe
Transcribe audio files to text using local Whisper (Docker). Use when receiving voice messages, audio files (.mp3, .m4a, .ogg, .wav, .webm), or when asked to transcribe audio content.