seedance-2-video-gen
v2.1.0
io.clawhub.evolinkai/seedance-2-video-gen
Seedance 2.0 AI video generation via EvoLink API. Three modes — text-to-video, image-to-video (1-2 images), reference-to-video (images + videos + audio). Aut...
“Audio Generation” 共 263 个结果
v2.1.0
io.clawhub.evolinkai/seedance-2-video-gen
Seedance 2.0 AI video generation via EvoLink API. Three modes — text-to-video, image-to-video (1-2 images), reference-to-video (images + videos + audio). Aut...
v1.3.14
io.clawhub.dlazyai/dlazy-gemini-2-5-tts
Generate multilingual, highly natural audio using Gemini 2.5 text-to-speech. 使用 Gemini 2.5 强大的文本转语音能力,生成多语言、高自然度的音频。
v1.0.0
io.clawhub.dagmawibabi/addis-assistant-stt
Provides Speech-to-Text (STT) and text Translation using the Addis Assistant API (api.addisassistant.com). Use when the user needs to convert an audio file to text (specifically Amharic), or translate text between languages (e.g., Amharic to English). Requires 'x-api-key'.
v1.0.0
io.clawhub.amstko/tts
Convert text to speech using Hume AI (or OpenAI) API. Use when the user asks for an audio message, a voice reply, or to hear something "of vive voix".
v1.0.1
io.clawhub.1999azzar/yt-dlp
A robust CLI wrapper for yt-dlp to download videos, playlists, and audio from YouTube and thousands of other sites. Supports format selection, quality control, metadata embedding, and cookie authentication.
v2.23.0
io.clawhub.michaelwang11394/video-agent
[DEPRECATED — uses outdated v1/v2 endpoints] Use `create-video` for prompt-based video generation (v3 Video Agent) or `avatar-video` for precise avatar/scene...
v1.0.0
io.clawhub.paki81/qwen-tts
Local text-to-speech using Qwen3-TTS-12Hz-1.7B-CustomVoice. Use when generating audio from text, creating voice messages, or when TTS is requested. Supports 10 languages including Italian, 9 premium speaker voices, and instruction-based voice control (emotion, tone, style). Alternative to cloud-based TTS services like ElevenLabs. Runs entirely offline after initial model download.
v0.1.5
io.clawhub.okaris/speech-to-text
Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large. Capabilities: transcription, translation,...
v1.0.0
io.clawhub.clawdbotborges/elevenlabs-stt
Transcribe audio files using ElevenLabs Speech-to-Text (Scribe v2).
v1.0.0
io.clawhub.hopyky/tts-whatsapp
Send high-quality text-to-speech voice messages on WhatsApp in 40+ languages with automatic delivery
v2.1.0
io.clawhub.jimliu/baoyu-image-gen
AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs...
v1.0.0
io.clawhub.rubenfb23/walkie-talkie
Handles voice-to-voice conversations on WhatsApp. Automatically transcribes incoming audio and responds with local TTS audio. Use when the user wants to "talk" instead of type.
v1.1.0
io.clawhub.araa47/gemini-stt
Transcribe audio files using Google's Gemini API or Vertex AI
v1.0.0
io.clawhub.robin797860/qwen-image
Generate images using Qwen Image API (Alibaba Cloud DashScope). Use when users request image generation with Chinese prompts or need high-quality AI-generated images from text descriptions.
v0.1.2
io.clawhub.thewulf7/avatar-video-messages
Generate and send video messages with a lip-syncing VRM avatar. Use when user asks for video message, avatar video, video reply, or when TTS should be delivered as video instead of audio.
v1.0.1
io.clawhub.clawdbotborges/elevenlabs-music
Generate music from text prompts using ElevenLabs Eleven Music API. Use when creating songs, soundtracks, jingles, lullabies, or any audio music from descriptions. Supports vocals with AI-generated lyrics, instrumental tracks, and multiple genres/styles. Requires paid ElevenLabs plan.
v1.0.0
io.clawhub.zhang-shubo/image2prompt
Analyze images and generate detailed prompts for image generation. Supports portrait, landscape, product, animal, illustration categories with structured or natural output.
v1.0.0
io.clawhub.fspecii/ace-music
Generate AI music using ACE-Step 1.5 via ACE Music's free API. Use when the user asks to create, generate, or compose music, songs, beats, instrumentals, or...
v1.3.15
io.clawhub.dlazyai/dlazy-generate
A comprehensive generation skill. Can generate images, videos, and audio by automatically selecting the appropriate dlazy CLI model. 综合生成技能。能够根据用户意图自动选择合适的 dlazy CLI 模型来生成图片、视频或音频。
v0.1.3
io.clawhub.tanchunsiong/zoom-meeting-assistance-with-rtms-unofficial-community-skill
Zoom RTMS Meeting Assistant — start on-demand to capture meeting audio, video, transcript, screenshare, and chat via Zoom Real-Time Media Streams. Handles meeting.rtms_started and meeting.rtms_stopped webhook events. Provides AI-powered dialog suggestions, sentiment analysis, and live summaries with WhatsApp notifications. Use when a Zoom RTMS webhook fires or the user asks to record/analyze a meeting.
v1.0.0
io.clawhub.instant-picture/clonev
Clone any voice and generate speech using Coqui XTTS v2. SUPER SIMPLE - provide a voice sample (6-30 sec WAV) and text, get cloned voice audio. Supports 14+ languages. Use when the user wants to (1) Clone their voice or someone else's voice, (2) Generate speech that sounds like a specific person, (3) Create personalized voice messages, (4) Multi-lingual voice cloning (speak any language with cloned voice).
v0.1.2
io.clawhub.kalvinrv/image-to-video-runcomfy
Image-to-video generation on RunComfy. This image-to-video skill turns any still image into a short video clip via the RunComfy Model API. The image-to-video...
v1.0.0
io.clawhub.ivangdavila/ffmpeg
Process video and audio with correct codec selection, filtering, and encoding settings.
v1.0.0
io.clawhub.yuf1011/xiaohongshu-publisher
Draft and publish posts to 小红书 (Xiaohongshu/RED). Use when creating content for 小红书, drafting posts, generating cover images, or publishing via browser automation. Covers the full workflow from content creation to browser-based publishing, including cover image generation with Pillow.