🫧 Seedance 2.0 Pro — Pro Pack on RunComfy
v0.1.2
io.clawhub.kalvinrv/seedance-v2
Seedance 2.0 Pro on RunComfy. Seedance 2.0 Pro (ByteDance Seedance v2) is a multi-modal cinematic short-form video model with native lip-sync audio. This ski...
“Audio Generation” 共 393 个结果
v0.1.2
io.clawhub.kalvinrv/seedance-v2
Seedance 2.0 Pro on RunComfy. Seedance 2.0 Pro (ByteDance Seedance v2) is a multi-modal cinematic short-form video model with native lip-sync audio. This ski...
v1.3.15
io.clawhub.dlazyai/dlazy-seedream-5-0-lite
Fast image generation with Doubao Seedream 5.0 Lite. Supports text-to-image and image-to-image. 使用豆包 Seedream 5.0 Lite 极速生成图像,支持文生图与图生图。
v1.0.0
io.clawhub.gugic/inworld-tts
Text-to-speech via Inworld.ai API. Use when generating voice audio from text, creating spoken responses, or converting text to MP3/audio files. Supports multiple voices, speaking rates, and streaming for long text.
v1.2.1
io.clawhub.iisweetheartii/agent-selfie
AI agent self-portrait generator. Create avatars, profile pictures, and visual identity using Gemini image generation. Supports mood-based generation, season...
v0.1.0
io.clawhub.apollo1234/yt-dlp-downloader-skill
Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp. Use when the user provides a video URL and wants to download it, extract audio (MP3), download subtitles, or select video quality. Triggers on phrases like "下载视频", "download video", "yt-dlp", "YouTube", "B站", "抖音", "提取音频", "extract audio".
v1.0.1
io.clawhub.xsir0/google-gemini-media
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
v1.0.0
io.clawhub.rhanbourinajd/ai-video-gen
End-to-end AI video generation - create videos from text prompts using image generation, video synthesis, voice-over, and editing. Supports OpenAI DALL-E, Replicate models, LumaAI, Runway, and FFmpeg editing.
v1.0.0
io.clawhub.al-one/edge-tts-uvx
Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
v1.1.1
io.clawhub.legendarylibr/seiso
Unified media generation gateway for agents. Discover tools dynamically, choose API key or x402 auth, invoke image/video/audio/music/3D/training tools, and h...
v2.2.1
io.clawhub.a1024708231/video-producer
短视频一键生成技能 v2.2。调用video-director进行画面规划,然后生成AI素材、TTS配音、视频渲染,输出完整MP4。
v0.2.4
io.clawhub.fossilizedcarlos/krea-api
Generate images via Krea.ai API (Flux, Imagen, Ideogram, Seedream, etc.)
v0.1.0
io.clawhub.agmmnn/fal-ai
Generate images, videos, and audio via fal.ai API (FLUX, SDXL, Whisper, etc.)
v1.0.1
io.clawhub.ivangdavila/podcast
Create and grow podcasts by planning episodes, producing audio or video, generating clips, and building audience across formats.
v1.3.15
io.clawhub.dlazyai/dlazy-jimeng-omnihuman-1-5
Generate realistic digital human broadcast videos from portrait images and audio/text using Jimeng OmniHuman 1.5. 使用即梦 (Jimeng) OmniHuman 1.5 模型,通过人像图片和音频/文本生成逼真的数字人播报视频。
v0.1.0
io.clawhub.delorenj/fal-text-to-image
Generate, remix, and edit images using fal.ai's AI models. Supports text-to-image generation, image-to-image remixing, and targeted inpainting/editing.
v0.1.0
io.clawhub.kalvinrv/lipsync
Lip-sync a face to a specific audio track on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar from a portrai...
v1.1.8
io.clawhub.matusvojtek/tubescribe
YouTube video summarizer with speaker detection, formatted documents, and audio output. Works out of the box with macOS built-in TTS. Optional recommended tools (pandoc, ffmpeg, mlx-audio) enhance quality. Requires internet for YouTube access. No paid APIs or subscriptions. Use when user sends a YouTube URL or asks to summarize/transcribe a YouTube video.
v1.0.0
io.clawhub.abhishek-official1/clawvox
ClawVox - ElevenLabs voice studio for OpenClaw. Generate speech, transcribe audio, clone voices, create sound effects, and more.
v1.1.5
io.clawhub.gizmogremlin/voice-ai-voices
High-quality voice synthesis with 9 personas, 11 languages, and streaming using Voice.ai API.
v1.0.1
io.clawhub.paulasjes/elevenlabs-transcribe
Transcribe audio to text using ElevenLabs Scribe. Supports batch transcription, realtime streaming from URLs, microphone input, and local files.
v0.6.0
io.clawhub.kkaticld/listenhub-ai
Turn ideas into podcasts, explainer videos, voice narration, and AI images via ListenHub. Use when the user wants to "make a podcast", "create an explainer v...
v0.1.0
io.clawhub.oconnell-carl/notebooklm-cli
Command-line interface to manage Google NotebookLM notebooks, sources, and generate audio, quizzes, reports, presentations, and visual study materials progra...
v1.0.0
io.clawhub.xtaq/liblib-ai-gen
Generate images with Seedream4.5 and videos with Kling via LiblibAI API. Use when user asks to generate/create images, pictures, illustrations, or videos using LiblibAI, Seedream, or Kling models.
v1.0.0
io.clawhub.liudu2326526/ffmpeg-master
Use when performing video/audio processing tasks including transcoding, filtering, streaming, metadata manipulation, or complex filtergraph operations with FFmpeg.