video-download
v0.1.6
io.clawhub.upupc/video-download
Download videos from 1800+ websites and generate subtitles using Faster Whisper AI. Use when user wants to download videos from YouTube, Bilibili, Twitter, T...
“Audio Generation” 共 824 个结果
v0.1.6
io.clawhub.upupc/video-download
Download videos from 1800+ websites and generate subtitles using Faster Whisper AI. Use when user wants to download videos from YouTube, Bilibili, Twitter, T...
v1.0.1
io.clawhub.merend/openai-tts-python
Text-to-speech conversion using OpenAI's TTS API for generating high-quality, natural-sounding audio. Supports 6 voices (alloy, echo, fable, onyx, nova, shimmer), speed control (0.25x-4.0x), HD quality model, multiple output formats (mp3, opus, aac, flac), and automatic text chunking for long content (4096 char limit per request). Use when: (1) User requests audio/voice output with triggers like "read this to me", "convert to audio", "generate speech", "text to speech", "tts", "narrate", "speak", or when keywords "openai tts", "voice", "podcast" appear. (2) Content needs to be spoken rather than read (multitasking, accessibility). (3) User wants specific voice preferences like "alloy", "echo", "fable", "onyx", "nova", "shimmer" or speed adjustments.
v1.3.15
io.clawhub.dlazyai/dlazy-kling-v3-omni
Versatile video generation with Kling v3 Omni. Supports multi-modal inputs to generate stunning dynamic videos. 使用可灵 (Kling) v3 Omni 全能视频生成模型,支持多模态输入(图片、提示词)生成震撼的动态视频。
v2.4.0
io.clawhub.shaharsha/elevenlabs-tts
ElevenLabs TTS - the best ElevenLabs integration for OpenClaw. ElevenLabs Text-to-Speech with emotional audio tags, ElevenLabs voice synthesis for WhatsApp,...
v1.3.14
io.clawhub.dlazyai/dlazy-seedance-2-0
ByteDance's latest video generation model. Supports multi-modal reference (images, video, audio) to generate videos, as well as first/last frame and text-to-video modes. 字节跳动最新视频生成模型 Seedance 2.0,支持多模态参考(图片 + 视频 + 音频)生视频、首尾帧及文生视频,适合高质量多样化视频创作。
v1.0.4
io.clawhub.xmanrui/voice-message
Send voice messages across chat channels (Telegram, Discord, Feishu/Lark, Signal, WhatsApp, and others) using edge-tts for text-to-speech and ffmpeg for audi...
v0.1.0
io.clawhub.thegovind/azure-ai-voicelive-py
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
v0.1.2
io.clawhub.kalvinrv/seedance-v2
Seedance 2.0 Pro on RunComfy. Seedance 2.0 Pro (ByteDance Seedance v2) is a multi-modal cinematic short-form video model with native lip-sync audio. This ski...
v1.3.15
io.clawhub.dlazyai/dlazy-seedream-5-0-lite
Fast image generation with Doubao Seedream 5.0 Lite. Supports text-to-image and image-to-image. 使用豆包 Seedream 5.0 Lite 极速生成图像,支持文生图与图生图。
v1.0.0
io.clawhub.gugic/inworld-tts
Text-to-speech via Inworld.ai API. Use when generating voice audio from text, creating spoken responses, or converting text to MP3/audio files. Supports multiple voices, speaking rates, and streaming for long text.
v1.2.1
io.clawhub.iisweetheartii/agent-selfie
AI agent self-portrait generator. Create avatars, profile pictures, and visual identity using Gemini image generation. Supports mood-based generation, season...
v0.1.0
io.clawhub.apollo1234/yt-dlp-downloader-skill
Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp. Use when the user provides a video URL and wants to download it, extract audio (MP3), download subtitles, or select video quality. Triggers on phrases like "下载视频", "download video", "yt-dlp", "YouTube", "B站", "抖音", "提取音频", "extract audio".
v1.0.1
io.clawhub.xsir0/google-gemini-media
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
v1.0.0
io.clawhub.rhanbourinajd/ai-video-gen
End-to-end AI video generation - create videos from text prompts using image generation, video synthesis, voice-over, and editing. Supports OpenAI DALL-E, Replicate models, LumaAI, Runway, and FFmpeg editing.
v1.0.0
io.clawhub.al-one/edge-tts-uvx
Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
v1.1.1
io.clawhub.legendarylibr/seiso
Unified media generation gateway for agents. Discover tools dynamically, choose API key or x402 auth, invoke image/video/audio/music/3D/training tools, and h...
v2.2.1
io.clawhub.a1024708231/video-producer
短视频一键生成技能 v2.2。调用video-director进行画面规划,然后生成AI素材、TTS配音、视频渲染,输出完整MP4。
v0.2.4
io.clawhub.fossilizedcarlos/krea-api
Generate images via Krea.ai API (Flux, Imagen, Ideogram, Seedream, etc.)
v0.1.0
io.clawhub.agmmnn/fal-ai
Generate images, videos, and audio via fal.ai API (FLUX, SDXL, Whisper, etc.)
v1.0.1
io.clawhub.ivangdavila/podcast
Create and grow podcasts by planning episodes, producing audio or video, generating clips, and building audience across formats.
v1.3.15
io.clawhub.dlazyai/dlazy-jimeng-omnihuman-1-5
Generate realistic digital human broadcast videos from portrait images and audio/text using Jimeng OmniHuman 1.5. 使用即梦 (Jimeng) OmniHuman 1.5 模型,通过人像图片和音频/文本生成逼真的数字人播报视频。
v0.1.0
io.clawhub.delorenj/fal-text-to-image
Generate, remix, and edit images using fal.ai's AI models. Supports text-to-image generation, image-to-image remixing, and targeted inpainting/editing.
v0.1.0
io.clawhub.kalvinrv/lipsync
Lip-sync a face to a specific audio track on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar from a portrai...
v1.1.8
io.clawhub.matusvojtek/tubescribe
YouTube video summarizer with speaker detection, formatted documents, and audio output. Works out of the box with macOS built-in TTS. Optional recommended tools (pandoc, ffmpeg, mlx-audio) enhance quality. Requires internet for YouTube access. No paid APIs or subscriptions. Use when user sends a YouTube URL or asks to summarize/transcribe a YouTube video.