text-to-speech
32 MCP servers and Agent skills related to text-to-speech, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Sag
v1.0.0
io.clawhub.steipete/sag
ElevenLabs text-to-speech with mac-style say UX.
Edge TTS
v2.0.0
io.clawhub.i3130002/edge-tts
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
elevenlabs-voices
v2.2.0
io.clawhub.robbyczgw-cla/elevenlabs-voices
High-quality voice synthesis with 18 personas, 32 languages, sound effects, batch processing, and voice design using ElevenLabs API.
OpenAI TTS
v1.0.0
io.clawhub.pors/openai-tts
Text-to-speech via OpenAI Audio Speech API.
Elevenlabs Tts
v2.4.0
io.clawhub.shaharsha/elevenlabs-tts
ElevenLabs TTS - the best ElevenLabs integration for OpenClaw.
Audio Generation
v1.0.17
io.clawhub.cellcog/audio-generation-cellcog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Three voice providers (OpenAI, ElevenLabs, MiniMax). Professional audio production from text prompts.
Mac TTS
v1.0.0
io.clawhub.kalijason/mac-tts
Text-to-speech using macOS built-in `say` command. Use for voice notifications, audio alerts, reading text aloud, or announcing messages through Mac speakers. Supports multiple languages including Chinese (Mandarin), English, Japanese, etc.
Elevenlabs
v1.3.4
io.clawhub.odrobnik/elevenlabs
Text-to-speech, sound effects, music generation, voice management, and quota checks via the ElevenLabs API.
Audio Cog
v1.0.12
io.clawhub.nitishgargiitd/audio-cog
AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue.
Youtube Factory
v1.3.0
io.clawhub.mayank8290/youtube-factory
Generate complete YouTube videos from a single prompt - script, voiceover, stock footage, captions, thumbnail.
macOS Local Voice
v1.0.0
io.clawhub.strrl/macos-local-voice
Local STT and TTS on macOS using native Apple capabilities. Speech-to-text via yap (Apple Speech.framework), text-to-speech via say + ffmpeg. Fully offline, no API keys required. Includes voice quality detection and smart voice selection.
Tts
v1.0.0
io.clawhub.amstko/tts
Convert text to speech using Hume AI (or OpenAI) API. Use when the user asks for an audio message, a voice reply, or to hear something "of vive voix".
Voice Reply
v1.0.0
io.clawhub.stolot0mt0m/voice-reply
Local text-to-speech using Piper voices via sherpa-onnx. 100% offline, no API keys required. Use when user asks for a voice reply, audio response, spoken answer, or wants to hear something read aloud. Supports multiple languages including German (thorsten) and English (ryan) voices. Outputs Telegram-compatible voice notes with [[audio_as_voice]] tag.
Sherpa ONNX TTS
v0.1.0
io.clawhub.danielsinewe/sherpa-onnx-tts
Local text-to-speech via sherpa-onnx (offline, no cloud)
Voice
v1.0.1
io.clawhub.zhaov1976/voice
Convert text to speech using Microsoft Edge's TTS engine with customizable voices, direct playback, and automatic temporary file cleanup.
Aliyun TTS
v1.0.0
io.clawhub.guang384/aliyun-tts
Alibaba Cloud Text-to-Speech synthesis service.
Voice Message
v1.0.4
io.clawhub.xmanrui/voice-message
Send voice messages across chat channels (Telegram, Discord, Feishu/Lark, Signal, WhatsApp.
AudioPod
v1.2.3
io.clawhub.rakesh1002/audiopod
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.
Qwen3-tts
v1.0.0
io.clawhub.paki81/qwen-tts
Local text-to-speech using Qwen3-TTS-12Hz-1.7B-CustomVoice. Use when generating audio from text, creating voice messages, or when TTS is requested. Supports 10 languages including Italian, 9 premium speaker voices, and instruction-based voice control (emotion, tone, style). Alternative to cloud-based TTS services like ElevenLabs. Runs entirely offline after initial model download.
TTS WhatsApp
v1.0.0
io.clawhub.hopyky/tts-whatsapp
Send high-quality text-to-speech voice messages on WhatsApp in 40+ languages with automatic delivery
Sapi Tts
v1.1.0
io.clawhub.korddie/sapi-tts
Windows SAPI5 text-to-speech with Neural voices. Lightweight alternative to GPU-heavy TTS - zero GPU usage, instant generation. Auto-detects best available voice for your language. Works on Windows 10/11.
openai-tts-python
v1.0.1
io.clawhub.merend/openai-tts-python
Text-to-speech conversion using OpenAI's TTS API for generating high-quality, natural-sounding audio. Supports 6 voices (alloy, echo, fable, onyx, nova, shimmer), speed control (0.25x-4.0x), HD quality model, multiple output formats (mp3, opus, aac, flac), and automatic text chunking for long content (4096 char limit per request). Use when: (1) User requests audio/voice output with triggers like "read this to me", "convert to audio", "generate speech", "text to speech", "tts", "narrate", "speak", or when keywords "openai tts", "voice", "podcast" appear. (2) Content needs to be spoken rather than read (multitasking, accessibility). (3) User wants specific voice preferences like "alloy", "echo", "fable", "onyx", "nova", "shimmer" or speed adjustments.
🗣️ Edge-TTS Skill using uvx
v1.0.0
io.clawhub.al-one/edge-tts-uvx
Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.
Video Producer
v2.2.1
io.clawhub.a1024708231/video-producer
短视频一键生成技能 v2.2。调用video-director进行画面规划,然后生成AI素材、TTS配音、视频渲染,输出完整MP4。
语音合成 Gemini 2.5 TTS
v1.3.22
io.clawhub.dlazyai/dlazy-gemini-2-5-tts
Generate multilingual, highly natural audio using Gemini 2.5 text-to-speech. 使用 Gemini 2.5 强大的文本转语音能力,生成多语言、高自然度的音频。
ElevenLabs
v1.2.6
io.clawhub.byungkyu/elevenlabs-api
ElevenLabs API integration with managed authentication. AI-powered text-to-speech, voice cloning, sound effects, and audio processing. Use this skill when users want to generate speech from text, clone voices, create sound effects, or process audio. For other third party apps, use the api-gateway skill (https://clawhub.ai/byungkyu/api-gateway). Calls run through the `maton` CLI with OAuth login, or over raw HTTP with a Maton API key where the CLI cannot be installed. Every call is authenticated as the user's connection and reaches only what that connection's authorization allows, which the provider enforces on every request; the endpoints documented here are the ones this skill uses, and any other endpoint of this app needs the user to ask for it by name. Default to read and list calls, and confirm every write or new connection with the user. This file also documents the three constructs that turn an ElevenLabs connection into automation, in the order they are used: the connection (the first step), a hosted f
Voice.ai Voices
v1.1.5
io.clawhub.gizmogremlin/voice-ai-voices
High-quality voice synthesis with 9 personas, 11 languages, and streaming using Voice.ai API.
Podcast Generation with Microsoft Foundry
v0.1.0
io.clawhub.thegovind/podcast-generation
Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Use when building text-to-speech features, audio narrative generation, podcast creation from content, or integrating with Azure OpenAI Realtime API for real audio output. Covers full-stack implementation from React frontend to Python FastAPI backend with WebSocket streaming.
it will help you to send voice messages to your AI Assistant and also can make it talk
v1.0.0
io.clawhub.amreahmed/elevenlabs-voice
Text-to-Speech and Speech-to-Text using ElevenLabs AI. Use when the user wants to convert text to speech, transcribe voice messages, or work with voice in multiple languages. Supports high-quality AI voices and accurate transcription.
Inworld TTS
v1.0.0
io.clawhub.gugic/inworld-tts
Text-to-speech via Inworld.ai API. Use when generating voice audio from text, creating spoken responses, or converting text to MP3/audio files. Supports multiple voices, speaking rates, and streaming for long text.
Pocket Tts
v1.0.1
io.clawhub.sherajdev/pocket-tts
Generate high-quality English speech offline on CPU using 8 built-in voices or custom voice cloning with Kyutai's Pocket TTS model.
Text To Speech
v0.1.5
io.clawhub.okaris/text-to-speech
Convert text to natural speech with DIA TTS, Kokoro, Chatterbox, and more via inference.sh CLI.