speech-to-text
21 MCP servers and Agent skills related to speech-to-text, each with install commands, source and popularity data, ready to paste into Cursor, Claude Code and other clients.
Openai Whisper
v1.0.0
io.clawhub.steipete/openai-whisper
Local speech-to-text with the Whisper CLI (no API key).
Local Whisper
v1.0.0
io.clawhub.araa47/local-whisper
Local speech-to-text using OpenAI Whisper. Runs fully offline after model download. High quality transcription with multiple model sizes.
Faster Whisper
v1.5.1
io.clawhub.theplasmak/faster-whisper
Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription.
macOS Local Voice
v1.0.0
io.clawhub.strrl/macos-local-voice
Local STT and TTS on macOS using native Apple capabilities. Speech-to-text via yap (Apple Speech.framework), text-to-speech via say + ffmpeg. Fully offline, no API keys required. Includes voice quality detection and smart voice selection.
ElevenLabs Speech-to-Text
v1.0.0
io.clawhub.clawdbotborges/elevenlabs-stt
Transcribe audio files using ElevenLabs Speech-to-Text (Scribe v2).
MLX STT
v1.0.7
io.clawhub.guoqiao/mlx-stt
Speech-To-Text with MLX (Apple Silicon) and opensource models (default GLM-ASR-Nano-2512) locally.
Mlx Whisper
v1.0.0
io.clawhub.kevin37li/mlx-whisper
Local speech-to-text with MLX Whisper (Apple Silicon optimized, no API key).
Speech To Text
v0.1.5
io.clawhub.okaris/speech-to-text
Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large.
AudioPod
v1.2.3
io.clawhub.rakesh1002/audiopod
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.
Gemini STT
v1.1.0
io.clawhub.araa47/gemini-stt
Transcribe audio files using Google's Gemini API or Vertex AI
AssemblyAI advanced speech transcription
v1.0.1
io.clawhub.tristanmanchester/assemblyai-transcribe
Transcribe, diarise, translate, post-process, and structure audio/video with AssemblyAI.
Local Whisper
v1.5.0
io.clawhub.impkind/whisper-mlx-local
Free local speech-to-text for Telegram and WhatsApp using MLX Whisper on Apple Silicon. Private, no API costs.
Speech is Cheap Transcribe
v1.2.0
io.clawhub.ilyakam/asr
Fast, affordable automatic speech-to-text transcription supporting 100 languages, speaker diarization, word timestamps, and customizable output formats.
Venice Ai
v2.1.1
io.clawhub.jonisjongithub/venice-ai
Complete Venice AI platform — text generation, vision/image analysis, web search, X/Twitter search, embeddings, TTS, speech-to-text, image generation.
Elevenlabs Transcribe
v1.0.1
io.clawhub.paulasjes/elevenlabs-transcribe
Transcribe audio to text using ElevenLabs Scribe. Supports batch transcription, realtime streaming from URLs, microphone input, and local files.
Addis Assistant
v1.0.0
io.clawhub.dagmawibabi/addis-assistant-stt
Provides Speech-to-Text (STT) and text Translation using the Addis Assistant API (api.addisassistant.com). Use when the user needs to convert an audio file to text (specifically Amharic), or translate text between languages (e.g., Amharic to English). Requires 'x-api-key'.
it will help you to send voice messages to your AI Assistant and also can make it talk
v1.0.0
io.clawhub.amreahmed/elevenlabs-voice
Text-to-Speech and Speech-to-Text using ElevenLabs AI. Use when the user wants to convert text to speech, transcribe voice messages, or work with voice in multiple languages. Supports high-quality AI voices and accurate transcription.
Parakeet Stt
v1.1.0
io.clawhub.carlulsoe/parakeet-stt
Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU). 30x faster than Whisper, 25 languages, auto-detection, OpenAI-compatible API. Use when transcribing audio files, converting speech to text, or processing voice recordings locally without cloud APIs.
Azure Ai Transcription Py
v0.1.0
io.clawhub.thegovind/azure-ai-transcription-py
Azure AI Transcription SDK for Python. Use for real-time and batch speech-to-text transcription with timestamps and diarization. Triggers: "transcription", "speech to text", "Azure AI Transcription", "TranscriptionClient".
Local Whisper (cpp)
v1.0.0
io.clawhub.wuxxin/local-whisper-cpp
Local speech-to-text using whisper-cli (whisper.cpp).
Voice Recognition
v1.0.0
io.clawhub.gykdly/voice-recognition
Local speech-to-text with OpenAI Whisper CLI. Supports Chinese, English, 100+ languages with translation and summarization.