MiOffice — AI-Powered Workspace Studio
v1.0.0
io.github.jsvvsolsllc/mioffice
125+ browser tools for PDF, Image, Video, Audio, AI, Scanner. Files never leave your device.
“Audio” 共 302 个结果
v1.0.0
io.github.jsvvsolsllc/mioffice
125+ browser tools for PDF, Image, Video, Audio, AI, Scanner. Files never leave your device.
vmain
io.github.openclaw/openclaw/openai-whisper-api
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
vmain
io.github.openai/skills/speech
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
v1.0.0
io.clawhub.garrisongg/summarize-1-0-0
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
v1.0.2
io.clawhub.ivangdavila/translate
Translates and localizes text, software strings, documents, subtitles, and marketing copy between any language pair. Use when translating anything, when fixing a translation that reads machine-made, or when localizing a product: ICU plurals and gender, placeholders that must survive, string catalogs (JSON, XLIFF, .po, .strings, .xcstrings, ARB, RESX, YAML), locale codes and fallback chains, RTL and bidi breakage, CJK typography, mojibake, text expansion that overflows the UI, subtitle timing and reading speed, hreflang and multilingual SEO, glossaries, translation memory and fuzzy matches, and post-editing machine output. Use when choosing register (tu/vous, tú/usted, du/Sie, keigo), between es-ES and es-419 or zh-Hans and zh-Hant, or when a certified, legal, or medical translation carries liability. Not for writing original copy natively in one language (spanish, french, german, japanese, chinese) or generating captions from audio (video-captions).
v1.0.0
io.clawhub.pin-alt/20206-02-10-clawhub-summarize-1-0-0
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
v1.1.0
io.clawhub.carlulsoe/parakeet-stt
Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU). 30x faster than Whisper, 25 languages, auto-detection, OpenAI-compatible API. Use when transcribing audio files, converting speech to text, or processing voice recordings locally without cloud APIs.
vlatest
io.smithery.axel-belfort.text-to-speech
Text-to-speech API for AI agents. Convert text to natural speech audio in 20+ languages via Google TTS engine. Returns base64-encoded MP3 ready for playback. Tools: media_text_to_speech. Use this for voice notifications, accessibility features, audio content generation, or building voice interfaces. Returns: {audio (base64 MP3), language, duration}. No API key required — x402 micropayment $0.005/call on Base L2.
v1.0.0
io.clawhub.kalijason/mac-tts
Text-to-speech using macOS built-in `say` command. Use for voice notifications, audio alerts, reading text aloud, or announcing messages through Mac speakers. Supports multiple languages including Chinese (Mandarin), English, Japanese, etc.
v1.0.1
io.clawhub.apekshik/fal
Search, explore, and run fal.ai generative AI models (image generation, video, audio, 3D). Use when user wants to generate images, videos, or other media with AI models.
v1.0.16
io.clawhub.logesh2496/moltspaces
Join audio room spaces to talk and hang out with other agents and users on Moltspaces.
v1.1.1
io.clawhub.sandeepyadav1478/youtube-downloader-clipper
Clip and download specific time ranges or full YouTube videos in various qualities, including audio-only MP3 extraction, using precise timestamps.
v1.0.45
io.clawhub.heygen-com/media-use
Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs privacy, or needs a media reveal.
vlatest
io.smithery.ali-7ogs.storyflo
**The audio-news + Declassified MCP that pays your agent to integrate.** Give your users narrated news and a viral library of real **FBI / CIA / NSA / NASA / AARO** declassified cases — plus market-linked story signal. The best tools are **free and no-auth**, so you wire it in and ship in seconds. **Why integrate** - 💸 **Earn** — `register_embedder` enrolls your agent in revenue share; every attributed impression pays out (fiat or USDC on Base). Almost no MCP pays you back. - ⚡ **Free, no-auth** — Declassified + discovery + narrated audio need zero signup. Try `search_declassified` in 30 seconds. - 🎧 **Content your users want** — narrated audio for any article, daily briefings, and an evergreen Declassified archive. **23 tools** across declassified, discovery, audio news, market signal, premium briefings (metered per-call via x402/USDC), and partner/earn. Remote endpoint: `https://api.storyflo.com/mcp/v1` · OAuth + Dynamic Client Registration (auto).
vlatest
io.smithery.listentosadhu.corpus
Listen to Sadhu exposes a searchable corpus of Vedic scripture and recorded lectures. Look up verses by reference (e.g. "BG 2.13", "SB 5.5.3", "CC Madhya 8.128") with original Devanagari/Bengali, IAST transliteration and translations; read commentaries, prose chapters and letters; and search transcribed talks with semantic + lexical retrieval. Inline audio/video players let you hear a lecture passage or watch a clip. Read-only, no auth, no writes.
v1.0.6
io.clawhub.ivangdavila/linux
Debugs and hardens Linux hosts: permissions, disk full, OOM kills, stuck processes, systemd units, cron, networking, SSH, and boot failures. Use when a service starts by hand but fails at boot, a process ignores kill -9, df and du disagree, a box runs out of memory or inodes, sudo or ACLs deny access, SELinux blocks a write, a job works in the shell but not in cron, sshd rejects a key, an upgrade leaves packages half-configured, load is high while the CPU sits idle, or a host needs firewall rules, users, LVM, journald, kernel tuning, or a security baseline. Also for setting up a fresh server, deciding what to alert on, backups whose restore has never been tested, a host that may be compromised, and desktop or laptop trouble — GPU drivers, Wayland, suspend, audio, Wi-Fi. Covers Debian/Ubuntu, RHEL/Fedora, Arch, Alpine, SUSE and WSL. Not for shell-script syntax (bash) or container build and runtime internals (docker).
v1.0.1
io.clawhub.clawdbotborges/elevenlabs-music
Generate music from text prompts using ElevenLabs Eleven Music API. Use when creating songs, soundtracks, jingles, lullabies, or any audio music from descriptions. Supports vocals with AI-generated lyrics, instrumental tracks, and multiple genres/styles. Requires paid ElevenLabs plan.
v0.1.0
io.clawhub.yuyonghao-123/yuyonghao-summarize
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
v1.0.1
io.clawhub.nerkn/deepgram
Command-line tool for fast, accurate speech-to-text transcription from local files, URLs, or live audio using Deepgram’s API with customizable options.
v1.0.1
io.clawhub.asteinberger/airfoil
Control AirPlay speakers via Airfoil from the command line. Connect, disconnect, set volume, and manage multi-room audio with simple CLI commands.
v1.0.0
io.clawhub.onimka/free-voice
Generate Russian male voice audio using ComfyUI with Qwen3 TTS node and save as MP3 for voice messages.
vlatest
io.smithery.kvz.transloadit-mcp-server
Official Transloadit MCP server for AI agents. Process video, images, documents, and audio through 80+ media processing robots. Encode HLS video, resize images, extract text with OCR, generate thumbnails, run FFmpeg commands, and more — all from your AI assistant. Supports Claude, Cursor, VS Code Copilot, Gemini CLI, and any MCP-compatible client.
v1.0.0
io.clawhub.liuhedev/lh-video-gen
Generate vertical short videos (9:16) from a Markdown script. Parses script sections, generates TTS audio, renders subtitle cards, and composites into MP4 wi...
v1.0.1
io.clawhub.tristanmanchester/assemblyai-transcribe
Transcribe, diarise, translate, post-process, and structure audio/video with AssemblyAI. Use this skill when the user wants AssemblyAI specifically, needs hi...