text-to-speech
Converts text content into natural, fluent speech output, suitable for various multimedia applications, enhancing user audio experience and content accessibility.
npx skills add inferen-sh/skills --skill text-to-speechBefore / After Comparison
1 组Recording narration or voice content requires professional voice actors and recording equipment, which is costly. Post-production modifications are inconvenient; any script change necessitates re-recording.
Using Text-to-Speech (TTS) technology, text can be quickly converted into natural and fluent speech, supporting various timbres, speeds, and language options. This not only significantly reduces production costs and time but also makes modifying voice content exceptionally convenient and fast.
description SKILL.md
text-to-speech
Text-to-Speech
Convert text to natural speech via inference.sh CLI.
Quick Start
Requires inference.sh CLI (infsh). Install instructions
infsh login
# Generate speech
infsh app run infsh/kokoro-tts --input '{"text": "Hello, welcome to our product demo."}'
Available Models
Model App ID Best For
ElevenLabs TTS
elevenlabs/tts
Premium quality, 22+ voices, 32 languages
DIA TTS
infsh/dia-tts
Conversational, expressive
Kokoro TTS
infsh/kokoro-tts
Fast, natural
Chatterbox
infsh/chatterbox
General purpose
Higgs Audio
infsh/higgs-audio
Emotional control
VibeVoice
infsh/vibevoice
Podcasts, long-form
Browse All Audio Apps
infsh app list --category audio
Examples
Basic Text-to-Speech
infsh app run infsh/kokoro-tts --input '{"text": "Welcome to our tutorial."}'
Conversational TTS with DIA
infsh app sample infsh/dia-tts --save input.json
# Edit input.json:
# {
# "text": "Hey! How are you doing today? I'm really excited to share this with you.",
# "voice": "conversational"
# }
infsh app run infsh/dia-tts --input input.json
Long-form Audio (Podcasts)
infsh app sample infsh/vibevoice --save input.json
# Edit input.json with your podcast script
infsh app run infsh/vibevoice --input input.json
Expressive Speech with Higgs
infsh app sample infsh/higgs-audio --save input.json
# {
# "text": "This is absolutely incredible!",
# "emotion": "excited"
# }
infsh app run infsh/higgs-audio --input input.json
Use Cases
-
Voiceovers: Product demos, explainer videos
-
Audiobooks: Convert text to spoken word
-
Podcasts: Generate podcast episodes
-
Accessibility: Make content accessible
-
IVR: Phone system voice prompts
-
Video Narration: Add narration to videos
Combine with Video
Generate speech, then create a talking head video:
# 1. Generate speech
infsh app run infsh/kokoro-tts --input '{"text": "Your script here"}' > speech.json
# 2. Use the audio URL with OmniHuman for avatar video
infsh app run bytedance/omnihuman-1-5 --input '{
"image_url": "https://portrait.jpg",
"audio_url": "<audio-url-from-step-1>"
}'
Related Skills
# ElevenLabs TTS (premium, 22+ voices)
npx skills add inference-sh/skills@elevenlabs-tts
# ElevenLabs dialogue (multi-speaker)
npx skills add inference-sh/skills@elevenlabs-dialogue
# Full platform skill (all 150+ apps)
npx skills add inference-sh/skills@infsh-cli
# AI avatars (combine TTS with talking heads)
npx skills add inference-sh/skills@ai-avatar-video
# AI music generation
npx skills add inference-sh/skills@ai-music-generation
# Speech-to-text (transcription)
npx skills add inference-sh/skills@speech-to-text
# Video generation
npx skills add inference-sh/skills@ai-video-generation
Browse all apps: infsh app list
Documentation
-
Running Apps - How to run apps via CLI
-
Audio Transcription Example - Audio processing workflows
-
Apps Overview - Understanding the app ecosystem
Weekly Installs4.4KRepositoryinferen-sh/skillsGitHub Stars159First Seen6 days agoSecurity AuditsGen Agent Trust HubPassSocketWarnSnykWarnInstalled onclaude-code3.5Kgemini-cli3.1Kcodex3.1Kamp3.1Kkimi-cli3.1Kgithub-copilot3.1K
forumUser Reviews (0)
Write a Review
No reviews yet
Statistics
User Rating
Rate this Skill