tools/audio/elevenlabs-voice-isolator/SKILL.md
ElevenLabs voice isolator - remove background noise and isolate vocals from audio via inference.sh CLI. Capabilities: noise removal, voice extraction, audio cleanup, background removal. Use for: podcast cleanup, interview audio, music vocals, noisy recordings, audio restoration. Triggers: voice isolator, noise removal, background removal, isolate voice, clean audio, remove background noise, audio cleanup, voice extraction, elevenlabs isolator, eleven labs noise, vocal isolation, denoise, audio restoration, voice separation
npx skillsauth add inf-sh/skills elevenlabs-voice-isolatorInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
Security scan pending...
This skill is queued for security scanning. Results will appear when the scan completes.
Install the belt CLI skill:
npx skills add belt-sh/cli
Remove background noise and isolate voices from audio via inference.sh CLI.

Requires inference.sh CLI (
belt). Install instructions
belt login
# Isolate voice from noisy audio
belt app run elevenlabs/voice-isolator --input '{"audio": "https://noisy-recording.mp3"}'
| Format | Max Size | Max Duration | |--------|----------|-------------| | WAV | 500MB | 1 hour | | MP3 | 500MB | 1 hour | | FLAC | 500MB | 1 hour | | OGG | 500MB | 1 hour | | AAC | 500MB | 1 hour |
# Remove background noise from a podcast recording
belt app run elevenlabs/voice-isolator --input '{"audio": "https://noisy-podcast.mp3"}'
# Isolate speaker from café background noise
belt app run elevenlabs/voice-isolator --input '{"audio": "https://cafe-interview.mp3"}'
# Separate vocals from instrumental
belt app run elevenlabs/voice-isolator --input '{"audio": "https://song.mp3"}'
# 1. Isolate voice from noisy recording
belt app run elevenlabs/voice-isolator --input '{
"audio": "https://noisy-meeting.mp3"
}' > cleaned.json
# 2. Transcribe the clean audio
belt app run elevenlabs/stt --input '{
"audio": "<cleaned-audio-url>",
"diarize": true
}'
# 1. Clean up the audio
belt app run elevenlabs/voice-isolator --input '{
"audio": "https://raw-recording.mp3"
}' > cleaned.json
# 2. Transform the voice
belt app run elevenlabs/voice-changer --input '{
"audio": "<cleaned-audio-url>",
"voice": "george"
}'
# 1. Clean the voiceover
belt app run elevenlabs/voice-isolator --input '{
"audio": "https://raw-voiceover.mp3"
}' > cleaned.json
# 2. Merge with video
belt app run infsh/media-merger --input '{
"media": ["video.mp4", "<cleaned-audio-url>"]
}'
# ElevenLabs voice changer (transform voice after cleaning)
npx skills add inference-sh/skills@elevenlabs-voice-changer
# ElevenLabs STT (transcribe clean audio)
npx skills add inference-sh/skills@elevenlabs-stt
# ElevenLabs TTS (generate clean speech from text)
npx skills add inference-sh/skills@elevenlabs-tts
# Full platform skill (all 250+ apps)
npx skills add inference-sh/skills@infsh-cli
Browse all audio apps: belt app store --category audio
data-ai
Generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching. Use when the user wants to create a podcast, talking head video, or multi-speaker conversation video.
tools
Generate videos with ByteDance Seedance 2.0 via inference.sh CLI. Unified model for text-to-video, image-to-video, and reference-to-video with synchronized audio, up to 1080p, 4-15s duration. Pro and Fast variants. Studio variants with private asset library for portrait consistency. Use for: social media videos, music videos, product demos, animated content, AI video with sound. Triggers: seedance, seedance 2, bytedance video, seedance t2v, seedance i2v, seedance r2v, video with audio, seedance 2.0, bytedance seedance, seedance studio
tools
Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative
tools
Generate and edit videos with Alibaba HappyHorse 1.0 models via inference.sh CLI. Models: HappyHorse T2V, I2V, R2V, Video Edit. Capabilities: text-to-video, image-to-video, reference-to-video, video editing with natural language, character preservation, 720P/1080P, up to 15 seconds. Use for: physically realistic video, video editing, character-consistent content, product demos, social media. Triggers: happyhorse, happy horse, alibaba video, happyhorse 1.0, dashscope video, alibaba happyhorse, video editing ai, ai video editor