Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

nuva-lab/align-captions

Name: align-captions
Author: nuva-lab

skills/align-captions/SKILL.md

npx skillsauth add nuva-lab/vibecut align-captions

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

align-captions

Align existing script text to audio for karaoke-style captions. Uses Qwen3-ForcedAligner-0.6B (~30ms precision) and jieba for Chinese word segmentation (groups characters into natural words).

Pipeline

Script + Audio
    ↓
Qwen3-ForcedAligner (character-level timestamps)
    ↓
Jieba word segmentation (characters → Chinese words)
    ↓
Position-based phrase matching (words → phrases)
    ↓
Output: phrases with embedded word timestamps

Usage

# Align script to audio (phrase-level output with word timestamps)
python skills/align-captions/align.py voiceover.wav --script "当全世界都在追AI的时候..."

# Save to file
python skills/align-captions/align.py voiceover.wav --script "..." --output captions.json

# Word-level only (no phrase grouping)
python skills/align-captions/align.py voiceover.wav --script "..." --word-level

Output Format

Designed for Remotion karaoke rendering:

{
  "segments": [
    {
      "text": "当全世界...",
      "startMs": 240, "endMs": 2080,
      "words": [
        {"text": "当", "startMs": 240, "endMs": 400},
        {"text": "全世界", "startMs": 400, "endMs": 880}
      ]
    }
  ],
  "word_segments": [...],
  "language": "Chinese",
  "model": "Qwen3-ForcedAligner-0.6B"
}

Programmatic Usage

from align import align_captions

# Get phrases with embedded word timestamps
result = align_captions(
    "voiceover.wav",
    script="当全世界都在追AI的时候...",
    language="Chinese"
)

# Each phrase has a 'words' array for karaoke highlighting
for phrase in result["segments"]:
    print(f"{phrase['text']}: {len(phrase['words'])} words")

Integration with make-video

The make_video.py script automatically uses align-captions when:

A voiceover file exists
A script is provided in project.json
caption_mode is "auto" or "asr"

The output is passed to Remotion's RollingCaption component for karaoke rendering.

Error Recovery

Audio file not found: Verify the path exists before calling. The script will raise FileNotFoundError -- check project.json for the correct voiceover path.
Model download fails: Qwen3-ForcedAligner (~1GB) downloads on first run. If it fails (network/disk), retry or manually download to the HuggingFace cache (~/.cache/huggingface/).
jieba not installed: Run pip install jieba. Without it, Chinese text falls back to character-level timestamps (no word grouping).

Notes

Supports 11 languages: Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish

nuva-lab/align-captions

skills/align-captions/SKILL.md

Generate karaoke-style word-level timestamps by aligning script text to audio using Qwen3-ForcedAligner + jieba for Chinese word segmentation. Use when the user says 'align captions', 'karaoke timestamps', 'word timestamps', 'caption alignment', 'sync text to audio'.

5 stars

content-media

Updated Apr 9, 2026

$ install --global

skillsauth

npx skillsauth add nuva-lab/vibecut align-captions

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 9, 2026, 2:19 AM6.8s2 files scanned

SKILL.md

name:: align-captions
description:: >

align-captions

Align existing script text to audio for karaoke-style captions. Uses Qwen3-ForcedAligner-0.6B (~30ms precision) and jieba for Chinese word segmentation (groups characters into natural words).

Pipeline

Script + Audio
    ↓
Qwen3-ForcedAligner (character-level timestamps)
    ↓
Jieba word segmentation (characters → Chinese words)
    ↓
Position-based phrase matching (words → phrases)
    ↓
Output: phrases with embedded word timestamps

Usage

# Align script to audio (phrase-level output with word timestamps)
python skills/align-captions/align.py voiceover.wav --script "当全世界都在追AI的时候..."

# Save to file
python skills/align-captions/align.py voiceover.wav --script "..." --output captions.json

# Word-level only (no phrase grouping)
python skills/align-captions/align.py voiceover.wav --script "..." --word-level

Output Format

Designed for Remotion karaoke rendering:

{
  "segments": [
    {
      "text": "当全世界...",
      "startMs": 240, "endMs": 2080,
      "words": [
        {"text": "当", "startMs": 240, "endMs": 400},
        {"text": "全世界", "startMs": 400, "endMs": 880}
      ]
    }
  ],
  "word_segments": [...],
  "language": "Chinese",
  "model": "Qwen3-ForcedAligner-0.6B"
}

Programmatic Usage

from align import align_captions

# Get phrases with embedded word timestamps
result = align_captions(
    "voiceover.wav",
    script="当全世界都在追AI的时候...",
    language="Chinese"
)

# Each phrase has a 'words' array for karaoke highlighting
for phrase in result["segments"]:
    print(f"{phrase['text']}: {len(phrase['words'])} words")

Integration with make-video

The make_video.py script automatically uses align-captions when:

A voiceover file exists
A script is provided in project.json
caption_mode is "auto" or "asr"

The output is passed to Remotion's RollingCaption component for karaoke rendering.

Error Recovery

Audio file not found: Verify the path exists before calling. The script will raise FileNotFoundError -- check project.json for the correct voiceover path.
Model download fails: Qwen3-ForcedAligner (~1GB) downloads on first run. If it fails (network/disk), retry or manually download to the HuggingFace cache (~/.cache/huggingface/).
jieba not installed: Run pip install jieba. Without it, Chinese text falls back to character-level timestamps (no word grouping).

Notes

Supports 11 languages: Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish

Related Skills

nuva-lab/write-script

tools

VerifiedTrustedCommunity

Generate voiceover scripts in Joyce's style for video clips

5SKILL.mdUpdated Apr 9, 2026

nuva-lab/write-script

nuva-lab/voice-clone

tools

VerifiedTrustedCommunity

Clone a voice using qwen3-tts and generate speech from text

5SKILL.mdUpdated Apr 9, 2026

nuva-lab/skills/validate-media

development

VerifiedTrustedCommunity

# Validate Media Skill Pre-flight media validation and diagnostics using ffprobe. ## Purpose Check video/audio files for common issues before rendering: - Duration mismatches between video and audio tracks - Missing audio tracks - Codec compatibility - Volume levels - Potential freeze points ## Usage ```bash python skills/validate-media/validate.py <video_file> [--verbose] ``` ## Output JSON report with issues and recommendations: ```json { "file": "video.mp4", "video_duration": 35.1

5SKILL.mdUpdated Apr 9, 2026

nuva-lab/skills/validate-media

nuva-lab/transcribe-clip

tools

VerifiedTrustedCommunity

Transcribe a video clip using Gemini to get timestamped segments for captions

5SKILL.mdUpdated Apr 9, 2026

nuva-lab/transcribe-clip

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/nuva-lab/vibecut.git

# Copy into Claude Code skills folder (global)
cp -r vibecut/skills/align-captions ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

nuva-lab/vibecut

5 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT