Commit Graph
1 Commits
Author SHA1 Message Date
osobhandClaude Opus 4.7 2feb5c9d67 rtx-csm: Phase 13.2 — audio→manifest pipeline (no Python sidecar)
Single command turns raw audio (podcast/audiobook/conversation) into a
training-ready JSONL manifest + per-segment clip wavs.

audio_to_manifest:
  1. Diarize (Phase 13.1: Silero V5 + WavLM-SV + clustering)
  2. Per segment: slice audio + Moonshine encode/decode → transcript
  3. Write `<stem>.spk{N}.{idx:04}.wav` + manifest.jsonl

Manifest rows match ManifestRow exactly (Phase 12.3), so the output
flows directly into lora_train_emotional / load_from_manifest.

Knobs: --segments-json (reuse precomputed diar), --emotion-tag and
--stage applied uniformly, --min-transcript-chars filters ASR failures,
plus all Phase 13.1 diarization knobs.

DiarizedSegment gained serde::Deserialize for the segments-json
reuse path.

Verified end-to-end: 2-speaker concat → 2 segments diarized in 192 ms
→ Moonshine transcribed → 2 manifest rows + 2 clips written → round-
trips through lora_train_emotional cleanly (LoRA injected with extended
coverage, adapter saved with embedded metadata, lib suite 104/104).

The full no-Python data-prep loop now reads:
  audio_to_manifest raw.wav → manifest.jsonl
  lora_train_emotional manifest.jsonl → voice.safetensors
  lora_eval base + lora for A/B
  generate / converse_server with --lora voice.safetensors

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-27 18:09:32 -07:00