osobh and Claude Opus 4.7
938b54a2b0
rtx-csm: AudioSeal watermark — Rust port end-to-end
...
SEANet generator + detector matching `facebook/audioseal` reference layout
(weight_norm-merged via pure-Rust pickle reader). Verified on real CSM
speech: mean_presence=0.9943, 16/16 message bits decoded.
- src/audioseal.rs: SeanetEncoder (4-stage strided downsample, 2-layer
LSTM bottleneck at 512 channels, 128-dim projection), MsgProcessor
(16-bit message via embedding sum + broadcast-add), SeanetDecoder,
Generator (encoder+msg+decoder), Detector (encoder + single 320×
reverse_convolution + 1×1 head). Padding mirrors audiocraft
_get_extra_padding_for_conv1d exactly.
- src/audioseal_convert.rs: candle_core::pickle reads .pth directly;
merge_weight_norm computes g*v/‖v‖ over all axes except 0; writes
flat safetensors keyed identically to what Generator/Detector read.
- examples/audioseal_inspect.rs: dumps tensor keys + shapes.
- examples/audioseal_convert.rs: HF download + convert CLI.
- examples/audioseal_demo.rs: load + embed + detect on real WAV or
synthetic burst, optionally writes watermarked WAV.
- audio_io.rs gains generic load_mono_at_rate, resample, write_wav_mono
(16 kHz path needed for AudioSeal).
12 new unit tests + 2 converter tests; 63 lib tests total green.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected] >
2026-04-25 19:26:13 -07:00
osobh and Claude Opus 4.7
15dd3575d4
Add rtx-csm: Rust-native port of Sesame CSM-1B with LoRA voice cloning
...
A new model crate at crates/models/rtx-csm implementing end-to-end
inference, quantization, and fine-tuning for Sesame's Conversational
Speech Model (CSM-1B). Built on candle 0.9 + Kyutai Mimi codec.
Key capabilities:
- Inference (FP F16 on Metal, F32 on CPU, BF16 on CUDA)
- Quantized inference (Q8_0 / Q4_K_M GGUF, ~3x speedup, ~50% memory)
- Streaming Mimi decode with proper StreamTensor state machine
- In-context voice cloning via SpeakerProfile
- Classifier-Free Guidance (Koel-TTS recipe)
- Long-form chunked generation with rolling context
- Audio post-processing (HPF + declick + EBU R128 LUFS)
- Text input normalization (brackets, times, unicode, length caps)
- Frame-level repetition guard (loop-escape)
- Top-k + top-p sampling
- LoRA fine-tuning end-to-end (training + inference, on FP and Q8 bases)
- In-process Whisper ASR via whisper-rs (under --features asr)
- Standalone TTS HTTP server (Axum)
- Bench harness with manifest export + per-prompt WER
Phases delivered: quantization, ASR/WER eval, LoRA voice cloning, HTTP
service. AudioSeal/WavLM/Unmute remain as documented future work.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected] >
2026-04-25 18:33:57 -07:00