Files
rustytorch/crates/models/rtx-csm/Cargo.toml
T
osobhandClaude Opus 4.7 df89372ad7 rtx-csm: Phase 10.4 — SilentCipher detect + Watermarker trait + apply CLI
End-to-end SilentCipher: bit-perfect round-trip on real LibriSpeech
audio. Sesame's actual production watermarker now works in pure
candle 0.9 + Metal.

New components in src/silentcipher.rs:

  detect(samples_16k) -> DetectResult
    1. RMS-normalize to VCTK baseline (matches embed pre-conditioning)
    2. STFT -> magnitude
    3. dec_m_0(magnitude) -> (B, message_dim, 1, T) logits
    4. argmax along message_dim -> (T,) per-frame predictions
    5. Truncate to multiple of message_len
    6. Reshape to (n_patches, message_len), per-column mode
    7. Find terminator (value 0), rotate so payload follows it
    8. Subtract +1 offset -> original codes

  encode_bits / decode_bits  (Phase 10.4 fix)
    Switched from base-4 (2 bits per code) to base-`(message_dim - 1)`.
    The 16 kHz model has message_dim=4 = 3 carrier values (1,2,3) +
    terminator (0), NOT 4 carrier values. Original base-4 packing
    occasionally produced value 3, which Python's
    `np.identity(4)[index+1]` would have crashed on. Real capacity:
    15 codes x log2(3) ~= 23.78 bits per patch.

  SilentCipherWatermark (impl Watermarker)
    Wraps a SilentCipherWatermarker with a fixed default_payload so
    it satisfies the existing Watermarker trait. Maps confidence ->
    DetectionResult.mean_presence and the lower-16-bits of the
    decoded payload -> DetectionResult.message (None below confidence
    0.7 to suppress false positives).

  examples/silentcipher_apply
    Mirrors audioseal_apply: --in / --out / --payload / --detect-only.
    Loads from sony/silentcipher HF repo, embeds, optionally
    resamples back to source rate, optionally re-detects to verify.

Verified end-to-end (LibriSpeech /tmp/asr_test.flac, 10.42 s @ 16 kHz):

  Build:        29 ms (3 .ckpt files from HF cache)
  Embed:      1213 ms = 0.116x realtime
  Detect:     1838 ms = 0.18x realtime
  payload:        0x00BC614E (in)
  recovered:      0x00BC614E (out)
  codes match:    15 / 15
  confidence:     1.0000

Clean (un-watermarked) audio: confidence 0.475, codes mostly 0 -
strong signal-vs-noise discrimination at the 0.7 threshold.

This closes the most surprising gap from the Sesame stack analysis:
rtx-csm now has the *literal* Sesame watermarker (not Meta's
AudioSeal) working in pure candle. AudioSeal stays available for
callers that prefer it.

Phase 10.5 (next): wire as a third option in converse_server alongside
AudioSeal, and a 24/16 kHz ResampledWatermarker for the CSM path.
Plus an A/B bench (SilentCipher vs AudioSeal).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-27 12:54:56 -07:00

280 lines
7.8 KiB
TOML

[package]
name = "rtx-csm"
version.workspace = true
edition.workspace = true
authors.workspace = true
license.workspace = true
repository.workspace = true
description = "Rust-native port of Sesame CSM-1B (Conversational Speech Model) on candle + moshi"
# NOTE: candle-transformers 0.8 (workspace pin) does NOT contain the `csm`
# model module — it was added in 0.9.0. We deliberately pull candle 0.9 + moshi
# 0.6 directly here, NOT via workspace deps. Cargo will compile candle 0.8 (for
# the rest of rustytorch) and candle 0.9 (for rtx-csm) side-by-side. No Tensor
# types are shared across that boundary today.
[dependencies]
# Candle 0.9 — required for the csm model module
candle-core = { version = "0.9.1", default-features = false }
candle-nn = { version = "0.9.1", default-features = false }
candle-transformers = { version = "0.9.1", default-features = false }
# Kyutai's moshi crate: provides streaming STT (asr.rs + lm.rs) on top of
# candle 0.9.1. We use moshi::{asr, lm, mimi} for STT integration. Note:
# moshi::mimi uses a different weight-key naming than HF's kyutai/mimi
# (older Kyutai split-format with weight_g/weight_v); we keep our existing
# Mimi loader on candle_transformers::models::mimi for the HF format. The
# STT path uses Kyutai's pytorch_mimi file which IS in moshi's expected
# naming, so they coexist cleanly in different model instances.
moshi = { version = "0.6.4", default-features = false }
# SentencePiece tokenizer for Kyutai STT detokenization (token IDs → text).
sentencepiece = "0.13"
# Mimi neural audio codec: we use the HF-compatible `candle-transformers::models::mimi`
# (not the `moshi` crate, which expects different weight-key naming).
# Tokenizer (Llama-3.2 BPE)
tokenizers = { version = "0.20", default-features = false, features = ["onig"] }
# HF Hub asset resolution (synchronous via ureq + rustls)
hf-hub = { version = "0.5", default-features = false, features = ["ureq", "rustls-tls"] }
# Audio I/O
hound = "3.5"
symphonia = { version = "0.5", features = ["all"] }
rubato = "0.15"
# FFT primitives for the SilentCipher STFT (Phase 10). Pure Rust,
# zero C linkage. Ships in the default build because it's <100 KB
# of compiled code.
rustfft = "6.2"
# Loudness normalization (EBU R128 / ITU-R BS.1770-4)
ebur128 = "0.1"
# In-process ASR via whisper.cpp bindings. Optional via the `asr` feature
# because it pulls a C++ build (cmake + clang). Provides Metal acceleration.
whisper-rs = { version = "0.16", default-features = false, optional = true }
# Silero V5 VAD via ONNX Runtime (`ort` crate). Optional via the `vad`
# feature because it introduces a second ML inference runtime alongside
# candle — Phase 8.1.2's `ort_conflict_probe` binary verifies it doesn't
# regress CSM Metal inference the way whisper-rs (ggml) did.
voice_activity_detector = { version = "0.2", optional = true }
# Text normalization
unicode-normalization = "0.1"
regex = "1"
# Weight loading
safetensors = "0.4"
# Errors / logging / serde
anyhow.workspace = true
thiserror.workspace = true
tracing.workspace = true
serde.workspace = true
serde_json.workspace = true
# Numerics
half = "2.3"
rand = "0.8"
bytemuck = { version = "1.14", features = ["derive"] }
# Async + HTTP for the LlmClient abstraction (Phase 6b). Promoted from
# dev-dependency to regular dependency so the trait is part of the public
# library surface.
tokio = { version = "1", features = ["macros", "rt-multi-thread", "sync"] }
futures-util = "0.3"
reqwest = { version = "0.12", default-features = false, features = ["json", "stream", "rustls-tls"] }
async-trait = "0.1"
eventsource-stream = "0.2"
[dev-dependencies]
clap = { version = "4.5", features = ["derive"] }
tempfile = "3.0"
approx = "0.5"
tracing-subscriber = "0.3"
# For the TTS HTTP server + converse_server WebSocket examples.
axum = { version = "0.7", features = ["multipart", "ws"] }
# WebSocket client for examples/converse_client.
tokio-tungstenite = { version = "0.24", default-features = false, features = ["connect", "rustls-tls-webpki-roots"] }
# tokio with extra features (signal handler) needed by tts_server.
tokio = { version = "1", features = ["macros", "rt-multi-thread", "signal", "sync"] }
tower = "0.5"
tower-http = { version = "0.6", features = ["trace"] }
# Multipart support added on top of the public reqwest dep for tts_server_bench.
reqwest = { version = "0.12", default-features = false, features = ["json", "multipart", "rustls-tls"] }
[features]
default = ["cpu"]
cpu = []
cuda = ["candle-core/cuda", "candle-nn/cuda", "candle-transformers/cuda"]
metal = ["candle-core/metal", "candle-nn/metal", "candle-transformers/metal"]
accelerate = ["candle-core/accelerate", "candle-nn/accelerate"]
mkl = ["candle-core/mkl", "candle-nn/mkl"]
# In-process Whisper ASR via whisper.cpp bindings. Brings in C++ build deps.
asr = ["dep:whisper-rs"]
asr-metal = ["asr", "whisper-rs/metal"]
asr-cuda = ["asr", "whisper-rs/cuda"]
# Silero V5 VAD via ONNX Runtime. Use `--features metal,vad` to enable
# alongside CSM. Phase 8.1.2 must pass before relying on this in
# production.
vad = ["dep:voice_activity_detector"]
[[example]]
name = "generate"
path = "examples/generate.rs"
[[example]]
name = "bench"
path = "examples/bench.rs"
[[example]]
name = "quantize"
path = "examples/quantize.rs"
[[example]]
name = "inspect_gguf"
path = "examples/inspect_gguf.rs"
[[example]]
name = "qmatmul_repro"
path = "examples/qmatmul_repro.rs"
[[example]]
name = "qmm_layer_diff"
path = "examples/qmm_layer_diff.rs"
[[example]]
name = "lora_train_step"
path = "examples/lora_train_step.rs"
[[example]]
name = "forward_loss_demo"
path = "examples/forward_loss_demo.rs"
[[example]]
name = "lora_finetune_step"
path = "examples/lora_finetune_step.rs"
[[example]]
name = "tts_server"
path = "examples/tts_server.rs"
[[example]]
name = "lora_train"
path = "examples/lora_train.rs"
[[example]]
name = "audioseal_inspect"
path = "examples/audioseal_inspect.rs"
[[example]]
name = "audioseal_convert"
path = "examples/audioseal_convert.rs"
[[example]]
name = "audioseal_demo"
path = "examples/audioseal_demo.rs"
[[example]]
name = "audioseal_apply"
path = "examples/audioseal_apply.rs"
[[example]]
name = "wavlm_sv_convert"
path = "examples/wavlm_sv_convert.rs"
[[example]]
name = "wavlm_sv_demo"
path = "examples/wavlm_sv_demo.rs"
[[example]]
name = "wavlm_sv_inspect"
path = "examples/wavlm_sv_inspect.rs"
[[example]]
name = "pipeline"
path = "examples/pipeline.rs"
[[example]]
name = "generate_long"
path = "examples/generate_long.rs"
[[example]]
name = "tts_server_bench"
path = "examples/tts_server_bench.rs"
[[example]]
name = "stt_demo"
path = "examples/stt_demo.rs"
[[example]]
name = "llm_chat"
path = "examples/llm_chat.rs"
[[example]]
name = "converse"
path = "examples/converse.rs"
[[example]]
name = "converse_server"
path = "examples/converse_server.rs"
[[example]]
name = "converse_client"
path = "examples/converse_client.rs"
[[example]]
name = "converse_server_bench"
path = "examples/converse_server_bench.rs"
[[example]]
name = "llm_extra_body_smoke"
path = "examples/llm_extra_body_smoke.rs"
[[example]]
name = "stt_profile"
path = "examples/stt_profile.rs"
[[example]]
name = "whisper_profile"
path = "examples/whisper_profile.rs"
required-features = ["asr"]
[[example]]
name = "ort_conflict_probe"
path = "examples/ort_conflict_probe.rs"
required-features = ["vad"]
[[example]]
name = "make_silence_test"
path = "examples/make_silence_test.rs"
[[example]]
name = "moonshine_inspect"
path = "examples/moonshine_inspect.rs"
[[example]]
name = "moonshine_smoke"
path = "examples/moonshine_smoke.rs"
[[example]]
name = "moonshine_transcribe"
path = "examples/moonshine_transcribe.rs"
[[example]]
name = "moonshine_profile"
path = "examples/moonshine_profile.rs"
[[example]]
name = "silentcipher_inspect"
path = "examples/silentcipher_inspect.rs"
[[example]]
name = "silentcipher_smoke"
path = "examples/silentcipher_smoke.rs"
[[example]]
name = "silentcipher_apply"
path = "examples/silentcipher_apply.rs"