Files
rustytorch/crates/models/rtx-csm/docs/clawsample_integration_plan.md
T
osobhandClaude Opus 4.7 83e80dc721 rtx-csm: add deferred clawsample-csm integration plan
Detailed plan for exposing rtx-csm as a managed service on the
clawsample platform: crate layout (mirrors clawsample-demucs +
clawsample-gen), HTTP routes (/v1/tts, /v1/tts/async, /v1/voice_profile,
WS /v1/converse), DB schema, R2 paths, webhook dispatch, ordered TDD
task breakdown, open decisions, acceptance criteria, reference commits.

Estimated effort: 5-10 days. Trigger to start: a real consumer for the
public TTS API, OR standalone converse_server hits a hard ceiling.
Until then, integration is product/platform work, not ML work.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-04-27 04:50:43 -07:00

13 KiB
Raw Blame History

clawsample-csm: integration plan

Status: planned, not started. Owner: deferred — pick up when there's a real consumer for the public TTS API. Estimated effort: 510 working days (TDD pace, no mocks per clawsample rules).

1. Goal

Expose rtx-csm (the local Rust voice stack) as a managed service on the clawsample platform, mirroring the DemucsService / GenService pattern. The crate is already deployable as a single-process binary via examples/converse_server. Integration adds: multi-tenancy, persistent async-job semantics, R2-backed audio storage, webhook dispatch, per-user voice profiles, billing/quota hooks.

2. Non-goals

  • New ML capabilities. Voice clone, watermark, streaming TTS, barge-in, Q8, etc. all already exist in rtx-csm — integration just routes them.
  • Lower latency. Platform overhead (DB writes, R2 upload, webhook dispatch) adds latency vs the standalone converse_server. The trade-off is durability + multi-tenancy.
  • Replacing examples/converse_server. The standalone binary stays as the canonical reference + dev tool.

3. Architecture

3.1 Crate layout

New crate at clawsample/backend/crates/clawsample-csm/. Mirrors clawsample-demucs / clawsample-gen exactly:

clawsample-csm/
├── Cargo.toml
└── src/
    └── lib.rs   (target ≤ 250 LOC per the 1250-line limit; split if needed)

Cargo.toml deps:

rtx-csm = { path = "../../../../rustytorch/crates/models/rtx-csm" }
clawsample-audio = { workspace = true }   # WAV encode/decode helpers
clawsample-r2    = { workspace = true }   # R2 uploads
bytes      = { workspace = true }
hound      = { workspace = true }
tokio      = { workspace = true }
thiserror  = { workspace = true }
tracing    = { workspace = true }
serde      = { workspace = true }
serde_json = { workspace = true }
uuid       = { workspace = true }

3.2 Service surface (lib.rs)

Mirror of DemucsService:

pub struct CsmServiceConfig {
    pub model_path: Option<PathBuf>,        // None = HF cache (default)
    pub quantized_gguf: Option<PathBuf>,    // Q8 GGUF for voice loops
    pub watermark_secret: Option<u16>,      // 16-bit message base
    pub default_speaker: u32,
    pub default_max_audio_ms: u32,
    pub r2_bucket: String,                  // where outputs land
}

pub struct CsmService {
    generator: Arc<Mutex<rtx_csm::Generator>>,
    config: CsmServiceConfig,
}

pub struct TtsResult {
    pub job_id: String,
    pub audio_url: String,             // pre-signed R2 URL
    pub duration_ms: u32,
    pub watermark_message: Option<u16>,
    pub tts_latency_ms: u128,
}

impl CsmService {
    pub fn new(config: CsmServiceConfig) -> Result<Self, CsmServiceError>;
    pub async fn synthesize(
        &self,
        text: &str,
        speaker: u32,
        voice_profile: Option<&VoiceProfile>,
    ) -> Result<TtsResult, CsmServiceError>;
    pub async fn voice_clone(
        &self,
        ref_audio: &[u8],   // user-uploaded WAV/FLAC bytes
        ref_text: &str,
    ) -> Result<VoiceProfile, CsmServiceError>;
}

pub struct VoiceProfile {
    pub id: String,                   // owns its R2 key for the LoRA adapter
    pub adapter_url: String,
    pub created_at: chrono::DateTime<Utc>,
}

pub enum CsmServiceError { /* ... */ }

Generator runs inside Arc<Mutex<>>. Inference dispatches via tokio::task::spawn_blocking — same pattern as DemucsService::separate which is already proven in clawsample-demucs/src/lib.rs:81.

3.3 HTTP routes (clawsample-api crate)

New file clawsample-api/src/routes/tts.rs:

Method Path Behavior
POST /v1/tts Sync. JSON {text, speaker?, voice_id?, max_audio_ms?} → 200 with {job_id, audio_url, duration_ms}. Caps at small clips (e.g. ≤ 30 s). Returns 400 on overflow.
POST /v1/tts/async Async. Same JSON → 202 {job_id}. Dispatches to background worker. Webhook fires on completion.
WS /v1/converse Full conversation. Auth required. Per-user history persisted in Postgres. Honors --stream-tts semantics from the standalone server.
POST /v1/voice_profile Multipart upload of ref_audio (WAV/FLAC) + ref_text (form field). Trains a small LoRA adapter, stores in R2, returns voice_id.
GET /v1/voice_profile/:id Returns metadata + a download URL for the user's own voice profiles.
DELETE /v1/voice_profile/:id Removes from R2 + DB.
GET /v1/voice_profile Lists current user's voices.

All routes go through clawsample's existing auth middleware (API key in Authorization: Bearer ..., scoped to a user/account row in the DB).

3.4 Database

Two new tables:

CREATE TABLE tts_jobs (
    id           UUID PRIMARY KEY,
    user_id      UUID NOT NULL REFERENCES users(id),
    voice_id     UUID REFERENCES voice_profiles(id),
    text         TEXT NOT NULL,
    status       VARCHAR(16) NOT NULL,  -- pending | running | done | failed
    audio_url    TEXT,
    duration_ms  INTEGER,
    error        TEXT,
    created_at   TIMESTAMPTZ NOT NULL DEFAULT now(),
    completed_at TIMESTAMPTZ
);

CREATE TABLE voice_profiles (
    id           UUID PRIMARY KEY,
    user_id      UUID NOT NULL REFERENCES users(id),
    name         VARCHAR(128),
    adapter_url  TEXT NOT NULL,        -- R2 key for safetensors
    created_at   TIMESTAMPTZ NOT NULL DEFAULT now()
);

Run as proper migrations under clawsample-db/migrations/.

3.5 R2 integration

  • TTS output WAVs → r2://{bucket}/tts/{user_id}/{job_id}.wav
  • LoRA adapters → r2://{bucket}/voice/{user_id}/{voice_id}.safetensors
  • Pre-signed URLs (existing helper in clawsample-r2) for downloads, 1-hour TTL.

3.6 Webhook dispatch

For /v1/tts/async: on job completion, fire webhook to the user's configured endpoint with {job_id, audio_url, duration_ms, status}. Reuse the existing webhook dispatcher (Phase 1 ships this; check clawsample-api for dispatch_webhook or similar).

4. Detailed task breakdown

Order matters: each step must have a real DB and real ML model running locally per clawsample's no-mocks rule.

Step 1 — crate skeleton (~0.5 day)

  • Create clawsample-csm/ with Cargo.toml + src/lib.rs stub
  • Add to workspace Cargo.toml members
  • cargo build -p clawsample-csm green

Step 2 — CsmService + sync synthesize() (~1.5 days, TDD)

  • Tests first: load real CSM-1B (HF cache must exist), synthesize a 1-sec clip, assert WAV bytes are valid + duration matches expectations.
  • CsmService::new boots Generator + applies optional Q8 GGUF + optional watermarker.
  • synthesize runs in spawn_blocking; returns Vec<f32> PCM.
  • Add audio_to_wav_bytes via clawsample-audio helper.

Step 3 — R2 upload glue (~0.5 day)

  • Test: synthesize → upload → fetch back via pre-signed URL → compare WAV bytes byte-for-byte.
  • Helper in clawsample-csm calls clawsample-r2::upload.

Step 4 — POST /v1/tts route (~1 day, TDD)

  • Test in clawsample-api/src/routes/tts.rs::tests:
    • 401 without auth
    • 400 on empty text
    • 200 returns valid audio_url that fetches a real WAV
    • text > limit → 400
  • Route handler enforces quota + writes a tts_jobs row.

Step 5 — POST /v1/tts/async + webhook (~1 day)

  • Test: submit async job → poll status → assert webhook fires (use the existing test webhook listener pattern).
  • Background worker pattern: tokio task pool reading from a job queue table or in-memory channel — match what clawsample-demucs does.

Step 6 — voice profile training (~2 days)

  • This is the trickiest: training a LoRA adapter takes minutes on Metal. Likely an async-only flow.
  • Tests: upload 30 sec ref audio + transcript → poll until done → assert generated audio with the new voice differs from default voice (via WavLM-SV cosine).
  • Reuse examples/lora_train.rs logic — extract the core into a public rtx_csm::training::train_voice_clone if not already exposed.

Step 7 — POST /v1/voice_profile route (~1 day)

  • Multipart upload handler (axum Multipart).
  • Schedule training as a long-running async job.
  • Webhook on completion.

Step 8 — WS /v1/converse route (~1 day)

  • Port the relevant code from examples/converse_server into the route handler, with platform auth on connect.
  • Persist conversation history in Postgres per user/session.

Step 9 — production hardening + docs (~12 days)

  • README in clawsample-csm/ documenting flags + perf numbers
  • Operational doc (Grafana metrics to watch, alert thresholds)
  • Load test against a real Z.AI account for end-to-end voice loop

5. Open decisions (resolve before starting)

  1. Synchronous TTS endpoint allowed? /v1/tts (sync) holds an HTTP connection for ~520 s. Long requests don't fit clawsample's architecture cleanly. Probably should be async-only, with sync as a shortcut for ≤ 5-second outputs.

  2. Voice training is async by necessity (minutes). What's the SLA? Probably "best-effort, webhook on completion, no time guarantee" — match clawsample-gen if it has long-running jobs.

  3. Per-user concurrency limit. Single-process Generator under Mutex serializes all inference. Either (a) run multiple worker processes with a load balancer, (b) use multiple Generator instances in a pool, or (c) hold the mutex per-job and accept queueing. Match the pattern clawsample-demucs already uses.

  4. Watermark policy for voice loops. Streaming TTS skips watermark (per --stream-tts semantics). For multi-tenant abuse traceability, we may need to either (a) require non-streaming for public APIs, (b) hash the request ID into the watermark message, or (c) ship raw streaming and rely on logging/account suspension.

  5. GPU placement. Production deploy needs Metal (Apple) or CUDA (Linux). Decide deployment target before implementation; affects feature flags and CI.

  6. CUDA driver issue (memory note project_rtx_csm.md): main rustytorch host needs 590.x driver. Resolve before deploying to that machine.

6. Test strategy (per clawsample TDD rule)

  • No mocks. Tests use real CSM-1B + real Postgres + real R2 (or MinIO local). Match the existing clawsample-demucs test pattern.
  • Test fixture audio: pre-generate a known-good 5-sec WAV and a 10-sec reference audio for voice cloning tests. Check into clawsample-csm/tests/fixtures/.
  • Smoke test: cargo test -p clawsample-csm runs every public method against real deps; CI must run on a host with HF cache + real Postgres.
  • Integration tests in clawsample-api/src/routes/tts.rs::tests drive the full HTTP surface.

7. Acceptance criteria

The integration is "done" when all of these pass:

  1. cargo build -p clawsample-csm --release green.
  2. cargo test -p clawsample-csm green (real deps, no mocks).
  3. cargo test -p clawsample-api routes::tts green.
  4. End-to-end: curl POST /v1/tts -H 'Authorization: Bearer <key>' -d '{"text":"Hello"}' returns a JSON with audio_url that GET-s to a valid 24 kHz mono WAV.
  5. Async + webhook: same with /v1/tts/async, webhook arrives at the configured endpoint within 30 s.
  6. Voice clone: upload 30 sec reference + transcript → wait for webhook → use the new voice_id in /v1/tts → output's WavLM-SV cosine to the reference > 0.6 (matches the LoRA-trained pattern from examples/lora_train).
  7. WebSocket /v1/converse works against an authenticated client + a real LLM (Z.AI thinking-disabled or equivalent).
  8. README documents production deploy: which flags, expected latency (~2 s TTFA per the rtx-csm bench), Grafana dashboards, runbook.

8. Reference points (rtx-csm side)

These are the rtx-csm capabilities the integration consumes. All exist already (commits cited as of 2026-04-27):

Capability rtx-csm entry point Reference commit
Q8 GGUF load Generator::load_csm_1b_quantized 8ea3055
Streaming TTS Converse::run_streaming ae400a7
LoRA voice clone add_lora_to_backbone + examples/lora_train 3162f0f
AudioSeal watermark Generator::set_watermarker 3162f0f
Z.AI extra_body GenConfig.extra_body 669319d
Phase tracing recv_phase / llm_to_first_audio / conv_total /metrics gauges 334d933
Boot warm-up done in-server, not in lib abc07ff (in examples/converse_server)
Bench harness examples/converse_server_bench 2ce6f8f

Production-validated voice-AI flag set (from memory project_rtx_csm.md part 42):

--quantized-gguf <gguf>     -- Q8, ~2 GB on disk
--stream-tts                -- sub-second TTFA
--llm-base ...              -- any OpenAI-compatible
--llm-model ...
--llm-extra-body '{"thinking":{"type":"disabled"}}'  -- mandatory for Z.AI
--auth-token <secret>
--rate-audio-secs-per-min 600
--rate-turns-per-min 60

Measured numbers:

  • Client TTFA p50: ~2 s (real Z.AI loop)
  • llm_to_first_audio: ~556 ms server-side (mock LLM, Q8 + stream)
  • Total turn: ~17 s (multi-sentence reply at Z.AI's stream rate)

9. When to actually pick this up

Don't start until at least one of these is true:

  • A real consumer is asking for /v1/tts (internal product or external dev).
  • The standalone converse_server is hitting a hard ceiling (rate limits, multi-tenant abuse, billing complexity).
  • A specific feature (voice library, multi-user history) requires the platform layer.

Until then, the standalone binary serves all the use cases the model itself supports. The integration is product/platform work, not ML work.