Detailed plan for exposing rtx-csm as a managed service on the clawsample platform: crate layout (mirrors clawsample-demucs + clawsample-gen), HTTP routes (/v1/tts, /v1/tts/async, /v1/voice_profile, WS /v1/converse), DB schema, R2 paths, webhook dispatch, ordered TDD task breakdown, open decisions, acceptance criteria, reference commits. Estimated effort: 5-10 days. Trigger to start: a real consumer for the public TTS API, OR standalone converse_server hits a hard ceiling. Until then, integration is product/platform work, not ML work. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
13 KiB
clawsample-csm: integration plan
Status: planned, not started. Owner: deferred — pick up when there's a real consumer for the public TTS API. Estimated effort: 5–10 working days (TDD pace, no mocks per clawsample rules).
1. Goal
Expose rtx-csm (the local Rust voice stack) as a managed service on the
clawsample platform, mirroring the DemucsService / GenService pattern.
The crate is already deployable as a single-process binary via
examples/converse_server. Integration adds: multi-tenancy, persistent
async-job semantics, R2-backed audio storage, webhook dispatch, per-user
voice profiles, billing/quota hooks.
2. Non-goals
- New ML capabilities. Voice clone, watermark, streaming TTS, barge-in,
Q8, etc. all already exist in
rtx-csm— integration just routes them. - Lower latency. Platform overhead (DB writes, R2 upload, webhook
dispatch) adds latency vs the standalone
converse_server. The trade-off is durability + multi-tenancy. - Replacing
examples/converse_server. The standalone binary stays as the canonical reference + dev tool.
3. Architecture
3.1 Crate layout
New crate at clawsample/backend/crates/clawsample-csm/. Mirrors
clawsample-demucs / clawsample-gen exactly:
clawsample-csm/
├── Cargo.toml
└── src/
└── lib.rs (target ≤ 250 LOC per the 1250-line limit; split if needed)
Cargo.toml deps:
rtx-csm = { path = "../../../../rustytorch/crates/models/rtx-csm" }
clawsample-audio = { workspace = true } # WAV encode/decode helpers
clawsample-r2 = { workspace = true } # R2 uploads
bytes = { workspace = true }
hound = { workspace = true }
tokio = { workspace = true }
thiserror = { workspace = true }
tracing = { workspace = true }
serde = { workspace = true }
serde_json = { workspace = true }
uuid = { workspace = true }
3.2 Service surface (lib.rs)
Mirror of DemucsService:
pub struct CsmServiceConfig {
pub model_path: Option<PathBuf>, // None = HF cache (default)
pub quantized_gguf: Option<PathBuf>, // Q8 GGUF for voice loops
pub watermark_secret: Option<u16>, // 16-bit message base
pub default_speaker: u32,
pub default_max_audio_ms: u32,
pub r2_bucket: String, // where outputs land
}
pub struct CsmService {
generator: Arc<Mutex<rtx_csm::Generator>>,
config: CsmServiceConfig,
}
pub struct TtsResult {
pub job_id: String,
pub audio_url: String, // pre-signed R2 URL
pub duration_ms: u32,
pub watermark_message: Option<u16>,
pub tts_latency_ms: u128,
}
impl CsmService {
pub fn new(config: CsmServiceConfig) -> Result<Self, CsmServiceError>;
pub async fn synthesize(
&self,
text: &str,
speaker: u32,
voice_profile: Option<&VoiceProfile>,
) -> Result<TtsResult, CsmServiceError>;
pub async fn voice_clone(
&self,
ref_audio: &[u8], // user-uploaded WAV/FLAC bytes
ref_text: &str,
) -> Result<VoiceProfile, CsmServiceError>;
}
pub struct VoiceProfile {
pub id: String, // owns its R2 key for the LoRA adapter
pub adapter_url: String,
pub created_at: chrono::DateTime<Utc>,
}
pub enum CsmServiceError { /* ... */ }
Generator runs inside Arc<Mutex<>>. Inference dispatches via
tokio::task::spawn_blocking — same pattern as DemucsService::separate
which is already proven in clawsample-demucs/src/lib.rs:81.
3.3 HTTP routes (clawsample-api crate)
New file clawsample-api/src/routes/tts.rs:
| Method | Path | Behavior |
|---|---|---|
POST |
/v1/tts |
Sync. JSON {text, speaker?, voice_id?, max_audio_ms?} → 200 with {job_id, audio_url, duration_ms}. Caps at small clips (e.g. ≤ 30 s). Returns 400 on overflow. |
POST |
/v1/tts/async |
Async. Same JSON → 202 {job_id}. Dispatches to background worker. Webhook fires on completion. |
WS |
/v1/converse |
Full conversation. Auth required. Per-user history persisted in Postgres. Honors --stream-tts semantics from the standalone server. |
POST |
/v1/voice_profile |
Multipart upload of ref_audio (WAV/FLAC) + ref_text (form field). Trains a small LoRA adapter, stores in R2, returns voice_id. |
GET |
/v1/voice_profile/:id |
Returns metadata + a download URL for the user's own voice profiles. |
DELETE |
/v1/voice_profile/:id |
Removes from R2 + DB. |
GET |
/v1/voice_profile |
Lists current user's voices. |
All routes go through clawsample's existing auth middleware (API key in
Authorization: Bearer ..., scoped to a user/account row in the DB).
3.4 Database
Two new tables:
CREATE TABLE tts_jobs (
id UUID PRIMARY KEY,
user_id UUID NOT NULL REFERENCES users(id),
voice_id UUID REFERENCES voice_profiles(id),
text TEXT NOT NULL,
status VARCHAR(16) NOT NULL, -- pending | running | done | failed
audio_url TEXT,
duration_ms INTEGER,
error TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
completed_at TIMESTAMPTZ
);
CREATE TABLE voice_profiles (
id UUID PRIMARY KEY,
user_id UUID NOT NULL REFERENCES users(id),
name VARCHAR(128),
adapter_url TEXT NOT NULL, -- R2 key for safetensors
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
Run as proper migrations under clawsample-db/migrations/.
3.5 R2 integration
- TTS output WAVs →
r2://{bucket}/tts/{user_id}/{job_id}.wav - LoRA adapters →
r2://{bucket}/voice/{user_id}/{voice_id}.safetensors - Pre-signed URLs (existing helper in
clawsample-r2) for downloads, 1-hour TTL.
3.6 Webhook dispatch
For /v1/tts/async: on job completion, fire webhook to the user's
configured endpoint with {job_id, audio_url, duration_ms, status}.
Reuse the existing webhook dispatcher (Phase 1 ships this; check
clawsample-api for dispatch_webhook or similar).
4. Detailed task breakdown
Order matters: each step must have a real DB and real ML model running locally per clawsample's no-mocks rule.
Step 1 — crate skeleton (~0.5 day)
- Create
clawsample-csm/withCargo.toml+src/lib.rsstub - Add to workspace
Cargo.tomlmembers cargo build -p clawsample-csmgreen
Step 2 — CsmService + sync synthesize() (~1.5 days, TDD)
- Tests first: load real CSM-1B (HF cache must exist), synthesize a 1-sec clip, assert WAV bytes are valid + duration matches expectations.
CsmService::newboots Generator + applies optional Q8 GGUF + optional watermarker.synthesizeruns inspawn_blocking; returnsVec<f32>PCM.- Add
audio_to_wav_bytesviaclawsample-audiohelper.
Step 3 — R2 upload glue (~0.5 day)
- Test: synthesize → upload → fetch back via pre-signed URL → compare WAV bytes byte-for-byte.
- Helper in
clawsample-csmcallsclawsample-r2::upload.
Step 4 — POST /v1/tts route (~1 day, TDD)
- Test in
clawsample-api/src/routes/tts.rs::tests:- 401 without auth
- 400 on empty text
- 200 returns valid
audio_urlthat fetches a real WAV - text > limit → 400
- Route handler enforces quota + writes a
tts_jobsrow.
Step 5 — POST /v1/tts/async + webhook (~1 day)
- Test: submit async job → poll status → assert webhook fires (use the existing test webhook listener pattern).
- Background worker pattern: tokio task pool reading from a job queue
table or in-memory channel — match what
clawsample-demucsdoes.
Step 6 — voice profile training (~2 days)
- This is the trickiest: training a LoRA adapter takes minutes on Metal. Likely an async-only flow.
- Tests: upload 30 sec ref audio + transcript → poll until done → assert generated audio with the new voice differs from default voice (via WavLM-SV cosine).
- Reuse
examples/lora_train.rslogic — extract the core into a publicrtx_csm::training::train_voice_cloneif not already exposed.
Step 7 — POST /v1/voice_profile route (~1 day)
- Multipart upload handler (axum
Multipart). - Schedule training as a long-running async job.
- Webhook on completion.
Step 8 — WS /v1/converse route (~1 day)
- Port the relevant code from
examples/converse_serverinto the route handler, with platform auth on connect. - Persist conversation history in Postgres per user/session.
Step 9 — production hardening + docs (~1–2 days)
- README in
clawsample-csm/documenting flags + perf numbers - Operational doc (Grafana metrics to watch, alert thresholds)
- Load test against a real Z.AI account for end-to-end voice loop
5. Open decisions (resolve before starting)
-
Synchronous TTS endpoint allowed?
/v1/tts(sync) holds an HTTP connection for ~5–20 s. Long requests don't fit clawsample's architecture cleanly. Probably should be async-only, with sync as a shortcut for ≤ 5-second outputs. -
Voice training is async by necessity (minutes). What's the SLA? Probably "best-effort, webhook on completion, no time guarantee" — match
clawsample-genif it has long-running jobs. -
Per-user concurrency limit. Single-process Generator under
Mutexserializes all inference. Either (a) run multiple worker processes with a load balancer, (b) use multipleGeneratorinstances in a pool, or (c) hold the mutex per-job and accept queueing. Match the patternclawsample-demucsalready uses. -
Watermark policy for voice loops. Streaming TTS skips watermark (per
--stream-ttssemantics). For multi-tenant abuse traceability, we may need to either (a) require non-streaming for public APIs, (b) hash the request ID into the watermark message, or (c) ship raw streaming and rely on logging/account suspension. -
GPU placement. Production deploy needs Metal (Apple) or CUDA (Linux). Decide deployment target before implementation; affects feature flags and CI.
-
CUDA driver issue (memory note
project_rtx_csm.md): main rustytorch host needs 590.x driver. Resolve before deploying to that machine.
6. Test strategy (per clawsample TDD rule)
- No mocks. Tests use real CSM-1B + real Postgres + real R2 (or
MinIO local). Match the existing
clawsample-demucstest pattern. - Test fixture audio: pre-generate a known-good 5-sec WAV and a
10-sec reference audio for voice cloning tests. Check into
clawsample-csm/tests/fixtures/. - Smoke test:
cargo test -p clawsample-csmruns every public method against real deps; CI must run on a host with HF cache + real Postgres. - Integration tests in
clawsample-api/src/routes/tts.rs::testsdrive the full HTTP surface.
7. Acceptance criteria
The integration is "done" when all of these pass:
cargo build -p clawsample-csm --releasegreen.cargo test -p clawsample-csmgreen (real deps, no mocks).cargo test -p clawsample-api routes::ttsgreen.- End-to-end:
curl POST /v1/tts -H 'Authorization: Bearer <key>' -d '{"text":"Hello"}'returns a JSON withaudio_urlthat GET-s to a valid 24 kHz mono WAV. - Async + webhook: same with
/v1/tts/async, webhook arrives at the configured endpoint within 30 s. - Voice clone: upload 30 sec reference + transcript → wait for webhook
→ use the new
voice_idin/v1/tts→ output's WavLM-SV cosine to the reference > 0.6 (matches the LoRA-trained pattern fromexamples/lora_train). - WebSocket
/v1/converseworks against an authenticated client + a real LLM (Z.AI thinking-disabled or equivalent). - README documents production deploy: which flags, expected latency (~2 s TTFA per the rtx-csm bench), Grafana dashboards, runbook.
8. Reference points (rtx-csm side)
These are the rtx-csm capabilities the integration consumes. All exist already (commits cited as of 2026-04-27):
| Capability | rtx-csm entry point | Reference commit |
|---|---|---|
| Q8 GGUF load | Generator::load_csm_1b_quantized |
8ea3055 |
| Streaming TTS | Converse::run_streaming |
ae400a7 |
| LoRA voice clone | add_lora_to_backbone + examples/lora_train |
3162f0f |
| AudioSeal watermark | Generator::set_watermarker |
3162f0f |
| Z.AI extra_body | GenConfig.extra_body |
669319d |
| Phase tracing | recv_phase / llm_to_first_audio / conv_total /metrics gauges |
334d933 |
| Boot warm-up | done in-server, not in lib | abc07ff (in examples/converse_server) |
| Bench harness | examples/converse_server_bench |
2ce6f8f |
Production-validated voice-AI flag set (from memory project_rtx_csm.md part 42):
--quantized-gguf <gguf> -- Q8, ~2 GB on disk
--stream-tts -- sub-second TTFA
--llm-base ... -- any OpenAI-compatible
--llm-model ...
--llm-extra-body '{"thinking":{"type":"disabled"}}' -- mandatory for Z.AI
--auth-token <secret>
--rate-audio-secs-per-min 600
--rate-turns-per-min 60
Measured numbers:
- Client TTFA p50: ~2 s (real Z.AI loop)
- llm_to_first_audio: ~556 ms server-side (mock LLM, Q8 + stream)
- Total turn: ~17 s (multi-sentence reply at Z.AI's stream rate)
9. When to actually pick this up
Don't start until at least one of these is true:
- A real consumer is asking for
/v1/tts(internal product or external dev). - The standalone
converse_serveris hitting a hard ceiling (rate limits, multi-tenant abuse, billing complexity). - A specific feature (voice library, multi-user history) requires the platform layer.
Until then, the standalone binary serves all the use cases the model itself supports. The integration is product/platform work, not ML work.