rtx-csm: gap analysis vs Sesame Labs full voice stack

Researched what Sesame has actually disclosed publicly (vs. marketed)
and compared systematically against rtx-csm's shipped surface.

Key findings:

1. **rtx-csm has shipped a SUPERSET of Sesame's open release.** CSM-1B
   inference + voice cloning + Q8 + production WebSocket server +
   three-backend STT pipeline + per-phase observability — Sesame ships
   inference code only.

2. **Sesame's deployed Maya is a cascaded STT->LLM->TTS pipeline**,
   same architecture pattern as rtx-csm. Their research blog explicitly
   states future work is "fully duplex models" — Maya today isn't
   duplex either. We're structurally equivalent at the pipeline level
   to Kyutai Unmute, Sesame's closest peer.

3. **Marketing latency claims are unverified.** "Sub-200 ms TTFA"
   appears in third-party blogs, not in any Sesame paper. Production
   benchmarks of similar cascaded stacks show 250-300 ms TTFT
   (gpt-realtime, Unmute) — our 280-380 ms TTS-side is competitive.
   Z.AI provider TTFT (~1.3 s) is the dominant cost in our 1.96 s
   end-to-end.

4. **Critical correction**: Sesame ships SilentCipher (their fork of
   Sony's), NOT AudioSeal (which is Meta's). Our Phase 4 AudioSeal
   work is functionally equivalent but isn't the *literal* Sesame
   watermarker. SilentCipher port is ~1-2 days.

5. **What's gated on Sesame**: CSM-3B / CSM-8B variants (trained but
   never released), Maya personality fine-tune dataset, distilled
   wearable variant. The Oct 2025 Series B + smart-glasses pivot
   suggests they're unlikely to release any of these.

Punch-list of remaining gaps captured in the doc with status (Closed/
Partial/Open/N/A) per capability.

Recommended next chunks (prioritized):
  1. SilentCipher port (~1-2 days) — literal Sesame watermarker parity
  2. clawsample-csm integration (~5-10 days) — separate plan exists
  3. Tier 2.1 VoXtream look-ahead (~5-7 days) — diminishing returns
     after Phase 9.2's chunk_frames tuning
  4. Tier 3 Frame-Stacked / VADUSA (training-required, ~3-6 weeks)
  5. Distilled CSM (speculative, wait for product target)

Defensible framing: "rtx-csm is Sesame's open release + voice cloning
+ production HTTP/WS server + Kyutai-style cascaded duplex. Remaining
gap to internal Maya is (a) SilentCipher watermarker, (b) Sesame's
proprietary fine-tune dataset, (c) larger CSM variants Sesame chose
not to release."

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-04-27 12:14:04 -07:00
co-authored by Claude Opus 4.7
parent bc57ebcde1
commit 1760309b39
@@ -0,0 +1,169 @@
# rtx-csm vs. Sesame Labs full voice stack — gap analysis
Research conducted 2026-04-27 to compare what we've shipped against
what Sesame Labs has actually disclosed publicly. The goal: an honest
punch-list of remaining gaps so we can prioritize what's worth porting
next vs. what's gated on Sesame releasing more.
## TL;DR
**rtx-csm has shipped a superset of Sesame's open release** (CSM-1B +
voice cloning + Q8 + production WebSocket server) and a rough peer of
Kyutai Unmute at the pipeline level.
The actual deployed Sesame product ("Maya") is, per their research
blog and absence of duplex paper, a **cascaded STT→LLM→TTS pipeline**
same architecture pattern as rtx-csm. There is no public Sesame paper
describing a duplex production system. Their published "Crossing the
uncanny valley" research blog explicitly states future work is
"fully duplex models."
Three concrete gaps remain to "literal Maya parity":
1. **SilentCipher** is what Sesame actually ships as their watermarker
(`SesameAILabs/silentcipher` on GitHub). rtx-csm has an AudioSeal
scaffold instead — different system, same purpose. **~1-2 days to
port** if literal parity matters.
2. **Maya's personality fine-tune dataset** — never released. Replicate
via our existing LoRA pipeline + a custom corpus. This is product
work, not infrastructure work.
3. **CSM-3B / CSM-8B variants** — Sesame trained them, never released.
The Oct 2025 Series B + smart-glasses pivot suggests they're
unlikely to ever release. Gated on Sesame, not us.
## What Sesame has actually disclosed
Substantiated:
- CSM-1B architecture (Llama-3.2 1B + 100M decoder + Mimi RVQ at 12.5 Hz)
- Apache-2.0 weights at `sesame/csm-1b` on HF
- ~1M hours predominantly-English training data
- 2048-token context (~2 min audio)
- Three new evaluation benchmarks (Homograph Disambiguation,
Pronunciation Continuation Consistency, Contextual Expressivity)
- They watermark via SilentCipher (the `watermarking.py` in the open
repo references it; `SesameAILabs/silentcipher` is the GitHub fork
of Sony's SilentCipher implementation)
- They explicitly admit the open release "can only model the text and
speech content in a conversation — not the structure" (i.e. no
turn-taking, VAD, or duplex in the open weights)
Unsubstantiated marketing:
- "Sub-200 ms TTFA" — appears in third-party blogs, NOT in Sesame's
own paper. Voice-agent benchmarks of cascaded stacks in 2025 show
~250-300 ms TTFT for production-tuned pipelines (OpenAI gpt-realtime,
Kyutai Unmute) and 1-2 s for self-hosted. Treat the 200 ms figure as
aspirational marketing.
- Maya's "interruption without stilted restart" — possible via barge-in
+ smart re-prompting. Not described in any technical doc. We have
acoustic barge-in + fade-out shipped.
## Pipeline gap matrix
| Capability | Status | Detail |
|---|---|---|
| CSM-1B inference | **Closed** | Phase 1, candle 0.9 + Q8 GGUF |
| Voice cloning (LoRA) | **Closed** | Phase 3 — exceeds open release; Sesame doesn't ship a finetune pipeline |
| Mimi streaming codec | **Closed** | `mimi.rs::decode_step` chain |
| Long-form generation | **Closed** | `generate_long` with rolling context |
| Cascaded STT→LLM→TTS pipeline | **Closed** | WebSocket server, three STT backends (Phase 6/8) |
| Semantic VAD / EOT detection | **Closed** | Kyutai head-2 EOT (Phase 6c.3b) |
| Energy VAD silence skip | **Closed** | Phase 8.1.3b |
| Barge-in (acoustic) + fade-out | **Closed** | Phase 6c.3c + 6c.3e |
| Multi-speaker context conditioning | **Closed** | `Segment` / `SpeakerProfile` / `build_prompt` |
| Boot warm-up | **Closed** | Phase 8.2 — closes -89 % TTFA |
| Watermarking — AudioSeal | **Closed** | Phase 4, our scaffold |
| **Watermarking — SilentCipher (Sesame's actual)** | **Open** | ~1-2 days to port `SesameAILabs/silentcipher` |
| TTFA latency (real LLM) | **Partial** | 1.96 s end-to-end on Z.AI; ~280-380 ms TTS-side. The Z.AI provider TTFT is the dominant cost. Closed against any public open-stack benchmark; open against marketing claims |
| Full-duplex (listen-while-speaking) | **Open** | Sesame admitted not in CSM-1B. Only Moshi delivers this open-source. Substantial CSM extension needed |
| CSM-3B / CSM-8B port | **N/A** | Weights not released, likely never will be |
| Maya personality fine-tune | **N/A** | Dataset not released; replicate via our LoRA + custom corpus |
| On-device / wearable variant | **Open** | Q8 GGUF gets us part-way; full distillation (smaller decoder, fewer codebooks) not done |
| Memory across conversations | **N/A** | Maya marketing feature; belongs in LLM layer / clawsample, not the speech stack |
| Production STT backends | **Closed** | 3 backends (Kyutai 1B / Moonshine-tiny / Whisper-tiny) — superset of Sesame open release |
## What we have that Sesame doesn't ship
- **`tracing` instrumentation** with /metrics gauges (recv / stt_post /
llm_to_first_audio / conv_total / total_turn) — Phase 6f.trace
- **Configurable LLM backend** via OpenAI-compatible HTTP — Sesame
doesn't expose a swappable LLM in their open release
- **Auth, rate-limit, graceful shutdown** — production hardening Sesame
doesn't ship
- **In-process Mimi (HF-named)** — rest of CSM stays on candle-0.9
while we resist the moshi crate's incompatible Mimi naming
- **Comprehensive perf history doc** capturing every shipped + rejected
optimization with reproducible benchmarks
## Closest open-source peers (anchor)
- **Kyutai Unmute**: closest analog to the deployed Maya — modular
cascaded with Kyutai STT + any LLM + Kyutai TTS 1.6 B. Our pipeline
is structurally equivalent.
- **Moshi**: the only open full-duplex system. If we ever need true
listen-while-speaking, this is the reference.
- **Orpheus TTS**: TTS-only competitor to CSM (Llama-3 3 B + SNAC).
Different codec; not directly portable.
- **OpenAI gpt-realtime**: closed; ~250-300 ms TTFT.
- **LiveKit Agents**: orchestration framework. Not a model. Could be
the answer for multi-tenant / multi-user voice agents downstream.
## Recommended next chunks (prioritized)
1. **Swap AudioSeal scaffold for a SilentCipher port** (~1-2 days).
This is the *actual* Sesame watermarker. Direct repo-to-port:
- `git clone https://github.com/SesameAILabs/silentcipher`
- Architecture is a small CNN — order of magnitude smaller than
AudioSeal. Likely a 1-day port at this point.
- Output: `--watermark-silentcipher` flag in converse_server,
`examples/silentcipher_apply` CLI parallel to `audioseal_apply`.
2. **clawsample-csm integration** (~5-10 days, separate plan exists at
`docs/clawsample_integration_plan.md`). Multi-tenant API surface,
billing, R2 storage. Product work rather than ML work.
3. **Tier 2.1 VoXtream look-ahead** (~5-7 days, plan exists at
`docs/voxtream_port_notes.md`). Closes the remaining ~270 ms TTS
first-chunk gap. **Likely diminishing returns** post Phase 9.2 chunk
tuning — TTS is already 280-380 ms; Z.AI TTFT (1300 ms) dominates.
4. **Tier 3 Frame-Stacked / VADUSA** (training-required, ~3-6 weeks).
Reserved for when a real consumer hits the current ceiling.
5. **Distilled CSM for wearable target** (10-30 days). Speculative;
probably wait until Sesame ships their distilled checkpoint or we
have a real on-device deploy target.
## Honest verdict
**rtx-csm is feature-complete relative to Sesame's open release.** The
remaining gaps to "literal deployed Maya" are:
- **Closed by porting SilentCipher** (1-2 days, recommended next chunk)
- **Gated on Sesame releasing more** (CSM-3B, CSM-8B, training data,
Maya fine-tune corpus) — probably never given the smart-glasses pivot
- **Architectural** (full-duplex via Moshi-style, distilled wearable
variant) — multi-week investments worth pursuing only when product
demand is concrete
Defensible product framing:
> rtx-csm is Sesame's open CSM-1B + voice cloning + production
> HTTP/WebSocket server + Kyutai-style cascaded duplex pipeline. The
> remaining gap to internal Maya is (a) SilentCipher watermarker
> (1-2 days), (b) Sesame's proprietary fine-tune dataset, (c) larger
> CSM variants Sesame chose not to release.
## Sources
- [Sesame — Crossing the uncanny valley of conversational voice](https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice)
- [SesameAILabs/csm](https://github.com/SesameAILabs/csm)
- [SesameAILabs/silentcipher](https://github.com/SesameAILabs/silentcipher)
- [SilentCipher paper, arXiv 2406.03822](https://arxiv.org/abs/2406.03822)
- [TechCrunch — Sesame raises $250M (Oct 2025)](https://techcrunch.com/2025/10/21/sesame-the-conversational-ai-startup-from-oculus-founders-raises-250m-and-launches-beta/)
- [TechCrunch — Sesame open-sources CSM-1B (Mar 2025)](https://techcrunch.com/2025/03/13/sesame-the-startup-behind-the-viral-virtual-assistant-maya-releases-its-base-ai-model/)
- [Kyutai Unmute](https://kyutai.org/unmute) / [unmute repo](https://github.com/kyutai-labs/unmute)
- [Moshi paper, arXiv 2410.00037](https://arxiv.org/html/2410.00037v2)
- [Voice agent latency benchmarks 2025](https://dev.to/cloudx/cracking-the-1-second-voice-loop-what-we-learned-after-30-stack-benchmarks-427)
- [HF sesame/csm-1b realtime discussion (200ms claim)](https://huggingface.co/sesame/csm-1b/discussions/10)