Files
rustytorch/crates/models/rtx-tts
osobhandClaude Sonnet 5 4aaa36a57a style: cargo fmt --workspace (whitespace/wrapping only, no semantic change)
Whole-workspace rustfmt pass picked up while iterating on Mamba GPU
backward work. Verified formatting-only via diff sampling; no logic
changed.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-08-10 07:09:36 -07:00
..
2026-03-04 00:08:42 +00:00
2026-03-04 00:08:42 +00:00
2026-03-04 00:08:42 +00:00
2026-03-04 00:08:42 +00:00

rtx-tts

Text-to-Speech (TTS) synthesis for RustyTorch with comprehensive text processing capabilities.

Features

  • Text Normalization: Advanced text preprocessing including:

    • Case normalization
    • Abbreviation expansion (Dr. → Doctor, St. → Street, etc.)
    • Number to word conversion (42 → "forty two")
    • Special character handling ($, %, °)
    • Whitespace normalization
  • Phonemization: Grapheme-to-phoneme (G2P) conversion:

    • Dictionary-based lookup for common English words
    • ARPAbet phoneme representation
    • Stress marker support
    • Rule-based fallback for unknown words
  • Prosody Prediction: Natural speech synthesis markers:

    • Pause detection at sentence and phrase boundaries
    • Emphasis prediction from capitalization
    • Intonation patterns (rising/falling)
    • Speaking rate estimation

Quick Start

use rtx_tts::{TextPipeline, TextNormalizerConfig};

// Create a text processing pipeline
let pipeline = TextPipeline::default();

// Process input text
let text = "Hello, Dr. Smith! I have 42 apples.";
let processed = pipeline.process_full(text)?;

println!("Normalized: {}", processed.normalized);
println!("Phonemes: {}", processed.phoneme_string(true));

// Prosody annotations
for annotation in &processed.prosody_annotations {
    println!("Segment: {} (markers: {:?})", annotation.text, annotation.markers);
}

Architecture

The TTS pipeline consists of three main stages:

  1. Text Normalization - Clean and normalize input text
  2. Phonemization - Convert text to phoneme sequences
  3. Prosody Prediction - Add timing and intonation markers

Examples

Text Normalization

use rtx_tts::TextNormalizer;

let normalizer = TextNormalizer::default();
let result = normalizer.normalize("Dr. Smith has 50 dollars and 3 cats")?;
// Output: "doctor smith has fifty dollars and three cats"

Phonemization

use rtx_tts::Phonemizer;

let phonemizer = Phonemizer::default();
let phonemes = phonemizer.g2p("hello world")?;
// Returns ARPAbet phoneme sequences with stress markers

Prosody Prediction

use rtx_tts::{Prosody, ProsodyMarker};

let prosody = Prosody::default();
let annotations = prosody.predict("How are you?")?;

// Check for rising intonation (questions)
let has_rising = annotations.iter().any(|a| {
    a.markers.iter().any(|m| matches!(m, ProsodyMarker::Rising))
});

Configuration

Each module can be configured independently:

use rtx_tts::{
    TextNormalizer, TextNormalizerConfig,
    Phonemizer, PhonemizerConfig, PhonemizerBackend,
    Prosody, ProsodyConfig,
    TextPipeline,
};

// Custom normalization
let norm_config = TextNormalizerConfig {
    lowercase: true,
    remove_punctuation: false,
    expand_abbreviations: true,
    normalize_numbers: true,
    normalize_whitespace: true,
};

// Custom phonemization
let phone_config = PhonemizerConfig {
    language: "en-us".to_string(),
    backend: PhonemizerBackend::Dictionary,
    include_stress: true,
};

// Custom prosody
let prosody_config = ProsodyConfig {
    use_punctuation_pauses: true,
    default_speaking_rate: 1.0,
    pause_at_commas: true,
    pause_at_sentences: true,
    detect_emphasis: true,
};

// Build pipeline
let pipeline = TextPipeline::new(
    TextNormalizer::new(norm_config),
    Phonemizer::new(phone_config),
    Prosody::new(prosody_config),
);

Testing

Run tests with:

cargo test

Run with coverage:

cargo test --all-features

Performance

All text processing operations are implemented in pure Rust with zero allocations where possible. The phonemizer uses efficient HashMap lookups for dictionary-based phonemization.

Future Work

  • Neural acoustic models for mel-spectrogram generation
  • GPU-accelerated neural vocoders (WaveGlow, HiFi-GAN)
  • Multi-language support
  • SSML markup support
  • Voice cloning capabilities

License

This project is licensed under the same terms as RustyTorch (MIT OR Apache-2.0).