Adds examples/make_silence_test.rs which builds a 10 s WAV that is
50% silence + 50% real speech (1s silence | 4s speech | 5s silence)
so we can see the VAD gate work clearly.
Bench A/B (mock LLM, Q8+stream, 3 turns each, M-series Metal):
Audio No VAD With VAD Δ recv_phase
90% speech (LibriSpeech) 4834 ms 4509 ms -7%
50% silence (synthetic) 6490 ms 2204 ms -66%
The structural win scales with silence content as expected. Real-world
voice-agent audio (30-50% silence per typical call-center / voice-bot
benchmarks) will see ~30-50% recv_phase reduction. The earlier 7% on
LibriSpeech wasn't a weak result — it accurately reflected the ~10%
silence in that recording.
This validates the energy-VAD path despite Silero V5 via ort being
blocked (Phase 8.1.3). Production voice loops should default to
--vad-gate.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>