Full encoder forward path: conv stem -> 6 transformer layers -> final
LayerNorm. Loads HF safetensors, runs end-to-end on Metal.
Components added to src/moonshine.rs:
RotaryCache partial RoPE (32 of 36 head_dim, theta=10000)
EncoderAttention MHA (8 heads, no bias), partial RoPE on q/k
EncoderMlp 288 -> 1152 -> 288 with bias, GELU(erf) activation
EncoderLayer Pre-LN attn + Pre-LN MLP (LayerNorm weight-only)
Encoder stem + 6 layers + final LayerNorm
load_encoder() VarBuilder convenience for the standalone smoke
Smoke test (`examples/moonshine_smoke`) verified end-to-end:
input (1, 1, 160000) -> output (1, 415, 288)
forward: 132 ms (10 s of audio at 0.013x realtime)
max abs: 6.67 (signal preserved, not zeros)
Implementation notes captured in the diff:
- candle Metal 4D batched matmul had shape-mismatch issues for our
(B, H, T, D) pattern. Switched to (B*H, T, D) 3D form which is
unambiguous and avoids the kernel bug.
- LayerNorm is weight-only (no bias tensors in safetensors); we
construct LayerNorm with a zeros bias to satisfy candle's API.
- rotary_dim = floor(head_dim * 0.9 / 2) * 2 = 32 (must be even).
The remaining 4 head_dim channels pass through unchanged via
`narrow + cat` on dim 3.
Numerical parity vs HF Python reference is NOT yet verified — that's
the next bounded chunk (Phase 8.7). Shape + signal correctness are
verified by the smoke test.
Next: decoder transformer block (self-attn + cross-attn + SwiGLU).
~3-4 h of focused work.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>