Files
rustytorch/docs/consolidation.md
T
osobhandClaude Fable 5 73102b71cf
Documentation / Build User Guide (push) Successful in 7s
CI / Build (macos-latest) (push) Failing after 9s
CI / Format Check (push) Failing after 12s
Documentation / Build API Documentation (push) Failing after 14s
Performance Benchmarks / Run Benchmarks (push) Failing after 31s
CI / Clippy Check (push) Failing after 30s
CI / Build (ubuntu-latest) (push) Failing after 36s
CI / Test (macos-latest) (push) Has been skipped
CI / Test (ubuntu-latest) (push) Has been skipped
CI / Python Bindings (maturin) (macos-latest) (push) Has been skipped
CI / Python Bindings (maturin) (ubuntu-latest) (push) Has been skipped
CI / WASM Build + Size Check (push) Has been skipped
CI / Distributed Training Tests (push) Has been skipped
CI / Build CPU-Only (Explicit) (push) Failing after 37s
CI / CI Success (push) Failing after 0s
feat(transformers): real compute in orchestrator modalities and attention planner
- orchestrator_core: execute_classical is genuine seeded QKV
  self-attention; execute_flash_attention delegates to
  rtx_flash_attention::flash_attention_forward (8-head reshape, clear
  Err on indivisible hidden); execute_hybrid composes the two;
  execute_edge/execute_distributed return explicit
  "modality not implemented" errors instead of fake success — which
  makes the previously-dead fallback_modalities retry loop real.
- tensor_core_kernels: AttentionComputationOptimizer::execute
  dispatches Standard/Online to scaled_dot_product_attention, Flash to
  rtx-flash-attention, Approximated to an explicit Err (no
  approximation kernel exists; refuses to compute exact attention
  under an approximated label).
- 10 new always-on tests incl. flash-vs-standard 1e-3 agreement and an
  Edge->Distributed->Classical fallback end-to-end.
- docs/consolidation.md updated: both entries moved from
  scaffolding/no-op-stub status to real dispatch descriptions.

982 rtx-transformers lib tests pass.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-07-10 07:06:50 -07:00

129 lines
7.5 KiB
Markdown

# Duplicate-implementation consolidation notes
Recorded during the 2026-07-09 wiring/dead-code sweep, updated 2026-07-09
during the deferred "rewrite call sites" pass. This pass audited every site
listed below; most turned out not to need (or not to be eligible for) a
call-site rewrite once actually read in full. Details per item below.
## Mixture of Experts
Canonical: `rtx-transformers/src/layers/mixture_of_experts/` (`MoEConfig`,
`Router`, `Expert`, `MixtureOfExperts`).
- `rtx-transformers/src/layers/metal_moe.rs` (Metal-specific) — **kept, no
change needed.** Its types already mirror the canonical naming 1:1
(`MetalMoEConfig`/`MetalExpert`/`MetalRouter`/`MetalMoE` vs.
`MoEConfig`/`Expert`/`Router`/`MixtureOfExperts`), and its config fields
(`num_experts`, `top_k`, `hidden_dim`, `expert_hidden_dim`,
`capacity_factor`, `dropout`, `aux_loss_weight`, `activation`, `bias`) are
identical to `MoEConfig`'s, plus Metal-only extras (`z_loss_weight`,
`jitter_noise`, `use_metal`). This is a backend specialization done right;
forcing a shared struct would just add an indirection layer for no benefit.
- `rtx-transformers/src/modular/router.rs`**not a duplicate; left as is.**
This is dynamic *module* selection (`ModuleRouter`, `TaskConditionedRouter`,
`AttentionRouter`, `RLRouter`) over `Vec<Box<dyn Module>>` for the modular
neural network system (`modular/modular_network.rs`), not token-to-FFN-expert
gating. It's wired into `ModuleRouter` trait impls in `modular_network.rs`
and has its own test suite (`modular_networks_tests.rs`). A genuinely
different abstraction — MoE's `Router` routes tokens to FFN experts by
learned gating logits; this routes whole inputs to heterogeneous modules
(including RL-policy-based and task-conditioned selection). Not merging.
- `rtx-transformers/src/architectures/glam.rs`**out of scope, dead code.**
`mod glam` is commented out in `architectures/mod.rs` ("Temporarily
disabled"), so this file is not part of the build at all. It also does not
currently compile on its own: `GLaMBlock::forward` references an undefined
`routing_info` binding (line ~229) — a pre-existing bug unrelated to
consolidation. Migrating its expert layer to
`layers::mixture_of_experts::{MixtureOfExperts, MoEConfig, Router}` is the
right call *when* someone re-enables this module, but doing that migration
on dead, non-compiling code with zero call sites and no test coverage today
would be unverifiable busywork. Left as a note for whoever re-enables GLaM.
Deleted in the previous sweep (were orphaned, never declared by `mod`):
`layers/moe_layer.rs`, `layers/moe_integration.rs`.
## Flash Attention
Canonical: the `rtx-flash-attention` crate (v2+v3 kernels, CUDA + Metal,
public API: `flash_attention_forward`/`flash_attention_backward` and the
`FlashAttention`/`FlashAttentionBackend` types over `rtx_tensor::Tensor`).
Re-audited the three non-JEPA sites previously flagged as reimplementations.
None of them actually contained duplicate attention math to delegate. Two of
the three (`tensor_core_kernels.rs`, `orchestrator_core.rs`) have since been
given real execution paths (2026-07-10); see below.
- `src/tensor_core_kernels.rs` — was config/strategy scaffolding only
(`AttentionConfig`, `SoftmaxOptimizationStrategy`,
`AttentionComputationOptimizer`) with no `compute`/`forward` method
anywhere in the file. **Now real:**
`AttentionComputationOptimizer::execute(&self, q, k, v: &Tensor) -> Result<Tensor>`
dispatches on `self.softmax_strategy`:
- `Standard` / `Online``Tensor::scaled_dot_product_attention`
(`rtx_tensor` matmul + softmax + matmul). Online softmax is
mathematically equivalent to standard softmax (same result, different
accumulation order) and no separate online-softmax kernel exists here,
so it intentionally shares the standard path rather than faking a
distinct computation.
- `FlashAttention` → delegates to `rtx_flash_attention::flash_attention_forward`
(the crate's own CUDA/Metal-or-naive-CPU-fallback logic applies
unchanged).
- `Approximated` → explicit `Err`: no approximation kernel is implemented,
so this strategy refuses to silently compute exact attention under an
"approximated" label.
- Accepts 3D `[batch, seq, head_dim]` or 4D `[batch, heads, seq, head_dim]`
q/k/v; other ranks are a clear `Err`.
- Tests (`tensor_core_kernels::real_execution_tests`, always-on, not
feature-gated): standard path is finite and input-dependent; flash vs.
standard agree within 1e-3 on a small `[1,2,4,8]` shape; `Approximated`
errors; wrong-rank input errors.
- `src/revolutionary/orchestrator_core.rs``execute_flash_attention` and
its siblings `execute_classical`/`execute_hybrid`/`execute_edge`/
`execute_distributed` were all identical no-op stubs (`Ok(input.clone())`)
while the scoring/decision machinery around them was real. **Now real:**
- `execute_classical`: genuine QKV self-attention — deterministic
fixed-seed weight matrices (`Tensor::randn_seeded`) project the input,
then `scaled_dot_product_attention`, then an output projection.
Input-dependent, reproducible, not an identity.
- `execute_flash_attention`: reshapes `[batch, seq, hidden]` into flash's
`[batch, heads, seq, head_dim]` convention (8 heads; `Err` if `hidden`
doesn't divide evenly) and calls
`rtx_flash_attention::flash_attention_forward` directly.
- `execute_hybrid`: real composition — `execute_classical`'s projection
feeds into `execute_flash_attention`'s attention step, not either path
run alone.
- `execute_edge` / `execute_distributed`: explicit
`Err("modality not implemented: ...")` — these need a target-device
runtime handle / cluster communicator that this orchestrator has no way
to obtain, so they no longer fake success. This also makes the
orchestrator's `fallback_modalities` machinery real: `execute_task`
already retried the next modality on `Err`, but previously nothing ever
errored, so the fallback path was dead code; now Edge/Distributed
failures genuinely exercise it.
- Tests (`orchestrator_core::real_modality_execution_tests`, always-on):
classical/flash/hybrid produce finite, input-dependent output (not equal
to the input, different for different inputs, deterministic for the same
input); flash rejects an indivisible hidden dim; edge/distributed error;
a decision with `primary_modality: Edge` and
`fallback_modalities: [Distributed, Classical]` ends up executed via
Classical with `result.success == true`.
- `src/training/training_loop.rs` — the only mention of `FlashAttention` is a
doc comment listing it as an example LLM op; no implementation present.
`src/ssl/jepa_gpu.rs` is intentionally **not** in this list: its attention is
part of the fused GPU-resident ViT block written for the JEPA platform, not
a standalone duplicate of `rtx-flash-attention`'s API — the fusion is the
point (avoids materializing intermediate tensors across the ViT block same
as flash attention's inner loop does, but tied to ViT-specific tensor
layouts). Revisit only if `rtx-flash-attention` grows a fused-block API that
covers this shape.
## Speculative decoding
Layering is intentional (documented in `rtx-inference/src/speculative/mod.rs`):
`speculative/` is the orchestration layer (traits/configs/trees);
`rtx-inference/src/medusa.rs` and `src/lookahead.rs` are the concrete
implementations. Not a duplication to remove, but keep the two `MedusaConfig`
types (orchestration vs implementation, re-exported as `MedusaHeadsConfig`)
from drifting.