Files
rustytorch/docs/consolidation.md
T
osobhandClaude Fable 5 73102b71cf
Documentation / Build User Guide (push) Successful in 7s
CI / Build (macos-latest) (push) Failing after 9s
CI / Format Check (push) Failing after 12s
Documentation / Build API Documentation (push) Failing after 14s
Performance Benchmarks / Run Benchmarks (push) Failing after 31s
CI / Clippy Check (push) Failing after 30s
CI / Build (ubuntu-latest) (push) Failing after 36s
CI / Test (macos-latest) (push) Has been skipped
CI / Test (ubuntu-latest) (push) Has been skipped
CI / Python Bindings (maturin) (macos-latest) (push) Has been skipped
CI / Python Bindings (maturin) (ubuntu-latest) (push) Has been skipped
CI / WASM Build + Size Check (push) Has been skipped
CI / Distributed Training Tests (push) Has been skipped
CI / Build CPU-Only (Explicit) (push) Failing after 37s
CI / CI Success (push) Failing after 0s
feat(transformers): real compute in orchestrator modalities and attention planner
- orchestrator_core: execute_classical is genuine seeded QKV
  self-attention; execute_flash_attention delegates to
  rtx_flash_attention::flash_attention_forward (8-head reshape, clear
  Err on indivisible hidden); execute_hybrid composes the two;
  execute_edge/execute_distributed return explicit
  "modality not implemented" errors instead of fake success — which
  makes the previously-dead fallback_modalities retry loop real.
- tensor_core_kernels: AttentionComputationOptimizer::execute
  dispatches Standard/Online to scaled_dot_product_attention, Flash to
  rtx-flash-attention, Approximated to an explicit Err (no
  approximation kernel exists; refuses to compute exact attention
  under an approximated label).
- 10 new always-on tests incl. flash-vs-standard 1e-3 agreement and an
  Edge->Distributed->Classical fallback end-to-end.
- docs/consolidation.md updated: both entries moved from
  scaffolding/no-op-stub status to real dispatch descriptions.

982 rtx-transformers lib tests pass.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-07-10 07:06:50 -07:00

7.5 KiB

Duplicate-implementation consolidation notes

Recorded during the 2026-07-09 wiring/dead-code sweep, updated 2026-07-09 during the deferred "rewrite call sites" pass. This pass audited every site listed below; most turned out not to need (or not to be eligible for) a call-site rewrite once actually read in full. Details per item below.

Mixture of Experts

Canonical: rtx-transformers/src/layers/mixture_of_experts/ (MoEConfig, Router, Expert, MixtureOfExperts).

  • rtx-transformers/src/layers/metal_moe.rs (Metal-specific) — kept, no change needed. Its types already mirror the canonical naming 1:1 (MetalMoEConfig/MetalExpert/MetalRouter/MetalMoE vs. MoEConfig/Expert/Router/MixtureOfExperts), and its config fields (num_experts, top_k, hidden_dim, expert_hidden_dim, capacity_factor, dropout, aux_loss_weight, activation, bias) are identical to MoEConfig's, plus Metal-only extras (z_loss_weight, jitter_noise, use_metal). This is a backend specialization done right; forcing a shared struct would just add an indirection layer for no benefit.
  • rtx-transformers/src/modular/router.rsnot a duplicate; left as is. This is dynamic module selection (ModuleRouter, TaskConditionedRouter, AttentionRouter, RLRouter) over Vec<Box<dyn Module>> for the modular neural network system (modular/modular_network.rs), not token-to-FFN-expert gating. It's wired into ModuleRouter trait impls in modular_network.rs and has its own test suite (modular_networks_tests.rs). A genuinely different abstraction — MoE's Router routes tokens to FFN experts by learned gating logits; this routes whole inputs to heterogeneous modules (including RL-policy-based and task-conditioned selection). Not merging.
  • rtx-transformers/src/architectures/glam.rsout of scope, dead code. mod glam is commented out in architectures/mod.rs ("Temporarily disabled"), so this file is not part of the build at all. It also does not currently compile on its own: GLaMBlock::forward references an undefined routing_info binding (line ~229) — a pre-existing bug unrelated to consolidation. Migrating its expert layer to layers::mixture_of_experts::{MixtureOfExperts, MoEConfig, Router} is the right call when someone re-enables this module, but doing that migration on dead, non-compiling code with zero call sites and no test coverage today would be unverifiable busywork. Left as a note for whoever re-enables GLaM.

Deleted in the previous sweep (were orphaned, never declared by mod): layers/moe_layer.rs, layers/moe_integration.rs.

Flash Attention

Canonical: the rtx-flash-attention crate (v2+v3 kernels, CUDA + Metal, public API: flash_attention_forward/flash_attention_backward and the FlashAttention/FlashAttentionBackend types over rtx_tensor::Tensor).

Re-audited the three non-JEPA sites previously flagged as reimplementations. None of them actually contained duplicate attention math to delegate. Two of the three (tensor_core_kernels.rs, orchestrator_core.rs) have since been given real execution paths (2026-07-10); see below.

  • src/tensor_core_kernels.rs — was config/strategy scaffolding only (AttentionConfig, SoftmaxOptimizationStrategy, AttentionComputationOptimizer) with no compute/forward method anywhere in the file. Now real: AttentionComputationOptimizer::execute(&self, q, k, v: &Tensor) -> Result<Tensor> dispatches on self.softmax_strategy:
    • Standard / OnlineTensor::scaled_dot_product_attention (rtx_tensor matmul + softmax + matmul). Online softmax is mathematically equivalent to standard softmax (same result, different accumulation order) and no separate online-softmax kernel exists here, so it intentionally shares the standard path rather than faking a distinct computation.
    • FlashAttention → delegates to rtx_flash_attention::flash_attention_forward (the crate's own CUDA/Metal-or-naive-CPU-fallback logic applies unchanged).
    • Approximated → explicit Err: no approximation kernel is implemented, so this strategy refuses to silently compute exact attention under an "approximated" label.
    • Accepts 3D [batch, seq, head_dim] or 4D [batch, heads, seq, head_dim] q/k/v; other ranks are a clear Err.
    • Tests (tensor_core_kernels::real_execution_tests, always-on, not feature-gated): standard path is finite and input-dependent; flash vs. standard agree within 1e-3 on a small [1,2,4,8] shape; Approximated errors; wrong-rank input errors.
  • src/revolutionary/orchestrator_core.rsexecute_flash_attention and its siblings execute_classical/execute_hybrid/execute_edge/ execute_distributed were all identical no-op stubs (Ok(input.clone())) while the scoring/decision machinery around them was real. Now real:
    • execute_classical: genuine QKV self-attention — deterministic fixed-seed weight matrices (Tensor::randn_seeded) project the input, then scaled_dot_product_attention, then an output projection. Input-dependent, reproducible, not an identity.
    • execute_flash_attention: reshapes [batch, seq, hidden] into flash's [batch, heads, seq, head_dim] convention (8 heads; Err if hidden doesn't divide evenly) and calls rtx_flash_attention::flash_attention_forward directly.
    • execute_hybrid: real composition — execute_classical's projection feeds into execute_flash_attention's attention step, not either path run alone.
    • execute_edge / execute_distributed: explicit Err("modality not implemented: ...") — these need a target-device runtime handle / cluster communicator that this orchestrator has no way to obtain, so they no longer fake success. This also makes the orchestrator's fallback_modalities machinery real: execute_task already retried the next modality on Err, but previously nothing ever errored, so the fallback path was dead code; now Edge/Distributed failures genuinely exercise it.
    • Tests (orchestrator_core::real_modality_execution_tests, always-on): classical/flash/hybrid produce finite, input-dependent output (not equal to the input, different for different inputs, deterministic for the same input); flash rejects an indivisible hidden dim; edge/distributed error; a decision with primary_modality: Edge and fallback_modalities: [Distributed, Classical] ends up executed via Classical with result.success == true.
  • src/training/training_loop.rs — the only mention of FlashAttention is a doc comment listing it as an example LLM op; no implementation present.

src/ssl/jepa_gpu.rs is intentionally not in this list: its attention is part of the fused GPU-resident ViT block written for the JEPA platform, not a standalone duplicate of rtx-flash-attention's API — the fusion is the point (avoids materializing intermediate tensors across the ViT block same as flash attention's inner loop does, but tied to ViT-specific tensor layouts). Revisit only if rtx-flash-attention grows a fused-block API that covers this shape.

Speculative decoding

Layering is intentional (documented in rtx-inference/src/speculative/mod.rs): speculative/ is the orchestration layer (traits/configs/trees); rtx-inference/src/medusa.rs and src/lookahead.rs are the concrete implementations. Not a duplication to remove, but keep the two MedusaConfig types (orchestration vs implementation, re-exported as MedusaHeadsConfig) from drifting.