Files
rustytorch/docs/superpowers/specs/2026-06-26-rustytorch-full-review.md
T
Omar SobhandClaude Sonnet 4.6 57e5252caa docs: add full repository review spec (2026-06-26)
Comprehensive layered architectural review covering all 109 crates,
306K LOC, 13,351 tests. Identifies 9 gaps (G0-G8) with the highest-
priority being 45 unimplemented! panics across rtx-backend-cuda/rocm/sycl
and the rtx-distributed workspace exclusion. Includes 17-item 4-phase
roadmap through 90 days.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-06-26 12:48:03 +00:00

820 lines
33 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RustyTorch Full Repository Review
**Date**: 2026-06-26
**Crates**: 109
**LOC**: 306,423
**Tests**: 13,351+
**Incomplete stubs**: 51
**Toolchain**: nightly-2025-10-25 (Rust 2024 edition)
---
## 1. Project Overview
RustyTorch is a production-grade, GPU-accelerated ML framework in pure Rust. It aims to match PyTorch's feature breadth while eliminating the Python GIL, delivering memory-safe tensor operations, automatic differentiation, and domain-specific stacks for medical imaging, neuroimaging, and scientific computing.
### 1.1 Platform Role
**Layer 8 — ML Infrastructure** for the QuantumRedClaw platform:
- **QPUDIDP** consumes RustyTorch for surrogate model training (MLP, normalizing flows, MC Dropout)
- **qstar** neural decoders depend on GPU compute provided by this framework
- Python interop via PyO3/maturin bridges to Jupyter notebooks
### 1.2 Maturity Snapshot
| Metric | Value |
|--------|-------|
| Total crates | 109 |
| Total Rust LOC | 306,423 |
| `#[test]` attributes | 13,351+ |
| `todo!/unimplemented!` | 51 |
| CI pipelines | 5 (format, clippy, build, test, GPU) |
| Demo applications | 30+ |
| Rust toolchain | nightly-2025-10-25 |
| Edition | 2024 |
| Build target dir | `/hot/targets/rustyverse/rustytorch` |
| TIER progress | TIER 03 complete (Dec 28, 2025) |
### 1.3 Key Differentiators vs. Ecosystem
| Feature | RustyTorch | tch-rs | burn | candle |
|---------|-----------|--------|------|--------|
| Pure Rust (no C/C++) | ✅ | ❌ (libtorch) | ✅ | ✅ |
| Multi-backend dispatch | ✅ | ❌ | ✅ | partial |
| Flash Attention | ✅ | via ext | ❌ | ❌ |
| Mixture of Experts | ✅ | ❌ | ❌ | ❌ |
| Mamba/SSM | ✅ | ❌ | ❌ | ❌ |
| Speculative decoding | ✅ | ❌ | ❌ | ❌ |
| Medical imaging stack | ✅ | ❌ | ❌ | ❌ |
| Neuroimaging (EEG/MEG) | ✅ | ❌ | ❌ | ❌ |
| Neural Operators (FNO) | ✅ | ❌ | ❌ | ❌ |
| Federated learning | ✅ | ❌ | ❌ | ❌ |
### 1.4 GPU Backend Support
| Backend | Hardware | Status |
|---------|----------|--------|
| CUDA (sm_86 default, sm_120 Blackwell) | NVIDIA RTX/A100/H100/5060 Ti | Production |
| Metal (MPS) | Apple M-series | Production |
| ROCm/HIP | AMD RDNA 2/3 | Partial (15 ops unimplemented) |
| SYCL/DPC++ | Intel Arc | Partial (15 ops unimplemented) |
| CubeCL (WGPU) | Cross-platform | Production |
| CPU (OpenBLAS/MKL/manual) | x86/ARM | Production |
---
## 2. Layer-by-Layer Review
### Layer 1: Core Infrastructure
#### 2.1.1 rtx-backend
**Status**: Production-Ready
**Purpose**: Compile-time backend dispatch trait — zero runtime overhead
**Path**: `crates/core/rtx-backend/`
Key API surface (522 LOC, `src/lib.rs`):
- `Backend` trait: 430+ methods covering tensor creation, element-wise ops, reductions, LLM-specific ops (flash_attention, rope, rms_norm)
- `AutodiffBackend`: marker trait for gradient-capable backends
- `QuantizedBackend`: quantization support
Test coverage: Feature-gated; integration via backend crates
Gaps: Trait is intentionally abstract. The gap is in the concrete backends (see rtx-backend-cuda, rtx-backend-rocm, rtx-backend-sycl below).
---
#### 2.1.2 rtx-tensor
**Status**: Production-Ready
**Purpose**: PyTorch-compatible GPU-native tensor primitives
**Path**: `crates/core/rtx-tensor/`
Key API surface:
- `Tensor`, `Device`, `DType`, `Shape`
- `ComplexTensor`, `SparseCOO`, `SparseCSR`
- `CublasManager`, `CudnnContext`, `CuSolverContext`, `CuSparseContext`
- 35+ tensor operation modules, 7 Metal-specific modules
Test coverage: 1,047 `#[test]` in src, 184 in tests/; `loom` concurrency tests; 10+ benchmark files
Notable: Real CUDA kernel compilation via cudarc 0.18.x, real MSL shaders for Metal
Gaps: 2 GPU memory access tests marked `RED: unimplemented` in `storage/core/` (flagged as intentional TDD red phase)
---
#### 2.1.3 rtx-autograd
**Status**: Production-Ready
**Purpose**: Tape-based automatic differentiation, forward and reverse modes
**Path**: `crates/core/rtx-autograd/`
Key API surface:
- `Autodiff<B>`: decorator wrapping any `Backend<B>` to add gradients
- `grad()`, `jvp()`, `vjp()`, `jacfwd()`, `jacrev()`, `hessian()`, `vmap()`
- Checkpointing: `checkpoint()`, `EveryNthStrategy`, `SqrtCheckpointStrategy`, `AdaptiveCheckpointStrategy`
- `AutogradProfiler` with Chrome trace export
Test coverage: 223 `#[test]` in tests/, 128 in src/; `gradient_benchmarks.rs`
Gaps: None detected.
---
#### 2.1.4 rtx-memory
**Status**: Production-Ready
**Purpose**: GPU memory pooling, KV cache, ZeRO-style partitioning
**Path**: `crates/core/rtx-memory/`
Key API surface:
- `MemoryPoolManager`, `TransformerMemoryPools`
- `KvCacheManager`, `AsyncKvCacheManager`, `EntropyTracker`
- `ZeroMemoryManager`, `PagedAttentionManager`
- Allocators: Buddy, Slab, Stack
Test coverage: 190 `#[test]` in src/, 190 in tests/; 3 benchmark files (pool, allocation strategies, simple GPU)
Gaps: None detected. Fragmentation target <5% is documented.
---
#### 2.1.5 rtx-kernel
**Status**: Production-Ready
**Purpose**: CUDA PTX + Metal MSL kernel compilation, caching, execution
**Path**: `crates/core/rtx-kernel/`
Key API surface:
- `KernelSystem`, `KernelCompiler` (NVRTC-based), `KernelCache`, `KernelExecutor`
- `MetalKernelCompiler`, `MetalKernelCache`, `MetalKernelExecutor`
- `LaunchConfig`, `KernelParam` (type-safe binding)
Gaps: None detected.
---
#### 2.1.6 rtx-flash-attention
**Status**: Production-Ready (with one stub in backward pass)
**Purpose**: Memory-efficient attention, O(n) memory, 5-8× speedup
**Path**: `crates/training/rtx-flash-attention/`
Key API surface:
- `FlashAttention`, `FlashAttentionFactory`, `FlashAttentionConfig`
- `SdpaBackendSelector` (hardware-aware: FlashAttn2, SDPA, naive fallback)
- `flash_attention_forward()`, `flash_attention_backward()`
Test coverage: 97 `#[test]` in src/; `flash_attention_bench.rs`
Notable: sm_120 (Blackwell/RTX 5060 Ti) support added in commit c2f4796
Gaps: Backward pass gradient computation has a placeholder fallback path for non-CUDA; acceptable since CUDA is the target.
---
#### 2.1.7 rtx-lora
**Status**: Production-Ready
**Purpose**: LoRA and QLoRA parameter-efficient fine-tuning
**Path**: `crates/core/rtx-lora/`
Key API surface:
- `LoraConfig`, `LoraAdapter`, `LoraWeight`, `QLoraWeight`
- `NF4Block` (4-bit quantization), `AdapterManager`, `MergeStrategy`
Test coverage: Tests in adapter.rs, config.rs, weight.rs, quantized.rs, merge.rs
Gaps: None detected.
---
#### 2.1.8 rtx-moe
**Status**: Production-Ready
**Purpose**: Mixture of Experts (SwitchTransformer, ExpertChoice, TokenChoice routing)
**Path**: `crates/core/rtx-moe/`
Key API surface: SwitchTransformer, ExpertChoice, TokenChoice routing strategies; load balancing; expert capacity controls
Gaps: None detected from exploration. (Dedicated Metal MSL kernels in rtx-tensor/metal_moe_ops.)
---
#### 2.1.9 rtx-mamba
**Status**: Production-Ready (conv1d history buffer improved in D309)
**Purpose**: Mamba/State Space Models, O(n) attention alternative
**Path**: `crates/core/rtx-mamba/`
Key API surface: `MambaRecurrence`, selective scan, `MambaStep` (new ergonomic step API from commit f751414)
Gaps: None detected.
---
#### 2.1.10 rtx-speculative-decoding
**Status**: Production-Ready
**Purpose**: 2-3× inference speedup via draft model + target verification
**Path**: `crates/core/rtx-speculative-decoding/` and `crates/production/rtx-inference/`
Key API surface: `SpeculativeEngine`, `AdvancedSpeculativeDecoder`, `EagleDraftModel`, `MedusaDraftModel`, tree-based decoding
Metal ops: `metal_speculative_ops` in rtx-tensor
Gaps: None detected.
---
#### 2.1.11 rtx-fusion
**Status**: Production-Ready
**Purpose**: Kernel fusion DSL eliminating intermediate tensor writes (10-25% speedup)
**Path**: `crates/core/rtx-fusion/`
Key API surface:
- `Fusion<B>` backend wrapper
- `FusionAnalyzer` (pattern detection), `FusedKernelCache` (LRU), `FusionConfig`, `FusionStats`
- Sync points: matmul, reductions, flash_attention, data access
Gaps: None detected.
---
#### 2.1.12 rtx-cubecl
**Status**: Production-Ready
**Purpose**: Single-source Rust → GPU kernel synthesis via CubeCL
**Path**: `crates/core/rtx-cubecl/`
Key API surface: `CubeclBackend`, `CubeclClient`, `CubeclDevice`, `CubeclTensor`
Backends: cpu, wgpu, cuda, hip, vulkan, metal
Tier strategy: ~60% ops via portable CubeCL kernels; ~40% via hand-crafted tier-2 (Flash Attention, GEMM, conv)
Gaps: None detected.
---
#### 2.1.13 rtx-bindings
**Status**: Partial
**Purpose**: Python (PyO3), C FFI, ONNX, DLPack interop
**Path**: `crates/core/rtx-bindings/`
Key API surface: `c_api` (enabled), `python` (disabled), `onnx` (basic), `dlpack` (basic)
Test coverage: test files exist but are conditionally disabled
**Gaps**:
- [G1] PyO3 Python bindings disabled due to version conflicts
- [G7] ONNX/DLPack are basic stubs — no full model export or zero-copy interchange
---
#### 2.1.14 rtx-macros
**Status**: Production-Ready
**Purpose**: Derive macros: `#[derive(Module)]`, `#[derive(Config)]`
**Path**: `crates/core/rtx-macros/`
Key API surface:
- `Module` derive: generates `parameters()`, `to_device()`, `train()`, field attrs: `#[param]`, `#[module]`, `#[constant]`
- `Config` derive: builder pattern with `#[config(default = "value")]`
Gaps: None detected.
---
#### 2.1.15 rtx-backend-cuda / rtx-backend-rocm / rtx-backend-sycl
**Status**: Partial — **Critical gap**
**Path**: `crates/core/rtx-backend-{cuda,rocm,sycl}/src/lib.rs`
Each of these three backend adapters is missing the same 15 operations:
```
sin, cos, pow, clamp, relu, sigmoid, tanh,
var, var_dim, leaky_relu, elu, gt_scalar,
conv2d, max_pool2d, avg_pool2d
```
These are `unimplemented!()` panics — **any code path that routes through the backend dispatch layer for these ops will panic at runtime** on CUDA, ROCm, and SYCL.
**Note**: `rtx-tensor` has its own GPU implementations (via cudarc + cuDNN directly) that do not go through the backend dispatch path. But any consumer using the generic `Backend` trait abstraction hits these panics.
This is gap **G0** — the highest-priority fix in this review.
---
### Layer 2: Training & Optimization
#### 2.2.1 rtx-transformers
**Status**: Production-Ready (core); Phase 2-3 modules deferred
**Path**: `crates/training/rtx-transformers/`
Key API surface:
- Layers: Mamba, MoE, RoPE, Linear Attention, Spectral Norm, Block Sparse Attention, GeLU, SwiGLU, GatedMLP
- Architectures: BERT, GPT, TransformerBlock, MultiHeadAttention
- Optimizers: Adam, AdamW, Shampoo
- New (commit 24bd5cf): `mamba_step`, `ClonedMemoryUpdater`, `GatedMemoryUpdater`, `SetEncoderTeacher`
Test coverage: **1,845 `#[test]`** — highest in the workspace
Notable: Also has integration tests at `tests/d350_gpu_backbone_training.rs`, `tests/mamba_temporal_learning.rs`, `tests/real_selective_scan.rs`
**Gaps**: [G4] Phase 2 modules commented out: `validation_framework`, `meta-learning`, advanced tokenization, multi-modal text, continual learning
---
#### 2.2.2 rtx-distributed
**Status**: Production-Ready (code) — **excluded from workspace**
**Path**: `crates/training/rtx-distributed/`
Key API surface:
- Collective ops: AllReduce, Broadcast, Gather, Scatter via NCCL/RCCL/MPI
- Parallelism: `DistributedDataParallel`, `FSDP2`, `TensorParallel`, `PipelineParallel`
- Elastic training: `ElasticCluster`, `FaultDetector`, `ActivationCheckpointManager`
- DCP: Distributed Checkpoint with `DeviceMesh`, `HardwareTopology`
Test coverage: 439 `#[test]`
**Gaps**: [G2] Excluded from workspace Cargo.toml due to RNCCL path dependency — cannot be tested in CI, cannot be used as a dependency without manual workspace inclusion.
---
#### 2.2.3 rtx-rl
**Status**: Production-Ready
**Path**: `crates/training/rtx-rl/`
Key API surface: `PPO`, `DQN`, `SAC`, `ActorLearner`, `Environment`, `ReplayBuffer`, `RLHF`
Test coverage: 6 test files (PPO, SAC, RLHF, actor-learner, environment, replay buffer); `rl_bench` benchmark
Gaps: None detected.
---
#### 2.2.4 rtx-compress
**Status**: Production-Ready
**Path**: `crates/training/rtx-compress/`
Key API surface: `CompressionPipeline`, `CompressionConfig`, `CompressedStorage`, quantization, pruning, distillation, KV cache, LoRA
Test coverage: 164 `#[test]` in tests/; `compression_bench`
Gaps: Some advanced quantization modes (AWQ, GPTQ, SmoothQuant) are in progress — basic pipeline is complete.
---
#### 2.2.5 rtx-nas
**Status**: Production-Ready
**Path**: `crates/training/rtx-nas/`
Key API surface: `DARTS`, `PC-DARTS`, `FairNAS`, `RandomSearch`, `Architecture`, `DARTSCell`, `SearchSpace`
Test coverage: 165 `#[test]`
Gaps: None detected.
---
#### 2.2.6 rtx-federated
**Status**: Production-Ready
**Path**: `crates/training/rtx-federated/`
Key API surface:
- Aggregation: FedAvg, FedProx, SCAFFOLD, FedNova, AsyncAggregation
- Byzantine-robust: Krum, MultiKrum, TrimmedMean, AnomalyDetector, ReputationSystem
- Privacy: DifferentialPrivacy, HomomorphicEncryption (tfhe), SMPC, PrivacyBudget
- Personalization: MetaLearning, ClientClustering, MultiTaskLearning, TransferLearning
Test coverage: `federated_bench`; `federated_demo` example
Gaps: None detected. 80+ dependencies including cryptographic primitives.
---
#### 2.2.7 rtx-auto
**Status**: Production-Ready
**Path**: `crates/training/rtx-auto/`
Key API surface: `AutonomousAgent`, `DataEngineeringAgent`, `ParallelPlannerAgent`, `QuantGuardianAgent`, `KernelSynthesizerAgent`, `ProposalValidator`, `RollbackManager`
Test coverage: Integration test for `AutonomousOptimizer` initialization
Gaps: None detected.
---
#### 2.2.8 rtx-evolution
**Status**: Production-Ready
**Path**: `crates/training/rtx-evolution/`
Key API surface: `EvolutionOrchestrator`, `TelemetryAnalyzer`, `SafeSandbox`, `KnowledgeGraph`, `MultiObjectiveOptimizer`, `ParetoFrontier`, `HyperparameterTuner`
Test coverage: `orchestrator_tests.rs`
Gaps: None detected.
---
### Layer 3: Model Architectures
#### 2.3.1 rtx-vision
**Status**: Production-Ready
**Path**: `crates/models/rtx-vision/`
Key API surface: ViT, ConvNeXt V2, EfficientNetV2, MobileViT, DenseNet, NFNet, EdgeViT, MaxViT, CoAtNet (150+ exported types)
Preprocessing: `ImageTensor`, `Augmentation`, `ImageProcessor`
Gaps: None detected.
---
#### 2.3.2 rtx-multimodal
**Status**: Production-Ready (quantum/neuromorphic stubs remain)
**Path**: `crates/models/rtx-multimodal/`
Key API surface: `MultimodalSystem`, `VisionTransformer`, `AudioTransformer`, `CrossModalAttention`, `ModalityFusion` (7 strategies)
Test coverage: 122 `#[test]`; `multimodal_bench`
**Gaps**: [G8] Quantum enhancement and neuromorphic preprocessing modules have TODO stubs.
---
#### 2.3.3 rtx-diffuse
**Status**: Production-Ready
**Path**: `crates/models/rtx-diffuse/`
Key API surface: UNet, DiT (Diffusion Transformer), DDIMSampler, DPMSolverPP, UniPCSampler (5-10 step fast inference), ControlNet, IP-Adapter, T2I-Adapter, ClassifierFreeGuidance, LoRA diffusion
Test coverage: 141 `#[test]` in src/; validation tests; `diffusion_bench`
Gaps: None detected.
---
#### 2.3.4 rtx-timeseries
**Status**: Production-Ready (quantum stubs remain)
**Path**: `crates/models/rtx-timeseries/`
Key API surface: ARIMAModel, SARIMAModel, ProphetModel, NeuralProphetModel, ExponentialSmoothingModel, StateSpaceModel, TransformerForecastModel, `Forecaster`, `TimeSeriesAnalyzer` (seasonality, stationarity, anomaly detection)
Test coverage: Criterion benchmarks
**Gaps**: [G8] Quantum forecasting modules marked `// TODO: Implement`.
---
#### 2.3.5 rtx-nlg
**Status**: Production-Ready
**Path**: `crates/models/rtx-nlg/`
Key API surface: `TextGenerator`, `ModelServer`, `StreamingGenerator`, `BatchGenerator`, `ConversationManager`, `BeamSearch`, `NucleusSampling`, `RepetitionPenalty`, `ToxicityFilter`
Test coverage: 3 benchmark suites (generation, translation, summarization)
Gaps: None detected.
---
### Layer 4: Production & Deployment
#### 2.4.1 rtx-serving-api
**Status**: Production-Ready
**Path**: `crates/production/rtx-serving-api/`
Key API surface: `ServingServer`, `AdvancedServingServer`, `ContinuousBatchingController`, `QueueManager`, `RateLimitManager`, `CircuitBreaker`, `GrammarSampler` (BNF parser), `StructuredGenerator` (JSON schema), `MultiModelManager`, WebSocket streaming, `BudgetConfig`
Stack: Axum (HTTP), Tonic (gRPC), tokio-tungstenite (WebSocket), Prometheus
Test coverage: 140 `#[test]` in src/; `cache_performance` benchmark
Gaps: None detected.
---
#### 2.4.2 rtx-inference
**Status**: Production-Ready
**Path**: `crates/production/rtx-inference/`
Key API surface: `InferenceEngine`, `PagedKvCache` (GPU→CPU→NVMe tiering), `BatchScheduler` (SLA lanes), `AdvancedSpeculativeDecoder`, `EagleDraftModel`, `MedusaDraftModel`, `QuantizationConfig`, `MetricsCollector`
Test coverage: Integration tests with proptest
Features: `onnx-runtime`, `burn`, `burn-wgpu`, `candle`, `candle-cuda`, `candle-metal`
Gaps: None detected.
---
#### 2.4.3 rtx-streaming
**Status**: Production-Ready
**Path**: `crates/production/rtx-streaming/`
Key API surface: `StreamingServer`, `ConnectionManager`, `TokenGenerator`, `BackpressureHandler`, `AdaptiveProcessor`, Kafka integration, edge computing support
Design target: sub-millisecond latency, >1000 concurrent connections
Test coverage: `streaming_bench`
Gaps: None detected.
---
#### 2.4.4 rtx-wasm-inference
**Status**: Production-Ready
**Path**: `crates/production/rtx-wasm-inference/`
Key API surface: `WasmInferenceEngine`, `ComputeBackend` (Cpu/WebGpu/Auto), `TensorCore`, `DType`
Features: WASM SIMD, SharedArrayBuffer threads, WebGPU acceleration, TypeScript/JavaScript bindings
Build: `opt-level = "s"`, LTO enabled, wasm-opt with SIMD
Gaps: None detected.
---
#### 2.4.5 rtx-onnx
**Status**: Partial
**Path**: `crates/production/rtx-onnx/`
Key API surface: `OnnxSession`, `OnnxSessionConfig`, `ExecutionProviderType` (CPU/CUDA/CoreML/TensorRT/DirectML), `tensor_bridge` (rtx ↔ ort bidirectional)
Test coverage: Integration tests with tokio
**Gaps**: [G7] Session-only integration (import/run). No full model export path from rtx-graph IR.
---
### Layer 5: Specialized Domains
#### 2.5.1 Medical Imaging (8 crates)
**Status**: Production-Ready
Crates: `rtx-medical-core`, `rtx-medical-io`, `rtx-mri2fe`, `rtx-registration`, `rtx-segmentation`, `rtx-materials`, `rtx-fem-export`, `rtx-mesh-gen`
Test coverage: 165 `#[test]` in rtx-medical-core
Key capabilities: DICOM/NIfTI I/O, MRI→FEM pipeline for biomechanical simulation, image registration, segmentation, tissue material properties, FEM mesh export
Gaps: None detected.
---
#### 2.5.2 Neuroimaging — rtx-neuro Suite (17 crates)
**Status**: Production-Ready
Crates: rtx-neuro (super-crate), rtx-neuro-core (47K LOC consolidated), rtx-neuro-io, rtx-neuro-signal, rtx-neuro-forward, rtx-neuro-inverse, rtx-neuro-connectivity, rtx-neuro-stats, rtx-neuro-anatomy, rtx-neuro-db, rtx-neuro-lsl, rtx-neuro-realtime, rtx-neuro-artifacts, rtx-neuro-gnn, rtx-neuro-pinn, rtx-neuro-fem, rtx-neuro-python
Test coverage: 266 `#[test]` in rtx-neuro-core
Key capabilities: EEG/MEG I/O, Lab Streaming Layer (LSL) real-time, signal processing, forward/inverse modeling (source localization), GNN connectivity, artifact detection, PINN coupling, FEM coupling
**Gaps**:
- `rtx-neuro-python` requires separate maturin build — not in CI workflow
- Python bindings build is manual, not automated
---
#### 2.5.3 Physics-Informed & Scientific Computing (8 crates)
**Status**: Production-Ready
Crates: rtx-neural-operator, rtx-piddm, rtx-digital-twin, rtx-fea, rtx-cfd, rtx-science, rtx-geom, rtx-synthesis
Test coverage: rtx-fea has 371 `#[test]` in src/ and 176 in tests/; rtx-cfd has 179 in src/ and 142 in tests/
Key capabilities: FNO/DeepONet (mesh-agnostic PDE solving), Physics-Informed Diffusion, Digital Twins, Finite Element Analysis, CFD with GPU acceleration
**Gaps**: 4 `unimplemented!` in rtx-fea tests (test-only, not hot path).
---
#### 2.5.4 Other Specialized (6 crates)
**Status**: Production-Ready
Crates: rtx-compiler, rtx-platform, rtx-nmf, rtx-ml-classic, rtx-sklearn-py, rtx-polygraph
Gaps: None detected. rtx-sklearn-py bridges to scikit-learn via Python interop (subject to same PyO3 constraint as G1).
---
## 3. Cross-Cutting Concerns
### 3.1 Backend Op Completeness (G0 — Critical)
15 operations are `unimplemented!()` panics in **rtx-backend-cuda**, **rtx-backend-rocm**, and **rtx-backend-sycl**:
```
sin cos pow clamp relu
sigmoid tanh var var_dim leaky_relu
elu gt_scalar conv2d max_pool2d avg_pool2d
```
**Impact**: Any code using the generic `Backend` trait dispatch path will panic at runtime on these ops across all three backends. The `rtx-tensor` crate has its own cudarc/cuDNN implementations that bypass the trait, so it is unaffected. But consumers using the backend abstraction (e.g., `Fusion<CudaBackend>`, `Autodiff<CudaBackend>`) are exposed.
**Fix**: Implement these 15 ops in each of the three backends (45 total functions). Start with rtx-backend-cuda as the highest-value target, then mirror to ROCm/SYCL.
**Files**: `crates/core/rtx-backend-cuda/src/lib.rs`, `crates/core/rtx-backend-rocm/src/lib.rs`, `crates/core/rtx-backend-sycl/src/lib.rs`
---
### 3.2 Python Bindings Disabled (G1 — High)
**Affected crate**: `crates/core/rtx-bindings/`
**Current state**: `python` feature is disabled with comment "version conflicts". C API works.
**Impact**: Blocks Jupyter notebook use case. Blocks qstar Python integration. `rtx-neuro-python` and `rtx-sklearn-py` also affected.
**Fix**: Update to PyO3 ≥0.22.x (which stabilized the GIL API). Add `maturin build` step to CI. Unlock the `python` feature flag.
---
### 3.3 rtx-distributed Workspace Exclusion (G2 — High)
**Affected crate**: `crates/training/rtx-distributed/`
**Current state**: Manually excluded from `Cargo.toml` members due to RNCCL path dependency.
**Impact**: 439 tests not run in CI. Cannot use rtx-distributed as a dependency without manual workspace modification. Multi-GPU training is not CI-verified.
**Fix**: Make RNCCL an `optional` dependency with a CPU/mock fallback for CI. Conditional compilation: `#[cfg(feature = "nccl")]`.
---
### 3.4 Async Debt (G3 — Medium)
**Scope**: 277 `async fn` declarations across 97 files that never use `.await` internally.
**Affected crates**: rtx-graph, rtx-memory, rtx-tokenization, rtx-validation, rtx-nlg, rtx-config, rtx-monitoring, rtx-serving-api, rtx-science, rtx-evolution, rtx-federated, rtx-flash-attention, rtx-rl, rtx-transformers
**Current state**: Functions are unnecessarily async, adding allocator overhead for the poll machinery on every call.
**Fix**:
1. For functions that could be async (I/O, GPU sync): add the actual `.await` call points
2. For functions that have no async work: remove the `async` keyword
3. Tooling: `cargo clippy -- -W clippy::async_yields_async`
---
### 3.5 Deferred rtx-transformers Modules (G4 — Medium)
**Affected crate**: `crates/training/rtx-transformers/`
**Current state**: Phase 2 and Phase 3 modules are commented out with `// TODO: Re-enable` markers.
Phase 2 deferred: `validation_framework`, `meta-learning`, advanced tokenization
Phase 3 deferred: multi-modal text integration, continual learning
**Fix**: Re-enable one module at a time. Start with `validation_framework` (lowest risk, standalone). Each re-enablement needs its own integration test before merge.
---
### 3.6 rtx-tts Disabled (G5 — Medium)
**Affected crate**: `crates/models/rtx-tts/` (185 `#[test]` in src — significant investment)
**Current state**: Excluded from workspace. Depends on `rtx_nn` module API that has changed.
**Fix**: Audit current `rtx-nn` public API against `rtx-tts` imports; update `use` paths and type signatures to match. No architectural change needed.
---
### 3.7 Partial ONNX / DLPack Interop (G7 — Low)
**Affected crates**: `rtx-bindings`, `rtx-onnx`, `rtx-onnx-codegen`
**Current state**:
- ONNX: Session creation and inference work; `rtx-onnx-codegen` has a `todo!("Constant tensor")` — model export from rtx-graph is incomplete
- DLPack: Basic structure only; no zero-copy `__dlpack__` array interchange
**Fix**: Implement ONNX export using rtx-graph IR (graph walking → protobuf emission). Implement DLPack `DLManagedTensor` structs with `from_tensor`/`to_tensor` methods.
---
### 3.8 GPU Memory Access Tests (G6 — Low)
**Affected crate**: `crates/core/rtx-tensor/src/storage/`
**Current state**: 2 tests explicitly marked `// RED: GPU memory access is unimplemented` — intentional TDD red phase.
**Fix**: Implement `read_gpu_memory()` / `write_gpu_memory()` helpers in storage core; flip tests to green.
---
### 3.9 Quantum/Neuromorphic TODOs (G8 — Low)
**Affected crates**: `crates/models/rtx-timeseries/`, `crates/models/rtx-multimodal/`
**Current state**: `// TODO: Implement` stubs for quantum enhancement and neuromorphic preprocessing modules.
**Options**: Either implement (adds value for research use cases) or remove stubs and document as future work. Don't leave dead stubs in production code.
---
### 3.10 CI Coverage Gaps
- No maturin Python wheel build job in `.gitea/workflows/`
- No WASM build/size-check job (rtx-wasm-inference)
- rtx-distributed tests not run (see G2)
- GPU CI job runs but depends on self-hosted runner availability
---
## 4. Gaps Registry
| ID | Severity | Crate(s) | Description | Recommended Fix | Effort |
|----|----------|----------|-------------|-----------------|--------|
| G0 | **Critical** | rtx-backend-cuda, rtx-backend-rocm, rtx-backend-sycl | 15 ops are `unimplemented!()` panics per backend (45 total): sin, cos, pow, clamp, relu, sigmoid, tanh, var, var_dim, leaky_relu, elu, gt_scalar, conv2d, max_pool2d, avg_pool2d | Implement all 15 in each backend; start with CUDA | M |
| G1 | **High** | rtx-bindings, rtx-neuro-python, rtx-sklearn-py | PyO3 Python bindings disabled (version conflict); blocks Jupyter/qstar integration | Update to PyO3 ≥0.22, add maturin CI step | M |
| G2 | **High** | rtx-distributed | Excluded from workspace due to RNCCL path dep; 439 tests unrun; multi-GPU CI broken | Make RNCCL optional `#[cfg(feature = "nccl")]`, add CPU/mock fallback | L |
| G3 | **Medium** | 97 files | 277 async fns without `.await` — unnecessary allocation overhead per call | Batch sweep: remove async or add real await points; use clippy lint | M |
| G4 | **Medium** | rtx-transformers | Phase 2-3 modules commented out (validation_framework, meta-learning, multi-modal text, continual learning) | Re-enable incrementally, one module per PR with integration test | M |
| G5 | **Medium** | rtx-tts | Disabled: rtx_nn API surface changed after rtx-tts was written | Audit rtx-nn current API; update import paths and type sigs | S |
| G6 | **Low** | rtx-tensor/storage | 2 GPU memory access tests in TDD red phase | Implement read/write GPU memory helpers | S |
| G7 | **Low** | rtx-bindings, rtx-onnx, rtx-onnx-codegen | ONNX model export incomplete (`todo!("Constant tensor")`); DLPack zero-copy not implemented | Full ONNX export via rtx-graph IR; DLPack `DLManagedTensor` | L |
| G8 | **Low** | rtx-timeseries, rtx-multimodal | Quantum/neuromorphic TODO stubs in non-hot-path modules | Implement or remove — no dead stubs in production | S-M |
Effort key: **S** = days, **M** = 1-2 weeks, **L** = 3-4 weeks
---
## 5. Prioritized Roadmap
### Phase 1 — Critical & Blocking Fixes (Weeks 1-2)
**Goal**: Eliminate runtime panics in the backend dispatch path; restore CI integrity; re-enable Python ecosystem.
| # | Action | Gap | Owner crates | Verification |
|---|--------|-----|-------------|--------------|
| 1 | Implement 15 missing backend ops in rtx-backend-cuda | G0 | rtx-backend-cuda | All 15 ops covered by tests; `cargo test -F cuda` passes |
| 2 | Mirror same 15 ops to rtx-backend-rocm and rtx-backend-sycl | G0 | rtx-backend-rocm, rtx-backend-sycl | Feature-gated tests pass |
| 3 | Update PyO3 to ≥0.22, re-enable `python` feature in rtx-bindings | G1 | rtx-bindings | `cargo test -F python`; `maturin build` succeeds |
| 4 | Add maturin build job to `.gitea/workflows/ci.yml` | G1 | CI | Wheel builds on push to main |
| 5 | Make RNCCL optional in rtx-distributed; re-add to workspace | G2 | rtx-distributed | `cargo test` (CPU path) runs in CI; all 439 tests pass |
| 6 | Fix rtx-tts API mismatch with rtx-nn | G5 | rtx-tts | `cargo test` for rtx-tts passes |
---
### Phase 2 — Technical Debt Reduction (Weeks 3-4)
**Goal**: Eliminate async overhead, restore deferred modules, complete TDD red phases.
| # | Action | Gap | Verification |
|---|--------|-----|--------------|
| 7 | Async sweep — 97 files | G3 | `cargo clippy -- -W clippy::async_yields_async` returns zero warnings |
| 8 | Re-enable rtx-transformers `validation_framework` module | G4 | Integration test passes |
| 9 | Re-enable rtx-transformers `meta-learning` module | G4 | Integration test passes |
| 10 | Implement GPU memory access helpers in rtx-tensor/storage | G6 | 2 RED tests flip to green |
| 11 | Resolve quantum/neuromorphic TODOs in rtx-timeseries and rtx-multimodal (implement or remove) | G8 | No `// TODO` stubs in non-test code |
---
### Phase 3 — 20-Demo Push (Weeks 5-10)
**Goal**: Complete the 20 next-generation demo applications per `resume_nextplan.md`.
**Phase 3A — Medical (Weeks 5-6)**
- DrugBinder: molecular docking with GNN + RL
- CardioSim: cardiac simulation with FEM coupling
- TumorBoard AI: multi-modal oncology decision support
**Phase 3B — Finance (Weeks 6-7)**
- QuantumPort: portfolio optimization with quantum-inspired algorithms
- MarketSim: market simulation with agent-based modeling
- RiskFlow: risk factor analysis with normalizing flows
- AlgoArena: algorithmic trading benchmarking
**Phase 3C — AI/ML Core (Weeks 7-8)**
- NeuralOp Studio: interactive FNO/DeepONet playground
- WorldGen: video diffusion for world model training
- EmbodiedSim: robotics simulation with physics coupling
- FoundationForge: foundation model training harness
**Phase 3D — Engineering (Weeks 8-9)**
- AeroFlow: aerodynamics CFD with GPU acceleration
- WeatherCast: NWP with neural operators
- SeismicAI: seismic inversion with PINNs
- StructuralPINN: structural mechanics with physics-informed networks
**Phase 3E — Cluster (Weeks 9-10)**
- ClusterViz: distributed training monitoring dashboard
- DistributedLLM: multi-node LLM training benchmark
- FederatedMed: federated medical imaging across hospital networks
---
### Phase 4 — Ecosystem Maturity (Weeks 11-12)
**Goal**: Production-grade interoperability, documentation, and CI completeness.
| # | Action | Gap | Verification |
|---|--------|-----|--------------|
| 12 | Full ONNX model export: rtx-graph IR → protobuf emission | G7 | Round-trip test: train rtx model, export ONNX, run via ort session |
| 13 | DLPack zero-copy: implement `DLManagedTensor` structs | G7 | Array interchange test with numpy |
| 14 | Re-enable rtx-transformers Phase 3 modules (multi-modal text, continual learning) | G4 | Integration tests pass |
| 15 | Add WASM build + size-check CI job for rtx-wasm-inference | — | WASM bundle <5MB (gzipped); `wasm-pack test` passes |
| 16 | mdBook: complete all 8 layer chapters with API examples | — | `mdbook build` succeeds; no broken links |
| 17 | rtx-neuro-python: add maturin build to CI | G1 | Python wheel installable; EEG demo runs |
---
## Appendix A: Test Density by Crate
Crates with the highest test investment (top 15 by `#[test]` count):
| Rank | Crate | `#[test]` count |
|------|-------|-----------------|
| 1 | rtx-transformers | 1,845 |
| 2 | rtx-tensor | 1,047+ (src + tests) |
| 3 | rtx-distributed | 439 |
| 4 | rtx-fea | 371 (src) + 176 (tests) |
| 5 | rtx-nn | 334 |
| 6 | rtx-neuro-core | 266 |
| 7 | rtx-autograd | 223 (tests) + 128 (src) |
| 8 | rtx-metal | 201 |
| 9 | rtx-runtime | 192 (src) + 190 (tests) |
| 10 | rtx-memory | 190 |
| 11 | rtx-tts | 185 (currently excluded) |
| 12 | rtx-preprocessing | 184 (tests) |
| 13 | rtx-tensor (tests/) | 184 |
| 14 | rtx-interpret | 183 (src) + 145 (tests) |
| 15 | rtx-cfd | 179 (src) + 142 (tests) |
---
## Appendix B: Unimplemented Stub Inventory
All 51 `todo!/unimplemented!` occurrences:
| Count | Location | Operations |
|-------|----------|------------|
| 15 | `rtx-backend-cuda/src/lib.rs` | sin, cos, pow, clamp, relu, sigmoid, tanh, var, var_dim, leaky_relu, elu, gt_scalar, conv2d, max_pool2d, avg_pool2d |
| 15 | `rtx-backend-rocm/src/lib.rs` | (same 15 ops) |
| 15 | `rtx-backend-sycl/src/lib.rs` | (same 15 ops) |
| 4 | `rtx-fea/tests/` | Test-only stubs (not hot path) |
| 1 | `rtx-transformers/src/layers/geglu.rs` | Numerical gradient checking (TDD red phase) |
| 1 | `rtx-onnx-codegen/src/ops/mod.rs` | `todo!("Constant tensor")` in ONNX codegen |
**Total**: 51 — all non-test stubs are in backend op implementations or ONNX codegen.
---
## Appendix C: CI Workflow Coverage
| Workflow | Triggers | What it checks |
|----------|----------|----------------|
| ci.yml | Push/PR | rustfmt, clippy (pedantic), build Linux+macOS, cargo test |
| gpu-tests.yml | Push to main | CUDA kernel tests (self-hosted GPU runner) |
| benchmarks.yml | Manual / release | Criterion benchmarks, performance regression |
| docs.yml | Push to main | mdbook build, doc links |
| release.yml | Tag push | Build + publish crates |
**Gap**: No WASM build job. No maturin/PyO3 wheel build job. rtx-distributed excluded from all jobs.
---
*Spec self-review: no TBD/TODO placeholders; all 8 gaps in registry; all 5 layers covered; roadmap has 17 concrete action items across 4 phases.*