Comprehensive layered architectural review covering all 109 crates, 306K LOC, 13,351 tests. Identifies 9 gaps (G0-G8) with the highest- priority being 45 unimplemented! panics across rtx-backend-cuda/rocm/sycl and the rtx-distributed workspace exclusion. Includes 17-item 4-phase roadmap through 90 days. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
33 KiB
RustyTorch Full Repository Review
Date: 2026-06-26
Crates: 109
LOC: 306,423
Tests: 13,351+
Incomplete stubs: 51
Toolchain: nightly-2025-10-25 (Rust 2024 edition)
1. Project Overview
RustyTorch is a production-grade, GPU-accelerated ML framework in pure Rust. It aims to match PyTorch's feature breadth while eliminating the Python GIL, delivering memory-safe tensor operations, automatic differentiation, and domain-specific stacks for medical imaging, neuroimaging, and scientific computing.
1.1 Platform Role
Layer 8 — ML Infrastructure for the QuantumRedClaw platform:
- QPUDIDP consumes RustyTorch for surrogate model training (MLP, normalizing flows, MC Dropout)
- qstar neural decoders depend on GPU compute provided by this framework
- Python interop via PyO3/maturin bridges to Jupyter notebooks
1.2 Maturity Snapshot
| Metric | Value |
|---|---|
| Total crates | 109 |
| Total Rust LOC | 306,423 |
#[test] attributes |
13,351+ |
todo!/unimplemented! |
51 |
| CI pipelines | 5 (format, clippy, build, test, GPU) |
| Demo applications | 30+ |
| Rust toolchain | nightly-2025-10-25 |
| Edition | 2024 |
| Build target dir | /hot/targets/rustyverse/rustytorch |
| TIER progress | TIER 0–3 complete (Dec 28, 2025) |
1.3 Key Differentiators vs. Ecosystem
| Feature | RustyTorch | tch-rs | burn | candle |
|---|---|---|---|---|
| Pure Rust (no C/C++) | ✅ | ❌ (libtorch) | ✅ | ✅ |
| Multi-backend dispatch | ✅ | ❌ | ✅ | partial |
| Flash Attention | ✅ | via ext | ❌ | ❌ |
| Mixture of Experts | ✅ | ❌ | ❌ | ❌ |
| Mamba/SSM | ✅ | ❌ | ❌ | ❌ |
| Speculative decoding | ✅ | ❌ | ❌ | ❌ |
| Medical imaging stack | ✅ | ❌ | ❌ | ❌ |
| Neuroimaging (EEG/MEG) | ✅ | ❌ | ❌ | ❌ |
| Neural Operators (FNO) | ✅ | ❌ | ❌ | ❌ |
| Federated learning | ✅ | ❌ | ❌ | ❌ |
1.4 GPU Backend Support
| Backend | Hardware | Status |
|---|---|---|
| CUDA (sm_86 default, sm_120 Blackwell) | NVIDIA RTX/A100/H100/5060 Ti | Production |
| Metal (MPS) | Apple M-series | Production |
| ROCm/HIP | AMD RDNA 2/3 | Partial (15 ops unimplemented) |
| SYCL/DPC++ | Intel Arc | Partial (15 ops unimplemented) |
| CubeCL (WGPU) | Cross-platform | Production |
| CPU (OpenBLAS/MKL/manual) | x86/ARM | Production |
2. Layer-by-Layer Review
Layer 1: Core Infrastructure
2.1.1 rtx-backend
Status: Production-Ready
Purpose: Compile-time backend dispatch trait — zero runtime overhead
Path: crates/core/rtx-backend/
Key API surface (522 LOC, src/lib.rs):
Backendtrait: 430+ methods covering tensor creation, element-wise ops, reductions, LLM-specific ops (flash_attention, rope, rms_norm)AutodiffBackend: marker trait for gradient-capable backendsQuantizedBackend: quantization support
Test coverage: Feature-gated; integration via backend crates
Gaps: Trait is intentionally abstract. The gap is in the concrete backends (see rtx-backend-cuda, rtx-backend-rocm, rtx-backend-sycl below).
2.1.2 rtx-tensor
Status: Production-Ready
Purpose: PyTorch-compatible GPU-native tensor primitives
Path: crates/core/rtx-tensor/
Key API surface:
Tensor,Device,DType,ShapeComplexTensor,SparseCOO,SparseCSRCublasManager,CudnnContext,CuSolverContext,CuSparseContext- 35+ tensor operation modules, 7 Metal-specific modules
Test coverage: 1,047 #[test] in src, 184 in tests/; loom concurrency tests; 10+ benchmark files
Notable: Real CUDA kernel compilation via cudarc 0.18.x, real MSL shaders for Metal
Gaps: 2 GPU memory access tests marked RED: unimplemented in storage/core/ (flagged as intentional TDD red phase)
2.1.3 rtx-autograd
Status: Production-Ready
Purpose: Tape-based automatic differentiation, forward and reverse modes
Path: crates/core/rtx-autograd/
Key API surface:
Autodiff<B>: decorator wrapping anyBackend<B>to add gradientsgrad(),jvp(),vjp(),jacfwd(),jacrev(),hessian(),vmap()- Checkpointing:
checkpoint(),EveryNthStrategy,SqrtCheckpointStrategy,AdaptiveCheckpointStrategy AutogradProfilerwith Chrome trace export
Test coverage: 223 #[test] in tests/, 128 in src/; gradient_benchmarks.rs
Gaps: None detected.
2.1.4 rtx-memory
Status: Production-Ready
Purpose: GPU memory pooling, KV cache, ZeRO-style partitioning
Path: crates/core/rtx-memory/
Key API surface:
MemoryPoolManager,TransformerMemoryPoolsKvCacheManager,AsyncKvCacheManager,EntropyTrackerZeroMemoryManager,PagedAttentionManager- Allocators: Buddy, Slab, Stack
Test coverage: 190 #[test] in src/, 190 in tests/; 3 benchmark files (pool, allocation strategies, simple GPU)
Gaps: None detected. Fragmentation target <5% is documented.
2.1.5 rtx-kernel
Status: Production-Ready
Purpose: CUDA PTX + Metal MSL kernel compilation, caching, execution
Path: crates/core/rtx-kernel/
Key API surface:
KernelSystem,KernelCompiler(NVRTC-based),KernelCache,KernelExecutorMetalKernelCompiler,MetalKernelCache,MetalKernelExecutorLaunchConfig,KernelParam(type-safe binding)
Gaps: None detected.
2.1.6 rtx-flash-attention
Status: Production-Ready (with one stub in backward pass)
Purpose: Memory-efficient attention, O(n) memory, 5-8× speedup
Path: crates/training/rtx-flash-attention/
Key API surface:
FlashAttention,FlashAttentionFactory,FlashAttentionConfigSdpaBackendSelector(hardware-aware: FlashAttn2, SDPA, naive fallback)flash_attention_forward(),flash_attention_backward()
Test coverage: 97 #[test] in src/; flash_attention_bench.rs
Notable: sm_120 (Blackwell/RTX 5060 Ti) support added in commit c2f4796
Gaps: Backward pass gradient computation has a placeholder fallback path for non-CUDA; acceptable since CUDA is the target.
2.1.7 rtx-lora
Status: Production-Ready
Purpose: LoRA and QLoRA parameter-efficient fine-tuning
Path: crates/core/rtx-lora/
Key API surface:
LoraConfig,LoraAdapter,LoraWeight,QLoraWeightNF4Block(4-bit quantization),AdapterManager,MergeStrategy
Test coverage: Tests in adapter.rs, config.rs, weight.rs, quantized.rs, merge.rs
Gaps: None detected.
2.1.8 rtx-moe
Status: Production-Ready
Purpose: Mixture of Experts (SwitchTransformer, ExpertChoice, TokenChoice routing)
Path: crates/core/rtx-moe/
Key API surface: SwitchTransformer, ExpertChoice, TokenChoice routing strategies; load balancing; expert capacity controls
Gaps: None detected from exploration. (Dedicated Metal MSL kernels in rtx-tensor/metal_moe_ops.)
2.1.9 rtx-mamba
Status: Production-Ready (conv1d history buffer improved in D309)
Purpose: Mamba/State Space Models, O(n) attention alternative
Path: crates/core/rtx-mamba/
Key API surface: MambaRecurrence, selective scan, MambaStep (new ergonomic step API from commit f751414)
Gaps: None detected.
2.1.10 rtx-speculative-decoding
Status: Production-Ready
Purpose: 2-3× inference speedup via draft model + target verification
Path: crates/core/rtx-speculative-decoding/ and crates/production/rtx-inference/
Key API surface: SpeculativeEngine, AdvancedSpeculativeDecoder, EagleDraftModel, MedusaDraftModel, tree-based decoding
Metal ops: metal_speculative_ops in rtx-tensor
Gaps: None detected.
2.1.11 rtx-fusion
Status: Production-Ready
Purpose: Kernel fusion DSL eliminating intermediate tensor writes (10-25% speedup)
Path: crates/core/rtx-fusion/
Key API surface:
Fusion<B>backend wrapperFusionAnalyzer(pattern detection),FusedKernelCache(LRU),FusionConfig,FusionStats- Sync points: matmul, reductions, flash_attention, data access
Gaps: None detected.
2.1.12 rtx-cubecl
Status: Production-Ready
Purpose: Single-source Rust → GPU kernel synthesis via CubeCL
Path: crates/core/rtx-cubecl/
Key API surface: CubeclBackend, CubeclClient, CubeclDevice, CubeclTensor
Backends: cpu, wgpu, cuda, hip, vulkan, metal
Tier strategy: ~60% ops via portable CubeCL kernels; ~40% via hand-crafted tier-2 (Flash Attention, GEMM, conv)
Gaps: None detected.
2.1.13 rtx-bindings
Status: Partial
Purpose: Python (PyO3), C FFI, ONNX, DLPack interop
Path: crates/core/rtx-bindings/
Key API surface: c_api (enabled), python (disabled), onnx (basic), dlpack (basic)
Test coverage: test files exist but are conditionally disabled
Gaps:
- [G1] PyO3 Python bindings disabled due to version conflicts
- [G7] ONNX/DLPack are basic stubs — no full model export or zero-copy interchange
2.1.14 rtx-macros
Status: Production-Ready
Purpose: Derive macros: #[derive(Module)], #[derive(Config)]
Path: crates/core/rtx-macros/
Key API surface:
Modulederive: generatesparameters(),to_device(),train(), field attrs:#[param],#[module],#[constant]Configderive: builder pattern with#[config(default = "value")]
Gaps: None detected.
2.1.15 rtx-backend-cuda / rtx-backend-rocm / rtx-backend-sycl
Status: Partial — Critical gap
Path: crates/core/rtx-backend-{cuda,rocm,sycl}/src/lib.rs
Each of these three backend adapters is missing the same 15 operations:
sin, cos, pow, clamp, relu, sigmoid, tanh,
var, var_dim, leaky_relu, elu, gt_scalar,
conv2d, max_pool2d, avg_pool2d
These are unimplemented!() panics — any code path that routes through the backend dispatch layer for these ops will panic at runtime on CUDA, ROCm, and SYCL.
Note: rtx-tensor has its own GPU implementations (via cudarc + cuDNN directly) that do not go through the backend dispatch path. But any consumer using the generic Backend trait abstraction hits these panics.
This is gap G0 — the highest-priority fix in this review.
Layer 2: Training & Optimization
2.2.1 rtx-transformers
Status: Production-Ready (core); Phase 2-3 modules deferred
Path: crates/training/rtx-transformers/
Key API surface:
- Layers: Mamba, MoE, RoPE, Linear Attention, Spectral Norm, Block Sparse Attention, GeLU, SwiGLU, GatedMLP
- Architectures: BERT, GPT, TransformerBlock, MultiHeadAttention
- Optimizers: Adam, AdamW, Shampoo
- New (commit
24bd5cf):mamba_step,ClonedMemoryUpdater,GatedMemoryUpdater,SetEncoderTeacher
Test coverage: 1,845 #[test] — highest in the workspace
Notable: Also has integration tests at tests/d350_gpu_backbone_training.rs, tests/mamba_temporal_learning.rs, tests/real_selective_scan.rs
Gaps: [G4] Phase 2 modules commented out: validation_framework, meta-learning, advanced tokenization, multi-modal text, continual learning
2.2.2 rtx-distributed
Status: Production-Ready (code) — excluded from workspace
Path: crates/training/rtx-distributed/
Key API surface:
- Collective ops: AllReduce, Broadcast, Gather, Scatter via NCCL/RCCL/MPI
- Parallelism:
DistributedDataParallel,FSDP2,TensorParallel,PipelineParallel - Elastic training:
ElasticCluster,FaultDetector,ActivationCheckpointManager - DCP: Distributed Checkpoint with
DeviceMesh,HardwareTopology
Test coverage: 439 #[test]
Gaps: [G2] Excluded from workspace Cargo.toml due to RNCCL path dependency — cannot be tested in CI, cannot be used as a dependency without manual workspace inclusion.
2.2.3 rtx-rl
Status: Production-Ready
Path: crates/training/rtx-rl/
Key API surface: PPO, DQN, SAC, ActorLearner, Environment, ReplayBuffer, RLHF
Test coverage: 6 test files (PPO, SAC, RLHF, actor-learner, environment, replay buffer); rl_bench benchmark
Gaps: None detected.
2.2.4 rtx-compress
Status: Production-Ready
Path: crates/training/rtx-compress/
Key API surface: CompressionPipeline, CompressionConfig, CompressedStorage, quantization, pruning, distillation, KV cache, LoRA
Test coverage: 164 #[test] in tests/; compression_bench
Gaps: Some advanced quantization modes (AWQ, GPTQ, SmoothQuant) are in progress — basic pipeline is complete.
2.2.5 rtx-nas
Status: Production-Ready
Path: crates/training/rtx-nas/
Key API surface: DARTS, PC-DARTS, FairNAS, RandomSearch, Architecture, DARTSCell, SearchSpace
Test coverage: 165 #[test]
Gaps: None detected.
2.2.6 rtx-federated
Status: Production-Ready
Path: crates/training/rtx-federated/
Key API surface:
- Aggregation: FedAvg, FedProx, SCAFFOLD, FedNova, AsyncAggregation
- Byzantine-robust: Krum, MultiKrum, TrimmedMean, AnomalyDetector, ReputationSystem
- Privacy: DifferentialPrivacy, HomomorphicEncryption (tfhe), SMPC, PrivacyBudget
- Personalization: MetaLearning, ClientClustering, MultiTaskLearning, TransferLearning
Test coverage: federated_bench; federated_demo example
Gaps: None detected. 80+ dependencies including cryptographic primitives.
2.2.7 rtx-auto
Status: Production-Ready
Path: crates/training/rtx-auto/
Key API surface: AutonomousAgent, DataEngineeringAgent, ParallelPlannerAgent, QuantGuardianAgent, KernelSynthesizerAgent, ProposalValidator, RollbackManager
Test coverage: Integration test for AutonomousOptimizer initialization
Gaps: None detected.
2.2.8 rtx-evolution
Status: Production-Ready
Path: crates/training/rtx-evolution/
Key API surface: EvolutionOrchestrator, TelemetryAnalyzer, SafeSandbox, KnowledgeGraph, MultiObjectiveOptimizer, ParetoFrontier, HyperparameterTuner
Test coverage: orchestrator_tests.rs
Gaps: None detected.
Layer 3: Model Architectures
2.3.1 rtx-vision
Status: Production-Ready
Path: crates/models/rtx-vision/
Key API surface: ViT, ConvNeXt V2, EfficientNetV2, MobileViT, DenseNet, NFNet, EdgeViT, MaxViT, CoAtNet (150+ exported types)
Preprocessing: ImageTensor, Augmentation, ImageProcessor
Gaps: None detected.
2.3.2 rtx-multimodal
Status: Production-Ready (quantum/neuromorphic stubs remain)
Path: crates/models/rtx-multimodal/
Key API surface: MultimodalSystem, VisionTransformer, AudioTransformer, CrossModalAttention, ModalityFusion (7 strategies)
Test coverage: 122 #[test]; multimodal_bench
Gaps: [G8] Quantum enhancement and neuromorphic preprocessing modules have TODO stubs.
2.3.3 rtx-diffuse
Status: Production-Ready
Path: crates/models/rtx-diffuse/
Key API surface: UNet, DiT (Diffusion Transformer), DDIMSampler, DPMSolverPP, UniPCSampler (5-10 step fast inference), ControlNet, IP-Adapter, T2I-Adapter, ClassifierFreeGuidance, LoRA diffusion
Test coverage: 141 #[test] in src/; validation tests; diffusion_bench
Gaps: None detected.
2.3.4 rtx-timeseries
Status: Production-Ready (quantum stubs remain)
Path: crates/models/rtx-timeseries/
Key API surface: ARIMAModel, SARIMAModel, ProphetModel, NeuralProphetModel, ExponentialSmoothingModel, StateSpaceModel, TransformerForecastModel, Forecaster, TimeSeriesAnalyzer (seasonality, stationarity, anomaly detection)
Test coverage: Criterion benchmarks
Gaps: [G8] Quantum forecasting modules marked // TODO: Implement.
2.3.5 rtx-nlg
Status: Production-Ready
Path: crates/models/rtx-nlg/
Key API surface: TextGenerator, ModelServer, StreamingGenerator, BatchGenerator, ConversationManager, BeamSearch, NucleusSampling, RepetitionPenalty, ToxicityFilter
Test coverage: 3 benchmark suites (generation, translation, summarization)
Gaps: None detected.
Layer 4: Production & Deployment
2.4.1 rtx-serving-api
Status: Production-Ready
Path: crates/production/rtx-serving-api/
Key API surface: ServingServer, AdvancedServingServer, ContinuousBatchingController, QueueManager, RateLimitManager, CircuitBreaker, GrammarSampler (BNF parser), StructuredGenerator (JSON schema), MultiModelManager, WebSocket streaming, BudgetConfig
Stack: Axum (HTTP), Tonic (gRPC), tokio-tungstenite (WebSocket), Prometheus
Test coverage: 140 #[test] in src/; cache_performance benchmark
Gaps: None detected.
2.4.2 rtx-inference
Status: Production-Ready
Path: crates/production/rtx-inference/
Key API surface: InferenceEngine, PagedKvCache (GPU→CPU→NVMe tiering), BatchScheduler (SLA lanes), AdvancedSpeculativeDecoder, EagleDraftModel, MedusaDraftModel, QuantizationConfig, MetricsCollector
Test coverage: Integration tests with proptest
Features: onnx-runtime, burn, burn-wgpu, candle, candle-cuda, candle-metal
Gaps: None detected.
2.4.3 rtx-streaming
Status: Production-Ready
Path: crates/production/rtx-streaming/
Key API surface: StreamingServer, ConnectionManager, TokenGenerator, BackpressureHandler, AdaptiveProcessor, Kafka integration, edge computing support
Design target: sub-millisecond latency, >1000 concurrent connections
Test coverage: streaming_bench
Gaps: None detected.
2.4.4 rtx-wasm-inference
Status: Production-Ready
Path: crates/production/rtx-wasm-inference/
Key API surface: WasmInferenceEngine, ComputeBackend (Cpu/WebGpu/Auto), TensorCore, DType
Features: WASM SIMD, SharedArrayBuffer threads, WebGPU acceleration, TypeScript/JavaScript bindings
Build: opt-level = "s", LTO enabled, wasm-opt with SIMD
Gaps: None detected.
2.4.5 rtx-onnx
Status: Partial
Path: crates/production/rtx-onnx/
Key API surface: OnnxSession, OnnxSessionConfig, ExecutionProviderType (CPU/CUDA/CoreML/TensorRT/DirectML), tensor_bridge (rtx ↔ ort bidirectional)
Test coverage: Integration tests with tokio
Gaps: [G7] Session-only integration (import/run). No full model export path from rtx-graph IR.
Layer 5: Specialized Domains
2.5.1 Medical Imaging (8 crates)
Status: Production-Ready
Crates: rtx-medical-core, rtx-medical-io, rtx-mri2fe, rtx-registration, rtx-segmentation, rtx-materials, rtx-fem-export, rtx-mesh-gen
Test coverage: 165 #[test] in rtx-medical-core
Key capabilities: DICOM/NIfTI I/O, MRI→FEM pipeline for biomechanical simulation, image registration, segmentation, tissue material properties, FEM mesh export
Gaps: None detected.
2.5.2 Neuroimaging — rtx-neuro Suite (17 crates)
Status: Production-Ready
Crates: rtx-neuro (super-crate), rtx-neuro-core (47K LOC consolidated), rtx-neuro-io, rtx-neuro-signal, rtx-neuro-forward, rtx-neuro-inverse, rtx-neuro-connectivity, rtx-neuro-stats, rtx-neuro-anatomy, rtx-neuro-db, rtx-neuro-lsl, rtx-neuro-realtime, rtx-neuro-artifacts, rtx-neuro-gnn, rtx-neuro-pinn, rtx-neuro-fem, rtx-neuro-python
Test coverage: 266 #[test] in rtx-neuro-core
Key capabilities: EEG/MEG I/O, Lab Streaming Layer (LSL) real-time, signal processing, forward/inverse modeling (source localization), GNN connectivity, artifact detection, PINN coupling, FEM coupling
Gaps:
rtx-neuro-pythonrequires separate maturin build — not in CI workflow- Python bindings build is manual, not automated
2.5.3 Physics-Informed & Scientific Computing (8 crates)
Status: Production-Ready
Crates: rtx-neural-operator, rtx-piddm, rtx-digital-twin, rtx-fea, rtx-cfd, rtx-science, rtx-geom, rtx-synthesis
Test coverage: rtx-fea has 371 #[test] in src/ and 176 in tests/; rtx-cfd has 179 in src/ and 142 in tests/
Key capabilities: FNO/DeepONet (mesh-agnostic PDE solving), Physics-Informed Diffusion, Digital Twins, Finite Element Analysis, CFD with GPU acceleration
Gaps: 4 unimplemented! in rtx-fea tests (test-only, not hot path).
2.5.4 Other Specialized (6 crates)
Status: Production-Ready
Crates: rtx-compiler, rtx-platform, rtx-nmf, rtx-ml-classic, rtx-sklearn-py, rtx-polygraph
Gaps: None detected. rtx-sklearn-py bridges to scikit-learn via Python interop (subject to same PyO3 constraint as G1).
3. Cross-Cutting Concerns
3.1 Backend Op Completeness (G0 — Critical)
15 operations are unimplemented!() panics in rtx-backend-cuda, rtx-backend-rocm, and rtx-backend-sycl:
sin cos pow clamp relu
sigmoid tanh var var_dim leaky_relu
elu gt_scalar conv2d max_pool2d avg_pool2d
Impact: Any code using the generic Backend trait dispatch path will panic at runtime on these ops across all three backends. The rtx-tensor crate has its own cudarc/cuDNN implementations that bypass the trait, so it is unaffected. But consumers using the backend abstraction (e.g., Fusion<CudaBackend>, Autodiff<CudaBackend>) are exposed.
Fix: Implement these 15 ops in each of the three backends (45 total functions). Start with rtx-backend-cuda as the highest-value target, then mirror to ROCm/SYCL.
Files: crates/core/rtx-backend-cuda/src/lib.rs, crates/core/rtx-backend-rocm/src/lib.rs, crates/core/rtx-backend-sycl/src/lib.rs
3.2 Python Bindings Disabled (G1 — High)
Affected crate: crates/core/rtx-bindings/
Current state: python feature is disabled with comment "version conflicts". C API works.
Impact: Blocks Jupyter notebook use case. Blocks qstar Python integration. rtx-neuro-python and rtx-sklearn-py also affected.
Fix: Update to PyO3 ≥0.22.x (which stabilized the GIL API). Add maturin build step to CI. Unlock the python feature flag.
3.3 rtx-distributed Workspace Exclusion (G2 — High)
Affected crate: crates/training/rtx-distributed/
Current state: Manually excluded from Cargo.toml members due to RNCCL path dependency.
Impact: 439 tests not run in CI. Cannot use rtx-distributed as a dependency without manual workspace modification. Multi-GPU training is not CI-verified.
Fix: Make RNCCL an optional dependency with a CPU/mock fallback for CI. Conditional compilation: #[cfg(feature = "nccl")].
3.4 Async Debt (G3 — Medium)
Scope: 277 async fn declarations across 97 files that never use .await internally.
Affected crates: rtx-graph, rtx-memory, rtx-tokenization, rtx-validation, rtx-nlg, rtx-config, rtx-monitoring, rtx-serving-api, rtx-science, rtx-evolution, rtx-federated, rtx-flash-attention, rtx-rl, rtx-transformers
Current state: Functions are unnecessarily async, adding allocator overhead for the poll machinery on every call.
Fix:
- For functions that could be async (I/O, GPU sync): add the actual
.awaitcall points - For functions that have no async work: remove the
asynckeyword - Tooling:
cargo clippy -- -W clippy::async_yields_async
3.5 Deferred rtx-transformers Modules (G4 — Medium)
Affected crate: crates/training/rtx-transformers/
Current state: Phase 2 and Phase 3 modules are commented out with // TODO: Re-enable markers.
Phase 2 deferred: validation_framework, meta-learning, advanced tokenization
Phase 3 deferred: multi-modal text integration, continual learning
Fix: Re-enable one module at a time. Start with validation_framework (lowest risk, standalone). Each re-enablement needs its own integration test before merge.
3.6 rtx-tts Disabled (G5 — Medium)
Affected crate: crates/models/rtx-tts/ (185 #[test] in src — significant investment)
Current state: Excluded from workspace. Depends on rtx_nn module API that has changed.
Fix: Audit current rtx-nn public API against rtx-tts imports; update use paths and type signatures to match. No architectural change needed.
3.7 Partial ONNX / DLPack Interop (G7 — Low)
Affected crates: rtx-bindings, rtx-onnx, rtx-onnx-codegen
Current state:
- ONNX: Session creation and inference work;
rtx-onnx-codegenhas atodo!("Constant tensor")— model export from rtx-graph is incomplete - DLPack: Basic structure only; no zero-copy
__dlpack__array interchange Fix: Implement ONNX export using rtx-graph IR (graph walking → protobuf emission). Implement DLPackDLManagedTensorstructs withfrom_tensor/to_tensormethods.
3.8 GPU Memory Access Tests (G6 — Low)
Affected crate: crates/core/rtx-tensor/src/storage/
Current state: 2 tests explicitly marked // RED: GPU memory access is unimplemented — intentional TDD red phase.
Fix: Implement read_gpu_memory() / write_gpu_memory() helpers in storage core; flip tests to green.
3.9 Quantum/Neuromorphic TODOs (G8 — Low)
Affected crates: crates/models/rtx-timeseries/, crates/models/rtx-multimodal/
Current state: // TODO: Implement stubs for quantum enhancement and neuromorphic preprocessing modules.
Options: Either implement (adds value for research use cases) or remove stubs and document as future work. Don't leave dead stubs in production code.
3.10 CI Coverage Gaps
- No maturin Python wheel build job in
.gitea/workflows/ - No WASM build/size-check job (rtx-wasm-inference)
- rtx-distributed tests not run (see G2)
- GPU CI job runs but depends on self-hosted runner availability
4. Gaps Registry
| ID | Severity | Crate(s) | Description | Recommended Fix | Effort |
|---|---|---|---|---|---|
| G0 | Critical | rtx-backend-cuda, rtx-backend-rocm, rtx-backend-sycl | 15 ops are unimplemented!() panics per backend (45 total): sin, cos, pow, clamp, relu, sigmoid, tanh, var, var_dim, leaky_relu, elu, gt_scalar, conv2d, max_pool2d, avg_pool2d |
Implement all 15 in each backend; start with CUDA | M |
| G1 | High | rtx-bindings, rtx-neuro-python, rtx-sklearn-py | PyO3 Python bindings disabled (version conflict); blocks Jupyter/qstar integration | Update to PyO3 ≥0.22, add maturin CI step | M |
| G2 | High | rtx-distributed | Excluded from workspace due to RNCCL path dep; 439 tests unrun; multi-GPU CI broken | Make RNCCL optional #[cfg(feature = "nccl")], add CPU/mock fallback |
L |
| G3 | Medium | 97 files | 277 async fns without .await — unnecessary allocation overhead per call |
Batch sweep: remove async or add real await points; use clippy lint | M |
| G4 | Medium | rtx-transformers | Phase 2-3 modules commented out (validation_framework, meta-learning, multi-modal text, continual learning) | Re-enable incrementally, one module per PR with integration test | M |
| G5 | Medium | rtx-tts | Disabled: rtx_nn API surface changed after rtx-tts was written | Audit rtx-nn current API; update import paths and type sigs | S |
| G6 | Low | rtx-tensor/storage | 2 GPU memory access tests in TDD red phase | Implement read/write GPU memory helpers | S |
| G7 | Low | rtx-bindings, rtx-onnx, rtx-onnx-codegen | ONNX model export incomplete (todo!("Constant tensor")); DLPack zero-copy not implemented |
Full ONNX export via rtx-graph IR; DLPack DLManagedTensor |
L |
| G8 | Low | rtx-timeseries, rtx-multimodal | Quantum/neuromorphic TODO stubs in non-hot-path modules | Implement or remove — no dead stubs in production | S-M |
Effort key: S = days, M = 1-2 weeks, L = 3-4 weeks
5. Prioritized Roadmap
Phase 1 — Critical & Blocking Fixes (Weeks 1-2)
Goal: Eliminate runtime panics in the backend dispatch path; restore CI integrity; re-enable Python ecosystem.
| # | Action | Gap | Owner crates | Verification |
|---|---|---|---|---|
| 1 | Implement 15 missing backend ops in rtx-backend-cuda | G0 | rtx-backend-cuda | All 15 ops covered by tests; cargo test -F cuda passes |
| 2 | Mirror same 15 ops to rtx-backend-rocm and rtx-backend-sycl | G0 | rtx-backend-rocm, rtx-backend-sycl | Feature-gated tests pass |
| 3 | Update PyO3 to ≥0.22, re-enable python feature in rtx-bindings |
G1 | rtx-bindings | cargo test -F python; maturin build succeeds |
| 4 | Add maturin build job to .gitea/workflows/ci.yml |
G1 | CI | Wheel builds on push to main |
| 5 | Make RNCCL optional in rtx-distributed; re-add to workspace | G2 | rtx-distributed | cargo test (CPU path) runs in CI; all 439 tests pass |
| 6 | Fix rtx-tts API mismatch with rtx-nn | G5 | rtx-tts | cargo test for rtx-tts passes |
Phase 2 — Technical Debt Reduction (Weeks 3-4)
Goal: Eliminate async overhead, restore deferred modules, complete TDD red phases.
| # | Action | Gap | Verification |
|---|---|---|---|
| 7 | Async sweep — 97 files | G3 | cargo clippy -- -W clippy::async_yields_async returns zero warnings |
| 8 | Re-enable rtx-transformers validation_framework module |
G4 | Integration test passes |
| 9 | Re-enable rtx-transformers meta-learning module |
G4 | Integration test passes |
| 10 | Implement GPU memory access helpers in rtx-tensor/storage | G6 | 2 RED tests flip to green |
| 11 | Resolve quantum/neuromorphic TODOs in rtx-timeseries and rtx-multimodal (implement or remove) | G8 | No // TODO stubs in non-test code |
Phase 3 — 20-Demo Push (Weeks 5-10)
Goal: Complete the 20 next-generation demo applications per resume_nextplan.md.
Phase 3A — Medical (Weeks 5-6)
- DrugBinder: molecular docking with GNN + RL
- CardioSim: cardiac simulation with FEM coupling
- TumorBoard AI: multi-modal oncology decision support
Phase 3B — Finance (Weeks 6-7)
- QuantumPort: portfolio optimization with quantum-inspired algorithms
- MarketSim: market simulation with agent-based modeling
- RiskFlow: risk factor analysis with normalizing flows
- AlgoArena: algorithmic trading benchmarking
Phase 3C — AI/ML Core (Weeks 7-8)
- NeuralOp Studio: interactive FNO/DeepONet playground
- WorldGen: video diffusion for world model training
- EmbodiedSim: robotics simulation with physics coupling
- FoundationForge: foundation model training harness
Phase 3D — Engineering (Weeks 8-9)
- AeroFlow: aerodynamics CFD with GPU acceleration
- WeatherCast: NWP with neural operators
- SeismicAI: seismic inversion with PINNs
- StructuralPINN: structural mechanics with physics-informed networks
Phase 3E — Cluster (Weeks 9-10)
- ClusterViz: distributed training monitoring dashboard
- DistributedLLM: multi-node LLM training benchmark
- FederatedMed: federated medical imaging across hospital networks
Phase 4 — Ecosystem Maturity (Weeks 11-12)
Goal: Production-grade interoperability, documentation, and CI completeness.
| # | Action | Gap | Verification |
|---|---|---|---|
| 12 | Full ONNX model export: rtx-graph IR → protobuf emission | G7 | Round-trip test: train rtx model, export ONNX, run via ort session |
| 13 | DLPack zero-copy: implement DLManagedTensor structs |
G7 | Array interchange test with numpy |
| 14 | Re-enable rtx-transformers Phase 3 modules (multi-modal text, continual learning) | G4 | Integration tests pass |
| 15 | Add WASM build + size-check CI job for rtx-wasm-inference | — | WASM bundle <5MB (gzipped); wasm-pack test passes |
| 16 | mdBook: complete all 8 layer chapters with API examples | — | mdbook build succeeds; no broken links |
| 17 | rtx-neuro-python: add maturin build to CI | G1 | Python wheel installable; EEG demo runs |
Appendix A: Test Density by Crate
Crates with the highest test investment (top 15 by #[test] count):
| Rank | Crate | #[test] count |
|---|---|---|
| 1 | rtx-transformers | 1,845 |
| 2 | rtx-tensor | 1,047+ (src + tests) |
| 3 | rtx-distributed | 439 |
| 4 | rtx-fea | 371 (src) + 176 (tests) |
| 5 | rtx-nn | 334 |
| 6 | rtx-neuro-core | 266 |
| 7 | rtx-autograd | 223 (tests) + 128 (src) |
| 8 | rtx-metal | 201 |
| 9 | rtx-runtime | 192 (src) + 190 (tests) |
| 10 | rtx-memory | 190 |
| 11 | rtx-tts | 185 (currently excluded) |
| 12 | rtx-preprocessing | 184 (tests) |
| 13 | rtx-tensor (tests/) | 184 |
| 14 | rtx-interpret | 183 (src) + 145 (tests) |
| 15 | rtx-cfd | 179 (src) + 142 (tests) |
Appendix B: Unimplemented Stub Inventory
All 51 todo!/unimplemented! occurrences:
| Count | Location | Operations |
|---|---|---|
| 15 | rtx-backend-cuda/src/lib.rs |
sin, cos, pow, clamp, relu, sigmoid, tanh, var, var_dim, leaky_relu, elu, gt_scalar, conv2d, max_pool2d, avg_pool2d |
| 15 | rtx-backend-rocm/src/lib.rs |
(same 15 ops) |
| 15 | rtx-backend-sycl/src/lib.rs |
(same 15 ops) |
| 4 | rtx-fea/tests/ |
Test-only stubs (not hot path) |
| 1 | rtx-transformers/src/layers/geglu.rs |
Numerical gradient checking (TDD red phase) |
| 1 | rtx-onnx-codegen/src/ops/mod.rs |
todo!("Constant tensor") in ONNX codegen |
Total: 51 — all non-test stubs are in backend op implementations or ONNX codegen.
Appendix C: CI Workflow Coverage
| Workflow | Triggers | What it checks |
|---|---|---|
| ci.yml | Push/PR | rustfmt, clippy (pedantic), build Linux+macOS, cargo test |
| gpu-tests.yml | Push to main | CUDA kernel tests (self-hosted GPU runner) |
| benchmarks.yml | Manual / release | Criterion benchmarks, performance regression |
| docs.yml | Push to main | mdbook build, doc links |
| release.yml | Tag push | Build + publish crates |
Gap: No WASM build job. No maturin/PyO3 wheel build job. rtx-distributed excluded from all jobs.
Spec self-review: no TBD/TODO placeholders; all 8 gaps in registry; all 5 layers covered; roadmap has 17 concrete action items across 4 phases.