# RustyTorch++ Technical Context ## Current Status: Rust 2024 Edition Migration Complete **Last Updated**: 2025-12-16 **Rust Edition**: 2024 (Rust 1.92+ nightly required) **Build Status**: ✅ `cargo check --workspace` passes (0 errors) ## Core Technologies ### Programming Languages - **Rust** (nightly 1.92+, 2024 edition) - Primary implementation language - Required features: const generics, GATs, async traits - Compilation targets: x86_64, aarch64, wasm32 - **Float comparisons**: Uses `total_cmp()` for NaN safety (Rust 2024 requirement) - **C++** (C++20) - CUDA kernel implementation - High-performance CPU kernels - Legacy system integration - **Python** (3.8+) - User-facing API bindings - Testing and benchmarking - Documentation examples ### GPU Technologies & Targets #### Primary Target: RTX 5090 (sm_120) - **Architecture**: NVIDIA Hopper/Ada successor - **Compiler**: rustg at `/home/osobh/projects/rust/rustg` - **Kernel Cache**: `target/kernel_cache/` versioned by arch #### CUDA (12.0+) ```toml [workspace.dependencies] # GPU acceleration (cuda-12060 for CUDA 12.6 compatibility) # f16 feature enables native FP16/BF16 GEMM with Tensor Cores (4-16x faster) cudarc = { version = "0.18.1", features = ["std", "driver", "runtime", "nvrtc", "cublas", "cublaslt", "nccl", "cudnn", "cusparse", "cusolver", "cufile", "curand", "cuda-12060", "f16"] } ``` - Kernel compilation via rustg + nvcc - CUDA Graphs for capture/replay - cuDNN integration for optimized ops - NCCL for multi-GPU communication - **cudarc 0.18.1**: Modern CUDA bindings with FP16/BF16 support #### ROCm (5.0+) ```toml [dependencies] rocm-sys = "0.1" hip = "0.2" # HIP runtime wrapper ``` - HIP kernel compilation via rustg - MIOpen for optimized operations - RCCL for AMD GPU communication #### Metal (macOS 13+) - Deferred ```toml [dependencies] metal = "0.27" metal-rs = "0.2" ``` - Metal Performance Shaders - Unified memory architecture - Darwin apps deferred to later phases ### Build System #### Cargo Configuration ```toml [workspace] members = [ "crates/rtx-compiler", "crates/rtx-runtime", "crates/rtx-kernel", "crates/rtx-tensor", "crates/rtx-autograd", "crates/rtx-ir", "crates/rtx-dist", "crates/rtx-synth", "crates/rtx-serve", "crates/rtx-evolve", "crates/rtx-profiler", "crates/rtx-bench", "crates/rtx-governance", ] [workspace.dependencies] rustg = { path = "../rust/rustg" } [profile.release] lto = "fat" codegen-units = 1 opt-level = 3 debug = false strip = true [profile.bench] inherits = "release" debug = true ``` #### Custom Build Scripts - `build.rs` for GPU kernel compilation - bindgen for C++ interop - CUDA/ROCm detection and configuration ### Dependencies #### Core Crates ```toml [dependencies] # Async runtime tokio = { version = "1.35", features = ["full"] } # Serialization serde = { version = "1.0", features = ["derive"] } bincode = "1.3" # Numerics ndarray = "0.15" num-traits = "0.2" half = "2.3" # f16/bf16 support # Error handling thiserror = "1.0" anyhow = "1.0" # Logging/Tracing tracing = "0.1" tracing-subscriber = "0.3" # Testing proptest = "1.4" criterion = "0.5" ``` #### FFI & Bindings ```toml [dependencies] # Python bindings pyo3 = { version = "0.20", features = ["extension-module"] } # C++ interop cxx = "1.0" bindgen = "0.69" # WebAssembly wasm-bindgen = "0.2" web-sys = "0.3" ``` ### Development Environment #### Required Tools ```bash # Rust toolchain rustup toolchain install nightly-2024-01-01 rustup component add rustfmt clippy miri # GPU toolkits # NVIDIA: CUDA Toolkit 12.0+ # AMD: ROCm 5.0+ # Apple: Xcode 15+ with Metal # Python environment python3 -m venv venv pip install maturin pytest numpy # Benchmarking cargo install cargo-flamegraph cargo install cargo-criterion ``` #### IDE Setup - **VS Code** with rust-analyzer - **CLion** with Rust plugin - **Neovim** with rust-tools.nvim ### Testing Infrastructure #### Test Frameworks ```rust // Unit tests #[cfg(test)] mod tests { use super::*; #[test] fn test_tensor_creation() { let tensor = Tensor::zeros([32, 64]); assert_eq!(tensor.shape(), &[32, 64]); } } // Property-based tests #[cfg(test)] mod prop_tests { use proptest::prelude::*; proptest! { #[test] fn test_matmul_associative( a in tensor_strategy(), b in tensor_strategy(), c in tensor_strategy() ) { // Test (A * B) * C == A * (B * C) } } } // Benchmarks #[bench] fn bench_matmul(b: &mut Bencher) { let a = Tensor::randn([1024, 1024]); let b = Tensor::randn([1024, 1024]); b.iter(|| a.matmul(&b)); } ``` #### CI/CD Pipeline ```yaml # .github/workflows/ci.yml name: CI on: [push, pull_request] jobs: test: strategy: matrix: os: [ubuntu-latest, macos-latest, windows-latest] rust: [stable, nightly] bench: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - run: cargo bench gpu-test: runs-on: [self-hosted, gpu] steps: - run: cargo test --features cuda ``` ### Deployment Targets #### Inference Runtime - **Native Binary**: Standalone executable - **Dynamic Library**: `.so/.dll/.dylib` - **Python Package**: Wheel distribution - **WebAssembly**: Browser/Node.js runtime - **Container**: Docker with GPU support #### Package Distribution ```toml # Cargo.toml [package] name = "rustytorch" version = "0.1.0" edition = "2021" license = "Apache-2.0" repository = "https://github.com/rustytorch/rustytorch" [package.metadata.docs.rs] all-features = true rustdoc-args = ["--cfg", "docsrs"] ``` ### Performance Profiling #### Tools - **perf**: Linux system profiler - **flamegraph**: Visualization of hot paths - **NVIDIA Nsight**: GPU kernel profiling - **AMD ROCProfiler**: ROCm profiling - **Intel VTune**: CPU optimization #### Metrics Collection ```rust use prometheus::{Counter, Histogram, register_counter, register_histogram}; lazy_static! { static ref TENSOR_OPS: Counter = register_counter!( "rustytorch_tensor_ops_total", "Total tensor operations" ).unwrap(); static ref OP_LATENCY: Histogram = register_histogram!( "rustytorch_op_latency_seconds", "Operation latency in seconds" ).unwrap(); } ``` ### Security Considerations #### Memory Safety - MIRI verification for unsafe code - Address sanitizer in debug builds - Fuzzing for input validation #### Supply Chain ```toml # Cargo.toml [dependencies] # Audit dependencies cargo-audit = "0.18" # Vendored dependencies for reproducible builds cargo-vendor = "0.12" ``` ### Documentation #### Generation ```bash # API documentation cargo doc --all-features --no-deps # Book generation mdbook build docs/ # Python API docs pdoc3 --html python/rustytorch ``` #### Hosting - API docs: https://docs.rs/rustytorch - User guide: https://rustytorch.ai/guide - Examples: https://github.com/rustytorch/examples ### Development Workflow #### Pre-commit Hooks ```bash #!/bin/bash # .git/hooks/pre-commit # Format code cargo fmt --all -- --check # Lint cargo clippy --all-targets --all-features -- -D warnings # Test cargo test --lib ``` #### Version Control ```gitignore # .gitignore target/ *.rs.bk Cargo.lock *.pyc __pycache__/ .venv/ .idea/ .vscode/ *.swp ``` ### External Integrations #### Model Formats - **ONNX**: Import/export support (Phase 7) - **SafeTensors**: Fast tensor serialization - **PyTorch**: `.pt` file compatibility (Phase 7) - **DLPack**: Zero-copy interop (Phase 2) #### Cluster Management - **Stratoswarm**: Primary orchestration at `/home/osobh/projects/stratoswarm` - Deployment manifests in `deploy/stratoswarm/` - Blue/green + canary rollouts (Phase 5) - Multi-region support (Phase 9) - **Kubernetes**: Via Stratoswarm integration - **SLURM**: HPC cluster integration (future) - **Ray**: Distributed compute framework (future) ### Technical Constraints #### Platform Limitations - **Linux**: Full feature support - **macOS**: Metal backend only - **Windows**: Limited CUDA support - **WebAssembly**: CPU-only, no threading #### Performance Targets (Phase-Specific) #### Phase 1 Targets - CUDA Graph capture hit-rate: ≥ 70% - Memory fragmentation: < 15% - Step-time improvement: ≥ 20% vs eager baseline - Determinism: fp32 ≤ 1e-6, bf16/fp16 ≤ 1e-3 #### Phase 2 Targets - Graph capture hit-rate: ≥ 75% - End-to-end step time: ≥ 20% reduction vs Phase 1 - Autograd correctness: Parity with fp32 goldens #### Phase 3 Targets - Single-node scaling: ≥ 0.8x efficiency (1→8 GPUs) - Multi-node scaling: ≥ 0.7x on IB/NVLink - FSDP memory reduction: ≥ 40% vs DP #### Phase 4 Targets - Auto-kernel speedup: ≥ 20-40% step-time reduction - Inference speedup: ≥ 1.5x tokens/sec - Kernel cache hit-rate: ≥ 80% #### Phase 5 Targets - Inference latency P99: < 150ms/token on RTX 5090 - KV cache hit-rate: ≥ 85% - Continuous batching utilization: ≥ 30% improvement #### Resource Limits - Max tensor size: 2^63 elements - Max dimensions: 32 - Max GPU memory: Device-dependent - Max nodes: 10,000 (Stratoswarm limit) ## Workspace Configuration ### Excluded Crates (December 2024) | Crate | Reason | Future Work | |-------|--------|-------------| | `integration_tests` | Tests reference unimplemented APIs (400+ errors) | Implement APIs when needed | | `rtx-flash-metal-attention` | macOS/Metal only - not available on Linux | Works on macOS systems | | `demos/ui/src-tauri` | Different MSRV and dependency requirements | Separate build process | ### Dependency Management #### nom Version Consolidation - **Removed**: nom 3.2.1 (legacy) - **Active versions**: nom 7.1.3, nom 8.0.0 - **Method**: Removed unused `npy` dependency from rtx-vision-advanced #### Float Comparison Safety (Rust 2024) ```rust // Before (Rust 2021 - panics on NaN) values.sort_by(|a, b| a.partial_cmp(b).unwrap()); // After (Rust 2024 - NaN-safe) values.sort_by(|a, b| a.total_cmp(b)); ``` **Pattern applied to**: 200+ files across the workspace ### Build Commands ```bash # Standard workspace build cargo check --workspace # Build specific crate cargo build -p rtx-tensor # Full release build cargo build --release --workspace # Run with CUDA features cargo build --workspace --features cuda ``` --- *Technical Context Last Updated: 2025-12-16*